Data Management · Data managers
CuRE Conduct
OMOP mapping, terminology authoring, governance, and the data plane every other product reads from.
What it does
OMOP mapping, terminology authoring, governance, and the data plane every other product reads from. Configure a pipeline, not a services project — FHIR, CRO CSVs, registries, labs, and claims map to one canonical OMOP data plane.
Key capabilities
- FHIR/HL7 → OMOP CDM v5.4 mapping engine
- ColumnMappingTemplate authoring (Mapping Workbench)
- Vocabulary management + concept/vocabulary API (search, hierarchy, multilingual synonyms, semantic search)
- Governed terminology-overlay authoring: your own concepts, typed concept properties, relationships and relationship types, and local source-code mappings
- Draft → steward review → approval lifecycle on every terminology edit, with hash-chained audit and content hashing
- Overlay-aware hierarchy closure — your assertions resolve alongside standard vocabulary without forking it
- Vocabulary release ingest + reconcile register: adopt a new standard release and see what your overlay has to reconcile
- Constrained source_to_concept_map interchange — import and export local mappings in a standard OMOP shape, with Usagi round-trip compatibility
- Governed CDISC Controlled Terminology register (versioned, hash-anchored)
- Governed platform terminology release channels with independent approval; organization-owned study channel selection and immutable run release bindings.
- End-to-end CDASH → SDTM → ADaM → define.xml traceability
- Consent scope + DUA enforcement, provenance/lineage, data enclaves
- AI structural ETL authoring (propose transform rules)
- OHDSI WebAPI-compatible source surface — concept sets and cohort definitions read by existing ATLAS/HADES tooling with no migration
- Clinical-notes NLP into omop.note / note_nlp
- Source-data profiling / scan report
- Vocabulary-refresh / concept-deprecation remediation
- Concept-accuracy QA harness — resolve, validity, and semantic-fit checks over every terminology binding
- Mapping-coverage drift monitoring
- Per-form UDM ↔ OMOP authority transitions with retained source custody
- Central-data-manager query review lane over site CRF data
- Cellular-product (graft) OMOP-extension model — product identity, processing methods, ISBT-128 identifiers
- CAR-T manufacturing + cold-chain projection onto OMOP, with lossless reconstruction
- Typed donor↔recipient person relationships on the EMPI substrate, DUA-gated
- HLA typing vocabulary plane — GL-string parsing and canonicalization, ARS group expansion, serologic equivalents, IMGT/HLA release pinning
- CIBMTR registry crosswalk families as governed mapping content — Forms 2400/2402/2100, HLA-Save, and AML/MM disease pathways
- Hybrid lexical + semantic concept retrieval with LLM re-rank and mandatory abstention, over a full Athena vocabulary load
- Blast-radius-proportional dual approval on mapping-template changes
- OHDSI Data Quality Dashboard checks run via a governed R sidecar, summarized into the pipeline UI
- Hash-sealed mapping-accuracy scoring harness — correct abstention scores as accuracy, over-reach as the costly error
- Accept-audit evidence in the review surface — steward acceptance measured over spot-checked rows with a Wilson confidence interval, reading "not yet measured" rather than a false 0%
- Per-template quality rollup across uploads, and per-row confidence tiers beside every row in the review grid
- Per-row classifier provenance on AI-resolved mappings — each AI-auto proposal records the classifying model and spec version, so a mapping's instrument identity is inspectable rather than inferred
- Synthetic-release operator workbench — inspect governed releases and their attestations, with design-driven creation
What sets it apart
- Configure a pipeline, not a services project — FHIR, CRO CSVs, registries, labs, and claims map to one canonical OMOP data plane.
- Vocabulary-as-a-service, and vocabulary-as-a-workbench: a governed CDISC Controlled Terminology register plus a concept API (search, hierarchy, multilingual synonyms, semantic search) that sibling apps read instead of forking a second copy of your standards — and a governed authoring layer where your terminologists write the extensions that register never had.
- Author terminology instead of filing a ticket for it: local concepts, typed properties, relationships, and source-code mappings are yours to write, each moving through draft → steward review → approval with a hash-chained audit trail, rather than waiting on a vendor's content team.
- Extend the standard without forking it: your assertions live in a tenant-scoped overlay that resolves alongside standard vocabulary in hierarchy closure, so upgrading to a new vocabulary release is a reconcile step against a release register — not a re-do of your local content.
- Your local extensions stay yours. Tenant-authored assertions are never published into the shared OMOP vocabulary and never reach cross-app views or the external WebAPI surface; only platform-curated and authority-ingested content is eligible for promotion into the canonical model.
- The standards/terminology plane of the platform Metadata Repository — CDISC CT, Biomedical Concepts, and the CDASH → SDTM → ADaM → define.xml traceability spine in one governed substrate.
- Governance is in the data plane, not an afterthought: row-level security + consent scope at every read.
- Research enclaves are the cross-org collaboration boundary — a first-class governed workspace (lifecycle-audited, snapshot-pinned) bound to a study or analysis project, where a collaborator reads a consent-scoped, DUA-gated slice of the record, never a copy of it.
- Versioned dataset releases, de-identification/PPRL, lakehouse deployment, and synthetic data run on the same governed plane.
- AI structural ETL authoring profiles an unmapped source and proposes the transform rules, grounded on the tenant's own accepted-mapping history and governed vocabulary.
- Mapping QA asks whether a concept is the RIGHT one, not just whether it resolves — a standing harness classifies every binding as unresolved, superseded, or semantically drifted, catching cases like a hemoglobin field bound to an arterial-blood LOINC that pass a resolve-and-validity check. Findings are reviewer-gated and never auto-applied.
- Migrate registry forms to OMOP progressively: parity-gated authority changes preserve field lineage and frozen historical report rows without an uncontrolled dual-write system.
- Workbench UI for data managers + service API for sibling apps — dual-nature by design.
- Cell-therapy products are first-class governed records: the graft itself — product identity, processing methods, and ISBT-128 identifiers — is a queryable OMOP-extension entity, not free text on a form (ADR-CON-061).
- CAR-T manufacturing and cold-chain data land on the patient's OMOP timeline and reconstruct back out losslessly, so release-checkpoint measurements, transit times, and temperature excursions align to outcomes (ADR-CON-038).
- Allogeneic donor↔recipient pairs are typed relationships on the EMPI substrate with their own DUA-gated disclosure scope — graft source, matching, and paired outcomes are queryable across the transplant, not a sidebar (ADR-CON-047).
- HLA is a first-class vocabulary plane, not free text: GL strings parse and canonicalize, ARS groups expand, serologic equivalents resolve, and every resolution pins to an IMGT/HLA release — transplant-registry machinery generic mappers don't carry.
- Registry crosswalks ship as governed, versioned mapping content — CIBMTR form families (2400/2402/2100, HLA-Save, AML/MM pathways) with concept-accuracy audits — not as a services engagement.
- Accuracy claims are scored, not asserted: a hash-sealed harness scores correct abstention as accuracy and over-reach as the costly error, so "no confident match" never silently becomes a wrong code.
- The accuracy evidence is visible in-product, not only in a report: steward-acceptance audits with a confidence interval, per-template quality rollups, and per-row confidence tiers sit beside the queue they describe — and a rate that has not been measured reads "not yet measured", never 0%.
Map FHIR to OMOP, field by field
Pick or edit a FHIR Condition resource and watch Conduct's mapper produce the OMOP condition_occurrence row live — resolving the source code to an OMOP standard concept (following Maps to when the source is non-standard), applying the onset-date priority chain, and flagging domain-routing or unmapped codes.
A non-standard ICD-10-CM source code — the resolver follows "Maps to" to the SNOMED standard concept, so source_concept_id ≠ condition_concept_id.
| condition_concept_id | 201826 |
| condition_source_concept_id | 45542738 |
| condition_source_value | http://hl7.org/fhir/sid/icd-10-cm|E11.9 |
| condition_start_date | 2022-11-30 |
| condition_end_date | ∅ |
| condition_type_concept_id | 32817 |
| condition_status_concept_id | 32902 |
| condition_status_source_value | active |
This is Conduct's real Condition mapper running client-side: the onset-date priority chain (onsetDateTime → onsetPeriod.start → recordedDate), the clinicalStatus→concept map, the EHR type concept, and the concept resolver that follows OMOP Maps to so a non-standard source code lands on the right standard concept while its origin is preserved in source_concept_id. Edit the JSON above and the row remaps instantly — no backend.
Why this is more than a toy
This is Conduct's actual Condition mapper, ported dependency-free from the platform monorepo (apps/conduct/…/mapping-engine): the same FHIR-system→OMOP-vocabulary table, vocabulary-priority coding selection, concept resolver, and field-for-field condition_occurrence construction. In the product Conduct owns the canonical FHIR→OMOP mapping (Conduit → Conduct → omop, per ADR-PLT-026) and resolves concepts against the full OMOP vocabularies; here it runs against a hardcoded vocabulary slice, entirely in your browser. No backend.
See CuRE Conduct in action
Every research ecosystem is unique. Let's discuss how CuRE can be configured for your needs.