Oncology lens
Each biomarker result and each stage comes out as a named column, ready to analyze.
| Group | Columns |
|---|---|
| Protein | ER · PR · HER2 · PD-L1 |
| Genomic | EGFR · KRAS · BRAF · ALK · MSI |
| Staging | TNM, clinical and pathologic |
For registries and research teams
Clinical Extract exports your cohort's records, keeps only the variables your protocol approved, and returns a CSV with its data dictionary. It runs from your environment, so patient data never reaches us.
Structured, coded data from your EHR's certified bulk FHIR API: ICD-10-CM, LOINC and RxNorm, ready for analysis.
EHR and EMR data extraction for one study at a time: a bulk FHIR export tool that runs in your environment.
Every certified EHR offers the bulk FHIR API Clinical Extract uses.1
We also handle EHR migration and archiving.
1Your EHR already has a certified bulk-export API. The certification criterion ONC §170.315(g)(10) requires certified EHR technology to export data for a group of patients through HL7 FHIR Bulk Data Access, authorized with SMART Backend Services. Under 45 CFR 170.404(b)(3), developers had to make it available to their customers by December 31, 2022. Your EHR team registers Clinical Extract once as a backend app, and later studies reuse that connection.
1 How it works
All three steps run in your environment. Each study's export covers only the patients and variables that study needs, so healthcare data extraction stays at the minimum necessary.
Your EHR team registers Clinical Extract once as a backend app. It signs a JWT with a key that stays on your server and gets a token with a read scope for each resource type the study needs.
Token request, abridgedPOST /oauth2/token grant_type=client_credentials client_assertion={JWT signed with your key} scope=system/Group.read system/Patient.read system/Condition.read system/Observation.read system/MedicationRequest.read
Pick the cohort and the variables the protocol names. The export covers the study cohort and its resource types only: a 212-patient study exports those 212 patients' records and no one else's. For a known patient list, roster pull matches the people first.
Export kickoffGET /Group/{cohort-id}/$export
?_type=Patient,Condition,
Observation,MedicationRequest
Accept: application/fhir+json
Prefer: respond-async The NDJSON lands in a folder you control. Clinical Extract keeps the variables the protocol approved, flattens them into one analysis-ready CSV and writes data-dictionary.json with each column's description, FHIR source and an example value.
Outputstudy-0142/ study.csv 212 rows data-dictionary.json 24 columns
| Column | Description | FHIR source | Example |
|---|---|---|---|
| patient_ref | Patient resource ID on your FHIR server | Patient.id | eKx3p9Qa |
| birth_year | Year of birth | Patient.birthDate | 1961 |
| gender | Administrative gender | Patient.gender | female |
| dx_code | Primary cancer diagnosis, ICD-10-CM | Condition.code | C50.412 |
| dx_date | Date of diagnosis | Condition. | 2024-03-14 |
| stage_group | AJCC clinical stage group (LOINC 21908-9) | Observation. | IIA |
| er_status | Estrogen receptor status (LOINC 16112-5) | Observation. | Positive |
| pr_status | Progesterone receptor status (LOINC 16113-3) | Observation. | Positive |
| her2_status | HER2 status (LOINC 48676-1) | Observation. | Negative |
| first_chemo | First chemotherapy agent ordered (RxNorm) | MedicationRequest. | paclitaxel |
The table scrolls sideways.
{ "er_status": { "description": "Estrogen receptor status (LOINC 16112-5)", "fhirSource": "Observation.valueCodeableConcept", "example": "Positive" } }
| Row | patient_ref | birth_year | gender | dx_code | dx_date | stage_group | er_status | pr_status | her2_status | first_chemo |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | eKx3p9Qa | 1961 | female | C50.412 | 2024-03-14 | IIA | Positive | Positive | Negative | paclitaxel |
| 2 | b7Tq2LmR | 1954 | female | C50.911 | 2024-05-02 | IIIB | Negative | Negative | Positive | docetaxel |
| 3 | Vn4c8WzE | 1970 | female | C50.212 | 2024-06-21 | IB | Positive | Negative | Negative | |
| 4 | Hm4r1XoP | 1948 | female | C50.511 | 2024-08-09 | IIB | Positive | Positive | Positive | doxorubicin |
The table scrolls sideways.
2 In the box
Count a cohort, match a patient list, get oncology columns and read the data dictionary without filing a request with your EHR team.
Each biomarker result and each stage comes out as a named column, ready to analyze.
| Group | Columns |
|---|---|
| Protein | ER · PR · HER2 · PD-L1 |
| Genomic | EGFR · KRAS · BRAF · ALK · MSI |
| Staging | TNM, clinical and pathologic |
Paste MRNs, or names with dates of birth. Each row comes back with a match grade, and you review only the uncertain ones.
| Pasted row | Match |
|---|---|
| MRN 0048152 | certain |
| Hale, Ruth 1958-04-02 | probable |
| Ortiz, M. 1972-11-19 | possible |
| Chen, Li 1966-07-30 | multiple (2) |
| Novak, Jan 1949-01-05 | none |
Rows you confirm join the study cohort. Multiple and none are held for your review.
Build the study cohort from coded criteria and see the feasibility count before any record is exported.
| Criterion | Patients |
|---|---|
| Seen 2021 to 2025 | 48,210 |
| ICD-10-CM C50.x | 1,904 |
| ER positive, HER2 negative | 1,112 |
| Stage II or III at diagnosis | 640 |
Counts update as new diagnoses and results arrive.
Every column in study.csv has an entry in data-dictionary.json with its description, its FHIR source path and an example value.
| Column | FHIR source |
|---|---|
| dx_code | Condition.code |
| er_status | Observation. |
| first_chemo | MedicationRequest. |
Table 1 above shows ten entries; the file carries every column.
3 Who it is for
A direct EHR feed for your registry. Pull treatment, biomarkers and TNM staging to complete cases, and run follow-back from a roster of MRNs. Cancer registries come first, and coded fields pre-fill the abstraction. Trauma, cardiac (NCDR) and quality registries use the same connection.
Minimum-necessary pulls for one study, without a place in the data warehouse queue. Every extract comes with a data dictionary that shows the IRB and the honest broker which patients and variables each file holds, and where each value came from.
SMART Backend Services auth, Group export, status polling, NDJSON downloads and retries are already built. You run it on your own infrastructure with read-only scopes, and the output schema and the data-dictionary format are documented for your pipeline.
Trial feasibility counts and cohort identification across your EHRs, including Epic Community Connect affiliates. See how many of your patients meet a protocol's criteria before you commit to the study.
4 Data path
Clinical Extract runs on a server you control, next to your EHR. It downloads the export to that server and writes the CSV to storage your analysts already use. Clinical data extraction and EHR data export happen on your servers. We receive no patient data.
Your EHR
Certified bulk FHIR API, on-premises or vendor-hosted
Your environment
Clinical Extract
Your server or cloud tenant
Study folder
Your analysts, your access controls
Releases
We deliver software releases. No data is sent back to us.
5 Evidence
Jones et al. (JAMIA, 2024) studied (g)(10) bulk export at five sites. Epic sites ran at 502 to 2,827 resources per minute, and Cerner above 8,000 per minute. At those rates, exporting a whole EHR takes months.
Clinical Extract requests only the study cohort and the resource types the protocol names, and that export finishes in hours. When a large Group export stalls, it switches to per-patient requests and keeps going. How a bulk FHIR export runs, step by step.
| Series | Value | Basis |
|---|---|---|
| Epic sites, throughput | 502 to 2,827 resources per minute | Reported, Jones et al. |
| Cerner (now Oracle Health), throughput | above 8,000 resources per minute | Reported, Jones et al. |
| 212 patients, the study cohort, 54,111 resources | 19 min to 1.8 h | Arithmetic on Epic’s range, the study’s manifest count |
| 1.2 million patients, the whole EHR, 300 million resources | 74 days to 1.1 years | Arithmetic on Epic’s range, 250 resources per patient |
6 By EHR
The registration, the export and how long a study takes on each platform.
7 Guides
Practical guides for the analyst who needs the data and the engineer who runs the export.
Guide
Bulk FHIR, EHI export, Clarity and Caboodle, SlicerDicer and Cosmos, compared in one sourced table.
Guide
From a patient list to an analysis-ready CSV: MRNs or names with dates of birth in, graded matches and a data dictionary out.
8 Request a demo
Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.
Works with Epic, Oracle Health, eClinicalWorks, athenahealth, MEDITECH, NextGen and Veradigm, through the bulk FHIR API that every certified EHR offers.