For registries and research teams

EHR data extraction for research through your EHR's certified API.

Clinical Extract exports your cohort's records, keeps only the variables your protocol approved, and returns a CSV with its data dictionary. It runs from your environment, so patient data never reaches us.

Structured, coded data from your EHR's certified bulk FHIR API: ICD-10-CM, LOINC and RxNorm, ready for analysis.

Works with Epic, Oracle Health, eClinicalWorks, athenahealth, MEDITECH, NextGen and Veradigm.

EHR and EMR data extraction for one study at a time: a bulk FHIR export tool that runs in your environment. Every certified EHR offers the bulk FHIR API Clinical Extract uses.1
We also handle EHR migration and archiving.

1Your EHR already has a certified bulk-export API. The certification criterion ONC §170.315(g)(10) requires certified EHR technology to export data for a group of patients through HL7 FHIR Bulk Data Access, authorized with SMART Backend Services. Under 45 CFR 170.404(b)(3), developers had to make it available to their customers by December 31, 2022. Your EHR team registers Clinical Extract once as a backend app, and later studies reuse that connection.

1 How it works

Three steps to an analysis-ready CSV.

All three steps run in your environment. Each study's export covers only the patients and variables that study needs, so healthcare data extraction stays at the minimum necessary.

  1. 1.1 Connect with a read-only backend app

    Your EHR team registers Clinical Extract once as a backend app. It signs a JWT with a key that stays on your server and gets a token with a read scope for each resource type the study needs.

    Token request, abridgedPOST /oauth2/token
    grant_type=client_credentials
    client_assertion={JWT signed with your key}
    scope=system/Group.read
          system/Patient.read
          system/Condition.read
          system/Observation.read
          system/MedicationRequest.read
  2. 1.2 Select the minimum necessary

    Pick the cohort and the variables the protocol names. The export covers the study cohort and its resource types only: a 212-patient study exports those 212 patients' records and no one else's. For a known patient list, roster pull matches the people first.

    Export kickoffGET /Group/{cohort-id}/$export
        ?_type=Patient,Condition,
               Observation,MedicationRequest
    Accept: application/fhir+json
    Prefer: respond-async
  3. 1.3 Receive the CSV and its data dictionary

    The NDJSON lands in a folder you control. Clinical Extract keeps the variables the protocol approved, flattens them into one analysis-ready CSV and writes data-dictionary.json with each column's description, FHIR source and an example value.

    Outputstudy-0142/
      study.csv              212 rows
      data-dictionary.json   24 columns
Three steps, all in your environmentThree numbered steps inside a dashed boundary labelled your environment. Your EHR, with its certified bulk FHIR API, sits on the boundary. Step 1, connect: Clinical Extract, a read-only backend app registered once, sends a signed JWT to the EHR and gets back an access token with read scopes only. Step 2, select: a grid of every patient and data element in which only the study cohort rows and the approved variable columns are kept; the export request is Group/{cohort-id}/$export. Step 3, receive: a stapled printout of study.csv, 212 rows, and the crimson coil-bound data dictionary, data-dictionary.json, with one entry per column.YOUR ENVIRONMENT1CONNECT2SELECT3RECEIVEYour EHRcertifiedbulk FHIR APIClinical Extractsigned JWTaccess tokenread-only backend appregistered onceapproved variablesstudy cohortGroup/{cohort-id}/$exportstudy.csv212 rowsdata-dictionary.jsonone entry per column
Three steps, all in your environmentThree numbered steps inside a dashed boundary labelled your environment. Your EHR, with its certified bulk FHIR API, sits on the boundary. Step 1, connect: Clinical Extract, a read-only backend app registered once, sends a signed JWT to the EHR and gets back an access token with read scopes only. Step 2, select: a grid of every patient and data element in which only the study cohort rows and the approved variable columns are kept; the export request is Group/{cohort-id}/$export. Step 3, receive: a stapled printout of study.csv, 212 rows, and the crimson coil-bound data dictionary, data-dictionary.json, with one entry per column.YOUR ENVIRONMENT1CONNECT2SELECT3RECEIVEYour EHRcertified bulk FHIR APIClinical Extractsigned JWTaccess tokenread-only backend appregistered onceapproved variablesstudy cohortGroup/{cohort-id}/$exportstudy.csv212 rowsdata-dictionary.jsonone entry per column
Figure 1. Three steps, all in your environment. Clinical Extract connects to your EHR once as a read-only backend app, exports only the study cohort, keeps only the approved variables, and returns study.csv and data-dictionary.json. The grid is schematic, and the 212-row study is illustrative.
Table 1. data-dictionary.json, rendered: 10 of 24 columns for a breast cancer study. Every value is synthetic.
ColumnDescriptionFHIR sourceExample
patient_refPatient resource ID on your FHIR serverPatient.ideKx3p9Qa
birth_yearYear of birthPatient.birthDate1961
genderAdministrative genderPatient.genderfemale
dx_codePrimary cancer diagnosis, ICD-10-CMCondition.codeC50.412
dx_dateDate of diagnosisCondition.onsetDateTime2024-03-14
stage_groupAJCC clinical stage group (LOINC 21908-9)Observation.valueCodeableConceptIIA
er_statusEstrogen receptor status (LOINC 16112-5)Observation.valueCodeableConceptPositive
pr_statusProgesterone receptor status (LOINC 16113-3)Observation.valueCodeableConceptPositive
her2_statusHER2 status (LOINC 48676-1)Observation.valueCodeableConceptNegative
first_chemoFirst chemotherapy agent ordered (RxNorm)MedicationRequest.medicationCodeableConceptpaclitaxel

The table scrolls sideways.

Listing 1. The er_status entry as it ships in data-dictionary.json.
{
  "er_status": {
    "description": "Estrogen receptor status (LOINC 16112-5)",
    "fhirSource": "Observation.valueCodeableConcept",
    "example": "Positive"
  }
}
Table 2. study.csv, rows 1 to 4 of 212. Same columns as Table 1. Synthetic values.
Row patient_refbirth_yeargenderdx_codedx_datestage_grouper_statuspr_statusher2_statusfirst_chemo
1 eKx3p9Qa1961femaleC50.4122024-03-14IIAPositivePositiveNegativepaclitaxel
2 b7Tq2LmR1954femaleC50.9112024-05-02IIIBNegativeNegativePositivedocetaxel
3 Vn4c8WzE1970femaleC50.2122024-06-21IBPositiveNegativeNegative
4 Hm4r1XoP1948femaleC50.5112024-08-09IIBPositivePositivePositivedoxorubicin

The table scrolls sideways.

2 In the box

Four tools that come with the export.

Count a cohort, match a patient list, get oncology columns and read the data dictionary without filing a request with your EHR team.

Oncology lens

Each biomarker result and each stage comes out as a named column, ready to analyze.

GroupColumns
ProteinER · PR · HER2 · PD-L1
GenomicEGFR · KRAS · BRAF · ALK · MSI
StagingTNM, clinical and pathologic

Oncology lens

Roster pull

Paste MRNs, or names with dates of birth. Each row comes back with a match grade, and you review only the uncertain ones.

Pasted rowMatch
MRN 0048152certain
Hale, Ruth 1958-04-02probable
Ortiz, M. 1972-11-19possible
Chen, Li 1966-07-30multiple (2)
Novak, Jan 1949-01-05none

Rows you confirm join the study cohort. Multiple and none are held for your review.

Roster pull

Cohort builder

Build the study cohort from coded criteria and see the feasibility count before any record is exported.

CriterionPatients
Seen 2021 to 202548,210
ICD-10-CM C50.x1,904
ER positive, HER2 negative1,112
Stage II or III at diagnosis640

Counts update as new diagnoses and results arrive.

Build a cohort

CSV and data dictionary

Every column in study.csv has an entry in data-dictionary.json with its description, its FHIR source path and an example value.

ColumnFHIR source
dx_codeCondition.code
er_statusObservation.valueCodeableConcept
first_chemoMedicationRequest.medicationCodeableConcept

Table 1 above shows ten entries; the file carries every column.

See the output files

3 Who it is for

Who uses Clinical Extract.

Registries

A direct EHR feed for your registry. Pull treatment, biomarkers and TNM staging to complete cases, and run follow-back from a roster of MRNs. Cancer registries come first, and coded fields pre-fill the abstraction. Trauma, cardiac (NCDR) and quality registries use the same connection.

EHR data for registries

Research informatics

Minimum-necessary pulls for one study, without a place in the data warehouse queue. Every extract comes with a data dictionary that shows the IRB and the honest broker which patients and variables each file holds, and where each value came from.

For informatics teams

Data engineering

SMART Backend Services auth, Group export, status polling, NDJSON downloads and retries are already built. You run it on your own infrastructure with read-only scopes, and the output schema and the data-dictionary format are documented for your pipeline.

Documentation

Research sites and physician networks

Trial feasibility counts and cohort identification across your EHRs, including Epic Community Connect affiliates. See how many of your patients meet a protocol's criteria before you commit to the study.

For research sites

4 Data path

Patient data never leaves your network.

Clinical Extract runs on a server you control, next to your EHR. It downloads the export to that server and writes the CSV to storage your analysts already use. Clinical data extraction and EHR data export happen on your servers. We receive no patient data.

Runs in your environment
On your server or in your cloud tenant, next to the EHR it reads from.
Read-only
Read-only scopes, so Clinical Extract cannot change the chart.
PHI stays with you
We never receive, store or see patient data.

Security overview

Your EHR

Certified bulk FHIR API, on-premises or vendor-hosted

Your environment

NDJSON,
read-only

Clinical Extract

Your server or cloud tenant

CSV + data dictionary

Study folder

Your analysts, your access controls

Releases

We deliver software releases. No data is sent back to us.

Figure 2. The data path. Patient data moves only between your EHR and your environment.

5 Evidence

Why a study export finishes in hours.

Jones et al. (JAMIA, 2024) studied (g)(10) bulk export at five sites. Epic sites ran at 502 to 2,827 resources per minute, and Cerner above 8,000 per minute. At those rates, exporting a whole EHR takes months.

Clinical Extract requests only the study cohort and the resource types the protocol names, and that export finishes in hours. When a large Group export stalls, it switches to per-patient requests and keeps going. How a bulk FHIR export runs, step by step.

Study pull vs whole-EHR pullTwo bar charts. Panel a, reported bulk export throughput in resources per minute: Epic sites ran at 502 to 2,827, and Cerner, now Oracle Health, ran above 8,000. Panel b, on a log time scale, the export time those Epic rates imply as arithmetic: the 212-patient study cohort, 54,111 resources by its manifest, takes 19 minutes to 1.8 hours, highlighted, and the whole EHR, 1.2 million patients at 250 resources per patient, 300 million resources, takes 74 days to 1.1 years.aThroughput at Epic and Cerner sitesresources per minute02,5005,0007,50010,000Epic range across Epic sitesEpic sites: 502 to 2,827 resources per minute502 to 2,827Cerner (now Oracle Health)Cerner: above 8,000 resources per minuteabove 8,000bExport time at Epic’s rangelog scale10 min1 h1 day1 wk1 mo1 yr212 patients the study cohort, 54,111 resources212 patients, 54,111 resources: 19 min to 1.8 h19 min to 1.8 h1.2 million patients the whole EHR, 300 million resources1.2 million patients, 300 million resources: 74 days to 1.1 years74 days to 1.1 years
Study pull vs whole-EHR pullTwo bar charts. Panel a, reported bulk export throughput in resources per minute: Epic sites ran at 502 to 2,827, and Cerner, now Oracle Health, ran above 8,000. Panel b, on a log time scale, the export time those Epic rates imply as arithmetic: the 212-patient study cohort, 54,111 resources by its manifest, takes 19 minutes to 1.8 hours, highlighted, and the whole EHR, 1.2 million patients at 250 resources per patient, 300 million resources, takes 74 days to 1.1 years.aThroughput at Epic and Cerner sitesresources per minute05,00010,000Epic range across Epic sitesEpic sites: 502 to 2,827 resources per minute502 to 2,827Cerner (now Oracle Health)Cerner: above 8,000 resources per minuteabove 8,000bExport time at Epic’s rangelog scale10 min1 h1 day1 mo1 yr212 patients the study cohort212 patients, 54,111 resources: 19 min to 1.8 h19 min to 1.8 h1.2 million patients the whole EHR1.2 million patients, 300 million resources: 74 days to 1.1 years74 days to 1.1 years
Figure 3 as a table
SeriesValueBasis
Epic sites, throughput502 to 2,827 resources per minuteReported, Jones et al.
Cerner (now Oracle Health), throughputabove 8,000 resources per minuteReported, Jones et al.
212 patients, the study cohort, 54,111 resources19 min to 1.8 hArithmetic on Epic’s range, the study’s manifest count
1.2 million patients, the whole EHR, 300 million resources74 days to 1.1 yearsArithmetic on Epic’s range, 250 resources per patient
Figure 3. Study pull versus whole-EHR pull. (a) Bulk export throughput reported at Epic and Cerner sites. (b) The export time those rates imply, on a log scale: a 212-patient study cohort exports in under two hours, the whole EHR takes months. Panel b is illustrative arithmetic on Epic’s reported range: the study cohort at its manifest count, the whole EHR at 250 resources per patient. Source for (a): Jones et al. J Am Med Inform Assoc. 2024. doi:10.1093/jamia/ocae040.

6 By EHR

How Clinical Extract connects to each EHR.

The registration, the export and how long a study takes on each platform.

All EHRs

7 Guides

Guides to getting research data out of an EHR.

Practical guides for the analyst who needs the data and the engineer who runs the export.

  1. Guide

    Every route out of Epic, compared

    Bulk FHIR, EHI export, Clarity and Caboodle, SlicerDicer and Cosmos, compared in one sourced table.

  2. Guide

    Running a chart review study with FHIR

    From a patient list to an analysis-ready CSV: MRNs or names with dates of birth in, graded matches and a data dictionary out.

8 Request a demo

See Clinical Extract run on one of your studies.

Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.

  1. We email you to agree on a meeting time.
  2. We walk through your EHR's bulk endpoint and your variable list.
  3. Then we plan a pilot on your study.

Works with Epic, Oracle Health, eClinicalWorks, athenahealth, MEDITECH, NextGen and Veradigm, through the bulk FHIR API that every certified EHR offers.

indicates a required field

EHRs (optional)

We use your details to arrange the demo, and for nothing else.