July 29, 2026
DeepECG-Tok
Rohan Banerjee, Juliette Beaulieu, Nicolas Dostie, Blandine Mondésert, Ram Ahuja, Alexis Nolin-Lapalme MD-PhD, Gilbert Jabbour, Shreya Shree Srikanth, Guillaume Marquis-Gravel, Olivier Tastet, Achille Sowa, Julia Cadrin-Tourigny MD-PhD, Jacques Delfrate MSc, Robert Avram MD-MSc
A unified 12-lead ECG-language model for interpretation and clinical-endpoint prediction

Automated ECG interpretation has traditionally been divided into narrow models: one model for rhythm classification, another for left ventricular dysfunction, another for atrial fibrillation risk and another for report generation. DeepECG-Tok replaces this fragmented approach with a single instruction-following system that can answer different clinical questions directly from the same 12-lead ECG.

The key idea is to treat an ECG as a language the model can learn. DeepECG-Tok converts continuous cardiac waveforms into compact, discrete tokens, aligns those tokens with clinical text and uses them to condition a medical large language model. The resulting model can generate free-text interpretations, produce structured reports, identify diagnostic findings and estimate clinically important endpoints.

From cardiac waveforms to language tokens

DeepECG-Tok uses a three-stage training framework:

  1. ECG tokenization: A residual vector-quantization model called QINCo compresses a 12-lead ECG into discrete, morphology-preserving tokens while a decoder reconstructs the original waveform.
  2. ECG-to-text alignment: A Q-Former learns to align the frozen ECG tokens with paired clinical reports.
  3. Instruction tuning: The aligned ECG representation conditions MedGemma 4B-IT, enabling the model to respond to natural-language clinical prompts.

The tokenizer was trained using 1.91 million ECGs from the Montreal Heart Institute, MIMIC-IV and CODE-15. Its frozen representations achieved a macro-averaged AUROC of 0.96 across 77 diagnostic conditions, outperforming the supervised and self-supervised DeepECG baselines. The representation transferred without retraining to the CLSA and Harvard-Emory cohorts, where macro-averaged AUROCs remained 0.88 and 0.90.

One model, many clinical questions

The instruction-tuned model was trained on more than 7.27 million question-answer pairs derived from over one million ECGs. The prompts covered free-text interpretation, structured JSON reporting, normal-versus-abnormal classification, rhythm and conduction disorders, ischemia and infarction, chamber enlargement, interval measurements and clinical endpoints.

Unlike conventional ECG classifiers, DeepECG-Tok can respond to questions such as:

This unified design removes the need to train and maintain a separate prediction head for every task.

Representative clinical questions and answers

The supplementary evaluation shows how the same model can move between focused diagnostic questions, structured outputs and open-ended interpretation. The examples below are adapted and shortened from Supplementary Table 20. They include both strong responses and an error to illustrate the model's present capabilities and limitations.

Clinical questionReference answerDeepECG-Tok answerScore
Can you assess for chamber abnormalities?Left atrial enlargement; left ventricular hypertrophyLeft atrial enlargement; left ventricular hypertrophy1.00
Can you identify the cardiac rhythm?Atrial tachycardia; ventricular rate 129 bpm; irregularly irregular rhythm; atrial fibrillation; premature ventricular complexAtrial tachycardia; irregularly irregular rhythm at 129 bpm; atrial fibrillation; premature ventricular complex1.00
Are there any fascicular blocks?Left anterior fascicular block; right bundle branch block; left axis deviationLeft anterior fascicular block; right bundle branch block; left axis deviation1.00
Can you identify any signs of myocardial injury?Widespread ST depression and T-wave inversion involving inferior, lateral, anterior and septal territoriesCorrectly identifies the principal ST-depression and T-wave-inversion pattern1.00
Are the T waves normal in morphology?No; T-wave inversions in inferior, lateral, anterior and septal leadsNo; T-wave inversions in inferior, lateral, anterior and septal leads1.00
Is urgent action needed for this ECG?Pathological; previous anterior and lateral myocardial infarction with T-wave inversionPathological; previous anterior myocardial infarction and lateral T-wave inversion1.00
Output JSON only with rhythm, conduction, chamber, ischemia, heart rate and overall classificationBradycardia with sinus rhythm and PVC; left anterior fascicular and first-degree AV block; LVH; Q waves and lateral ST depression; heart rate 56; pathologicalCaptures the major findings, but omits the PVC and one anterior Q-wave finding; heart rate 57; pathological0.94
Is there atrial enlargement?Yes; left atrial enlargementNo; no atrial enlargement0.00

These examples are illustrative rather than a performance sample. They show that DeepECG-Tok can reproduce complex multi-finding interpretations and constrained JSON-like outputs, while also making clinically meaningful errors that require expert oversight.

External validation across four cohorts

DeepECG-Tok was evaluated without retraining on four independent external cohorts comprising 20,755 ECGs from 16,767 patients. These cohorts tested both conventional interpretation and linked clinical endpoints across different institutions, populations and acquisition systems.

Diagnostic transfer remained strong, with frozen-tokenizer macro-averaged AUROCs of 0.88 in CLSA and 0.90 in Harvard-Emory. For left ventricular ejection fraction (LVEF), the model supported both threshold-based detection and continuous estimation:

CohortAUROC for LVEF ≤40%AUROC for LVEF <50%LVEF MAEPearson rICC
Montreal Heart Institute0.80 (0.77-0.82)0.77 (0.75-0.79)8.41 percentage points (8.11-8.73)0.56 (0.52-0.59)0.68 (0.64-0.71)
MIMIC-LVEF0.76 (0.75-0.78)0.73 (0.72-0.74)10.21 percentage points (9.98-10.44)0.46 (0.43-0.48)0.60 (0.58-0.63)
EchoNext0.74 (0.72-0.75)0.71 (0.69-0.73)10.22 percentage points (9.97-10.47)0.41 (0.38-0.44)0.51 (0.47-0.54)

The continuous LVEF estimates were most accurate in the internal cohort and remained informative in both external datasets, although the higher MAE and lower correlation and agreement metrics show a measurable generalization gap. Structural heart disease prediction also transferred from the Montreal Heart Institute to EchoNext with comparable evaluation scores.

Beyond LVEF: AF risk, coronary occlusion and structural disease

The model's clinical-endpoint capabilities extend well beyond ventricular function. From the same ECG representation and without separate task-specific prediction heads, DeepECG-Tok assessed future atrial fibrillation risk, structural heart disease, acute coronary occlusion and the likely culprit coronary artery.

Clinical endpointEvaluation resultClinical interpretation
Five-year incident atrial fibrillationAUROC 0.61; LLM-as-a-judge score 0.63Predicted future AF among patients initially in sinus rhythm. The MHI evaluation included 1,160 events among 3,121 labeled ECGs.
Structural heart diseaseAUROC 0.72 internally; LLM-as-a-judge score 0.67 at MHI and 0.68 at EchoNextPerformance transferred to EchoNext without significant degradation in the generated endpoint response. Sensitivity/specificity were 63.8%/69.7% at MHI and 61.2%/73.8% at EchoNext.
Acute coronary occlusionAUROC 0.85; sensitivity 63.0%; specificity 92.6%Identified angiographically linked acute coronary occlusion in the MHI cohort, correctly detecting 225 of 357 positive cases and 615 of 664 negative cases.
Culprit coronary arteryLLM-as-a-judge score 0.58Answered localization questions about the likely culprit artery from the presenting ECG.

These endpoints are particularly important because they span different clinical time horizons: immediate triage for acute coronary occlusion, detection of existing structural disease and longer-term prediction of incident atrial fibrillation. AF risk, acute coronary occlusion and culprit-artery localization were developed and evaluated using linked MHI data and therefore still require prospective, multicenter validation.

Across the full instruction-following evaluation, the ontology-grounded composite score was 0.71 internally and 0.50-0.53 in the external interpretation cohorts. Endpoint-style questions transferred more consistently than open-ended or rigidly structured report generation, highlighting both the model's generalization and the remaining effect of institutional reporting differences.

A reusable evaluation framework for ECG-language models

Free-text ECG reports cannot be evaluated reliably by word overlap alone. Clinically equivalent statements such as "atrial fibrillation with rapid ventricular response" and "AFib with RVR" may receive poor BLEU, ROUGE or METEOR scores despite expressing the same diagnosis. The study therefore introduced an open-source, ontology-grounded LLM-as-a-judge framework designed as a reusable evaluation standard for ECG-language models.

The framework routes each model response through one of two complementary pathways:

  1. Deterministic evaluation for structured outputs: JSON fields, binary labels and numeric measurements are compared field by field. Clinically defined tolerances include heart rate within 5 beats per minute, PR and QT intervals within 20 milliseconds and LVEF within 5 percentage points.
  2. Semantic evaluation for free-text reports: A deterministic LLM adjudicator operating at temperature zero compares the generated report with the reference interpretation. It uses an ECG ontology containing 847 validated term mappings across 12 diagnostic categories to resolve synonyms and related concepts.

For each response, the judge separates the clinical content into findings that were identified, missed or hallucinated. The score is calculated as identified findings divided by the total number of identified, missed and hallucinated findings, then aggregated across 19 evaluation categories. This makes the score sensitive to both omissions and unsupported additions while avoiding the false penalties produced by literal word matching.

Two board-certified cardiologists independently re-scored a random sample of 500 question-answer pairs using the same finding-level rubric. Agreement between the automated framework and the cardiologists reached a mean Cohen's kappa of 0.82, supporting its use for scalable evaluation. The framework is publicly available at github.com/HeartWise-AI/ECG_LLM_Judge, and the planned PhysioNet release will provide 25,000 ECGs with standardized images, raw signals and clinically validated question-answer pairs for reproducible benchmarking.

Because the evaluator is independent of DeepECG-Tok's architecture, other groups can reuse it to compare ECG-language models, test new prompting strategies and measure cross-institutional generalization using clinically meaningful criteria rather than lexical similarity alone.

Reports evaluated by cardiologists

A separate blinded reader study compared DeepECG-Tok reports with reference clinician reports. In single-report review, the model report was rated better in 32% of comparisons, tied in 35% and rated lower in 33%. In forced-choice review, readers preferred the model in 31% of cases, judged the reports equivalent in 42% and preferred the reference in 28%. Among decided cases, the model was preferred 53% of the time, with no significant difference between model and reference reports.

Toward general-purpose ECG intelligence

DeepECG-Tok shows that a discrete ECG vocabulary can support more than classification. One shared representation can connect cardiac waveforms to natural-language interpretation, structured reasoning and endpoint prediction, providing a foundation for ECG systems that clinicians can query directly.

The study remains retrospective. Prospective, multicenter validation is needed before clinical deployment, particularly for endpoints derived from the Montreal Heart Institute cohort. The model also requires an upstream signal-quality gate and a calibrated abstention mechanism: when an ECG is absent or corrupted, structured prompts can still elicit a plausible but unsupported answer. These safeguards are important next steps toward safe clinical use.

Open evaluation resources

The anonymized ECG evaluation dataset supporting the LLM-as-a-judge analysis is planned for release on PhysioNet following peer-reviewed publication, subject to credentialing and a data-use agreement.

This work was supported by the Fonds de recherche du Québec through grant FRQ 5232 and by FRQS career awards.