Why the CER fails more SaMD applications than any other document
In every technical file we audit at TrustedTraceMed, the Clinical Evaluation Report is the section most likely to contain major gaps. Not because it is the hardest document to write — but because most SaMD teams misunderstand what is required.
The most common failure: treating the CER as a literature summary. Finding a few papers showing that algorithms like yours work, citing them, and calling it a clinical evaluation. This is not a CER. It is a bibliography. A real CER under MDCG 2020-1 is a structured, systematic evaluation of clinical evidence that demonstrates your specific device is safe and performs as intended for its specific intended use.
The MDCG 2020-1 structure for software clinical evaluation
Step 1: Define the intended use precisely
The clinical evaluation begins and ends with the intended use. Every piece of evidence you collect must be evaluated against the specific intended use, patient population, user, and clinical context you have defined. If your intended use is vague — "a software tool to support clinical decision-making" — you cannot evaluate whether your evidence is sufficient, because you have not defined what you are trying to demonstrate.
Intended use for SaMD must specify: the medical condition or disease targeted, the patient population (age range, disease stage, contraindications), the intended users (trained clinicians, general practitioners, patients), the clinical context (point of care, hospital, home use), and the specific clinical output (diagnosis, prognosis, treatment planning, monitoring).
Step 2: Identify and evaluate clinical evidence
MDCG 2020-1 identifies three types of clinical data for SaMD:
- Clinical performance data from your specific device: Performance studies using your actual software on relevant patient populations. This is the highest-quality evidence and carries the most weight. It includes validation studies, retrospective analysis of clinical outcomes, and registry data.
- Literature data: Published evidence from equivalent or similar devices, or from the clinical domain underpinning your software's algorithm. This requires a systematic literature search with documented search strategy, inclusion/exclusion criteria, and quality assessment of included papers.
- Analytical performance data: Algorithm validation data demonstrating that your software's outputs are accurate, reproducible, and robust. For diagnostic AI, this includes sensitivity, specificity, AUC, and calibration data from appropriately designed validation datasets.
Step 3: Assess clinical equivalence (with caution)
EU MDR allows clinical evaluation by equivalence — demonstrating that your device is clinically, technically, and biologically equivalent to a device with established clinical evidence. However, for SaMD, equivalence is extremely difficult to establish under MDR for three reasons:
First, you need access to the equivalent device's technical documentation — which typically requires a formal legal agreement with the competitor manufacturer. Few competitors will agree to this. Second, technical equivalence for software requires that algorithms are substantially similar, not just in the same clinical domain. Third, NBs are significantly more sceptical of equivalence claims under MDR than they were under MDD. Unless you have a strong equivalence case with documented access to the predicate's technical file, build your CER on direct clinical evidence for your device.
Step 4: Benefit-risk analysis
The CER must conclude with a structured benefit-risk analysis that explicitly states: the clinical benefits demonstrated by your evidence, the residual risks identified in your risk management file, and the conclusion that benefits outweigh risks for the intended use. This conclusion must be specific — not a generic statement that the benefits outweigh risks, but a quantified or well-substantiated argument that the performance demonstrated by your clinical evidence is sufficient for the clinical context.
Clinical evidence for AI medical software — specific requirements
AI and machine learning SaMD creates additional clinical evaluation challenges that MDCG 2020-1 and the emerging AI Act guidance are beginning to address explicitly.
Training and validation dataset documentation: AI SaMD must document the datasets used for training and validation, including size, demographic composition, labelling methodology, and the source clinical settings. NBs want to know whether your validation dataset is representative of the patient population your device will be used on. A model validated only on data from one tertiary hospital may not generalise to the community settings where it will be deployed.
Performance in subgroups: Clinical evaluation for AI SaMD should include analysis of algorithm performance in clinically relevant subgroups — by age, sex, disease severity, and device/imaging equipment. Performance disparities across subgroups are both a clinical safety issue and an EU AI Act concern (algorithmic bias obligations).
Continuous learning models: For AI models that update based on post-deployment data, the CER must address how the clinical evaluation framework will be applied to updated model versions and how significant performance changes will trigger CER updates.
The PMCF plan — the CER's ongoing companion
Under EU MDR, the CER is not a one-time document. It must be updated throughout your device's lifecycle as new clinical evidence accumulates. The Post-Market Clinical Follow-up (PMCF) plan defines how you will systematically collect post-market clinical data — through literature surveillance, registry participation, post-market studies, or analysis of real-world data from your installed base.
For most SaMD, the PMCF plan should include: ongoing systematic literature search on an annual basis, analysis of complaint and vigilance data for clinical performance signals, and a defined threshold for when PMCF data would trigger a major CER update. NBs check that your PMCF plan is specific to your device — generic PMCF plans that do not define data collection methods or analysis frequency generate major non-conformities.