For AI-enabled medical devices, one of the main regulatory difficulties is the gap between what a manufacturer wants to claim about the product and what the evidence can support.
Under the EU MDR, a claim is more than a marketing statement. It shapes the intended purpose, influences qualification and classification, drives the clinical evaluation strategy, affects risk management, and helps determine what will need to be monitored once the device is on the market.
This is where many start-ups developing AIaMD or SaMD products begin to struggle.
Product development often moves faster than evidence planning. Teams focus on technical performance and start talking about better decisions, earlier detection, smoother workflows, less admin and greater consistency. That is perfectly understandable, but these points can quickly turn into clinical claims. Once they do, they need to be supported within the structure of the EU MDR technical file. When that is not recognised early enough, clinical evidence planning often starts too late.
By the time some companies start preparing the Clinical Evaluation, the important decisions have already been made elsewhere. The intended purpose may be too broad, the classification rationale may be weak, the software features may not be clearly tied to the General Safety and Performance Requirements (GSPRs), and the available evidence may not match the claim being made. The result is a file where the documents all exist, but do not support one another particularly well.
From a consulting perspective, that is the real issue to solve. For an EU MDR submission to stand up well, the reviewer needs to be able to follow a clear line from intended purpose to clinical claims, to clinical evidence, to residual risk, to post-market follow-up. Where that line is weak, the file is more likely to generate questions, delay and closer scrutiny from the Notified Body.
A stronger approach is to build the clinical evidence strategy around five practical questions: intended purpose, technical features, benchmarks, risks and quality.
A Mantra Systems x Cyber Alchemy Perspective - Episode 4
This article is published in partnership with Cyber Alchemy. Mantra Systems takes medical devices through UK and EU MDR/ IVDR, from regulatory strategy to technical documentation. Cyber Alchemy focuses on cybersecurity, helping teams develop and evidence security for software-enabled and connected medical devices. Together, we're producing a practical series for MedTech teams: what to build, what to defer, and how to avoid rework when moving between UK, NHS procurement, and EU routes.
1. Intended Purpose: where do classification and evidence begin?
The first question is simple: what is the device intended to do in clinical practice?
For AI-enabled software as a medical device, this matters immediately because the intended purpose is closely tied to qualification, classification and evidence burden. It should define the target population, intended user, use environment, inputs, outputs and clinical role of the software. It should also make clear where the boundaries are.
This is especially important for AI products, because claims often drift. A tool described internally as helping workflow may later be described as helping clinicians identify risk, prioritise cases or support diagnosis. Those shifts affect classification and change the type of clinical evidence that will be expected.
Where the intended purpose is tightly framed, the evidence strategy is easier to control. Where it is vague or commercially stretched, the technical file often starts to pull in different directions. For a Notified Body reviewer, intended purpose is one of the first anchors in the file. If that anchor is unstable, the classification rationale, clinical evaluation and benefit-risk discussion all become harder to defend.
2. Technical features: how do they map to the GSPRs?
The second question is how the device’s technical features relate to the General Safety and Performance Requirements (GSPRs) under the EU MDR.
The GSPRs are set out in Annex I of Regulation (EU) 2017/745. These are the core safety and performance requirements that manufacturers must assess against their device. In practice, the manufacturer needs to review the GSPRs one by one, determine which requirements apply, explain how conformity with each applicable requirement has been achieved, and identify the evidence that supports that conclusion. Where a requirement does not apply, the technical documentation should provide a clear justification.
This assessment is then reflected in the manufacturer’s technical documentation under Annex II, usually through a GSPR checklist or compliance matrix. For each applicable requirement, the file should show the route to conformity and cross-reference the supporting evidence. That evidence may include design documentation, risk management records, software verification and validation, usability engineering, clinical evaluation, labelling, cybersecurity documentation and other relevant materials. Where conformity also depends on post-market activities, supporting documentation may sit within Annex III, which covers post-market surveillance.
For AI-enabled medical devices, this exercise often requires closer attention to software-driven risks and evidence types. Manufacturers need to consider how the device handles input data, how outputs are generated and presented, what review or approval steps are built into the workflow, how usability and human oversight are maintained, how cybersecurity is addressed, and how output reliability is monitored over time. These factors can affect how conformity with the relevant GSPRs is demonstrated, and they often mean that evidence has to be drawn from several parts of the technical file rather than a single test report or checklist entry.
Not sure if your intended purpose is too broad? We'll tell you in 15 minutes
3. Benchmarks: what does good performance actually look like?
The third question is how the manufacturer uses the state-of-the-art (SOTA) to define clinically meaningful performance for the device in its intended context of use.
Under the EU MDR and the relevant clinical evaluation guidance, a SOTA review does more than describe competing products or current clinical practice. It establishes the clinical and technical background against which the subject device should be assessed. This includes the intended purpose of comparable devices or approaches, the benefits and risks already recognised in the field, the main areas of uncertainty, and the safety and performance expectations that are relevant to the subject device.
A well-conducted SOTA review also helps identify clinical outcome parameters reported consistently across the literature, for example, patients’ or HCPs’ satisfaction, earlier detection rate, improved note completeness or reduced time to review. Through critical appraisal and synthesis of the literature, these outcome parameters can then be used to derive safety and performance objectives for the subject device.
For AI-enabled devices, this is especially important because internal validation results alone do not show whether performance is clinically meaningful. The relevant question is whether the device performs at a level that is acceptable in light of current practice, recognised benefits and risks, and the remaining uncertainties within the intended use setting. SOTA, therefore, helps frame both the clinical relevance of the claim and the standard against which the subject device should be judged.
4. Risks: what can go wrong, and how does the evidence address it?
The fourth question is what can go wrong if the device output is incorrect, absent, delayed, misinterpreted or inappropriately acted upon.
For AIaMD and SaMD, the risk management file (according to ISO 14971) must address failure modes specific to algorithmic outputs. These include hazardous situations such as clinically consequential false positives and false negatives, unreliable outputs from out-of-distribution inputs, and hallucinated content in generative models, as well as foreseeable contributing factors including training data bias, distributional shift, model drift, and use-related risks such as automation bias.
Each identified hazard should trace through the risk management process: from hazardous situation and harm estimation, through risk control measures (including algorithmic design choices, interface safeguards and workflow constraints), to residual risk evaluation.
The critical link is that each residual risk must map to a specific type of evidence that demonstrates its acceptability:
- Verification and validation data (sensitivity, specificity, AUC) address false positive/negative risks at the algorithm level.
- Human factors validation (per IEC 62366-1) addresses risks of misinterpretation, over-reliance and unsafe workflow integration.
- Clinical evaluation provides the overall judgement on whether residual risks are acceptable in light of intended clinical benefit.
- PMCF addresses risks that cannot be fully closed pre-market, such as long-term drift, real-world generalisability and rare failure modes.
A technically sound file uses a traceability matrix linking each hazard ID through its risk controls and residual risk to the specific evidence source and endpoint that supports its acceptability. This creates a closed evidence chain in which potential risks are identified and assessed, risk control measures are defined and implemented, residual risks are evaluated, and the supporting verification, validation, clinical evidence or post-market evidence is clearly referenced to show whether the device remains acceptable for its intended use.
Relevant article: Understanding Risk Management for SaMD
5. Quality: can your QMS support the device for its whole lifecycle?
The fifth question is whether the manufacturer has a quality management system (QMS) capable of supporting the device throughout its lifecycle.
Under the EU MDR, the QMS provides the framework for controlling the device across its lifecycle, from design and development through to post-market activities, so that conformity with the applicable regulatory requirements is maintained. In practice, this is often built in line with ISO 13485, which gives manufacturers a recognised framework for controlling design and development, document management, change control, risk management, supplier oversight, corrective and preventive action, and post-market activities. For AI-enabled devices, this is particularly important because software changes, data-related updates, post-market findings and evolving clinical use can all affect safety and performance over time.
An electronic quality management system (eQMS) can help by making these controls more structured and traceable, especially for software-driven products where documentation, version control and cross-functional review need to be maintained consistently. Whether the system is electronic or manual, the key point is that the QMS should support software lifecycle management, change control, PMS and PMCF in a controlled way.
Companion perspective
This article explains how to build a coherent clinical evidence and regulatory strategy for medical device software under EU MDR. Cyber Alchemy’s companion article examines the cybersecurity and governance evidence needed alongside it, including AI boundaries, threat modelling, auditability, rollback, change control and post-market monitoring.
Together, the articles show how clinical, regulatory and cybersecurity evidence can be developed as one traceable, lifecycle-based assurance story. Read Cyber Alchemy’s expert perspective: Security and Governance Evidence for AI Features
Case studies
Here are two case studies about AI-enabled medical devices, which will be helpful if you are developing this type of product.
Case Study 1: Ambient Scribe
An ambient scribe is a good example of how these five questions work together.
Intended purpose
The intended purpose needs to stay tightly framed. If the software is intended to generate a draft note for healthcare professional review, that position should remain clear. If the wording starts to imply diagnostic support or autonomous identification of clinically important findings, the evidence burden rises quickly.
The technical features likely include speech capture, transcription, speaker attribution, summarisation and structured note generation. Those features need to be linked to the relevant GSPRs and assessed in terms of their safety and performance implications.
The benchmark question is less about diagnostic accuracy and more about documentation quality. Relevant benchmarks may include note completeness, transcription accuracy, time reduced for documenting, healthcare professionals’ satisfaction rates, word error rate, and omission rates.
The main risks include clinically significant omissions, hallucinated content, incorrect attribution and over-reliance on draft outputs.
Case Study 2: AI-assisted imaging system for cancer diagnosis
An AI-assisted imaging system for cancer diagnosis sits in a much higher-stakes area.
Intended purpose
The intended purpose should define the imaging modality, target condition, intended user, patient population, clinical setting and role of the software output. Under the EU MDR, this links closely to classification, particularly under Rule 11, and directly affects the level of evidence expected.
Where the software enhances or analyses imaging data to support clinician interpretation, while the final diagnosis remains with the healthcare professional, it may often fall within Class IIa or Class IIb, depending on the significance of the information provided and the consequences if that information is wrong. In this setting, the evidence package is likely to include analytical and clinical performance validation, comparison with clinician interpretation, external validation, usability evidence and clinical evaluation showing that the output supports the intended workflow. Common benchmarks include sensitivity, specificity, false positive and false negative rates, ROC AUCs, reader comparison studies and comparison with current clinical practice.
Where the software generates an automated diagnostic conclusion or drives a clinical decision with limited meaningful clinician interpretation, the classification position may move towards Class III. In that setting, the evidence burden becomes heavier, with stronger clinical performance evidence, clearer justification of clinical benefit, and more extensive human factors and PMCF planning.
The risks are also more direct. False negatives may delay diagnosis, false positives may lead to unnecessary follow-up, and automation bias may affect interpretation. These risks should be reflected in the clinical evaluation, human factors evidence and post-market planning.
Summary
AI-enabled medical devices are hardest to defend under EU MDR when the claim gets ahead of the evidence. Intended purpose is the anchor: it shapes classification, clinical evaluation, risk management, GSPR alignment, and what has to be monitored after launch. The strongest technical files follow a clear line from intended purpose to evidence, residual risk, and post-market follow-up, so a reviewer can see exactly why the device is safe, performs as claimed, and is supported by the right data.
A practical way to structure that evidence is around five regulatory questions:
- What the device is intended to do?
- How its technical features map to the GSPRs?
- What good performance looks like in the real world?
- What can go wrong and how those risks are controlled?
- Whether the QMS can support the product across its lifecycle?
That’s the difference between a file that just contains documents and one that actually holds together under Notified Body review.
Next step: Get your evidence strategy reviewed
Developing an AI-enabled medical device? In a free 30-minute joint review, specialists from Mantra Systems and Cyber Alchemy will assess how well your regulatory, clinical and cybersecurity evidence fits together.
We’ll help you identify the most important gaps across intended purpose, claims, clinical evidence, risk, AI governance and cybersecurity, then clarify what to prioritise next for your market entry.
Book your free 30-minute review call
Articles in this series
- Episode 1 from Mantra Systems
- Where to Launch First? A MedTech Founder's Regulatory Roadmap to the EU, UK and US
- Episode 1 from Cyber Alchemy
- EU MDR & NHS DTAC Cybersecurity Requirements for UK Market Entry
- Episode 2 from Mantra Systems
- How to handle non-conformities and get back on track
- Episode 2 from Cyber Alchemy
- EU MDR, FDA 510(k) and DTAC Cybersecurity Nonconformities: How to Recover
- Episode 3 from Mantra Systems
- Understanding Risk Management for SaMD
- Episode 3 from Cyber Alchemy
- Security Risk Controls in the Medical Device Technical File
- Episode 4 from Mantra Systems
- Preparing AI-Enabled Medical Devices for EU MDR: Building Evidence Around the Right Regulatory Questions
- Episode 4 from Cyber Alchemy
- Security and Governance Evidence for AI Features