Governance DocsGovernance Docs
Browse Toolkits

CART

No products in the cart.

ISO Compliance Insights & Best Practices

IVDR performance evaluation — scientific validity, analytical and clinical performance

IVDR Performance Evaluation: Scientific Validity, Analytical and Clinical Performance

IVDR performance evaluation is the part of Regulation (EU) 2017/746 with no medical device equivalent to borrow from. There is no clinical evaluation report here, and a file built on that assumption fails structurally rather than in detail.

Article 56(3) requires three separate demonstrations. Annex XIII then requires a report for each, and a fourth report that contains all three. This guide covers what each limb has to establish, the sources the Regulation permits for each, and the sequencing trap that turns an analytical gap into a mandatory clinical study.

Three limbs, three reports

Article 56(3) requires a defined and methodologically sound procedure for the demonstration of:

Limb The question it answers
Scientific validity Is this analyte or marker actually associated with the clinical condition or physiological state?
Analytical performance Does this device measure that analyte correctly?
Clinical performance Do this device’s results correlate with the condition in the intended population?

The data and the conclusions drawn from assessing those three constitute the clinical evidence for the device. That evidence must scientifically demonstrate, by reference to the state of the art in medicine, that the intended clinical benefits will be achieved and that the device is safe.

Annex XIII Section 1.3.2 then sets the structural requirement that catches repurposed files: the performance evaluation report shall include the scientific validity report, the analytical performance report, the clinical performance report and an assessment of those reports. It contains them. A report that cites three others as annexes held elsewhere has not met the section.

Specify the level of evidence first

Article 56(1) requires the manufacturer to specify and justify the level of clinical evidence necessary to demonstrate conformity, appropriate to the characteristics of the device and its intended purpose.

That judgement belongs at the start. Recording it afterwards, calibrated to whatever the data turned out to support, is visible and gets challenged. It is skipped about as universally in IVD files as its counterpart is in medical device files.

Limb one: scientific validity

Scientific validity asks a question that is prior to the device: is this marker associated with the condition at all? A device can measure a marker impeccably and still fail here, if the marker turns out not to be associated with the condition the intended purpose claims.

Annex XIII Section 1.2.1 permits the demonstration to rest on one or a combination of:

  • relevant information on the scientific validity of devices measuring the same analyte or marker;
  • scientific peer-reviewed literature;
  • consensus expert opinions or positions from relevant professional associations;
  • results from proof of concept studies;
  • results from clinical performance studies.

The third source is distinctively IVD and badly under-used. Where a professional association has published a guideline, consensus statement or position on a marker’s clinical utility, that is an expressly permitted source — and for an established marker it is often the strongest evidence available.

The first source is the pragmatic route for a marker other devices already measure. Note carefully what it does and does not carry: it establishes the marker’s association with the condition; it says nothing about this device’s ability to measure it.

The failure mode here is subtler than absence. An association can be validly established and still not support the claimed use — an association strong enough for monitoring an established diagnosis may be nowhere near strong enough for population screening.

Limb two: analytical performance

Annex XIII Section 1.2.2 sets a stronger default than the other two limbs: analytical performance shall always be demonstrated on the basis of analytical performance studies as a general rule. Literature will not do the work here.

The demonstration must cover all the parameters in Annex I Section 9.1(a) unless an omission can be justified as not applicable:

analytical sensitivity; analytical specificity; trueness (bias); precision (repeatability and reproducibility); accuracy; limits of detection and quantitation; measuring range; linearity; cut-off; criteria for specimen collection and handling; control of known relevant endogenous and exogenous interference; and cross-reactions.

Two of those are routinely under-done. Specimen collection and handling criteria are a listed performance parameter, not a labelling afterthought — the claim that a result is valid depends on the specimen having been collected and handled the way the study assumed. And interference and cross-reactivity must cover the known relevant substances for the intended population, which means naming them and justifying the list.

The trueness hierarchy, and where it leads

Annex XIII Section 1.2.2 anticipates that for novel markers, or markers without available certified reference materials or reference measurement procedures, trueness may not be demonstrable. It then sets a hierarchy:

  1. a certified reference material or reference measurement procedure exists — demonstrate trueness against it;
  2. none exists, but a comparative approach can be shown appropriate — for example comparison to another well-documented method, or a composite reference standard;
  3. no such approach is available — a clinical performance study comparing the device to current clinical standard practice is required.

That third position is the provision that catches programmes late. An analytical shortfall escalates into a mandatory clinical study, which changes the timeline, the cost, and possibly whether Articles 57 to 77, an ethics committee and a Member State authorisation are engaged.

It should be established at plan stage. Annex XIII Section 1.1 requires the performance evaluation plan to identify the certified reference materials and reference measurement procedures in the first place, which is exactly the point at which the answer becomes knowable.

Metrological traceability

Annex I Section 9.3 requires the metrological traceability of values assigned to calibrators and control materials to be assured through suitable reference measurement procedures or reference materials of a higher metrological order — and, where available, to certified reference materials or reference measurement procedures.

Two obligations in order, and the second is the one that gets missed: choosing an in-house standard when a certified one exists does not satisfy it. EN ISO 17511 is the harmonised standard behind this, and it has no medical device counterpart at all.

It also has a labelling consequence. Annex I Section 20.4.1(u) requires the instructions for use to state the traceability, identify the reference materials or higher-order procedures applied, and give the maximum self-allowed batch-to-batch variation with figures and units. That published number is one the manufacturing process then has to hold, so it should be set where it is sustainable rather than where it is aspirational.

Limb three: clinical performance

Article 56(4) and Annex XIII Section 1.2.3 both say the same thing: clinical performance studies shall be carried out unless it is duly justified to rely on other sources. The default is a study, and not running one is a decision that has to be justified in writing.

The permitted alternatives are exactly two: scientific peer-reviewed literature, and published experience gained by routine diagnostic testing. The second is distinctively IVD and under-used — but note the word published. Internal experience from your own customers is post-market data, not this.

The parameters come from Annex I Section 9.1(b): diagnostic sensitivity; diagnostic specificity; positive predictive value; negative predictive value; likelihood ratio; and expected values in normal and affected populations.

Three things that inflate a clinical performance claim

Prevalence. Predictive values depend on it. A device studied in an enriched population and claimed for a screening population will have predictive values in use that its study never demonstrated. Record the prevalence in the studied population and in the intended-use population, separately.

Cut-off derivation. A cut-off optimised on the same dataset that then reports the performance overstates it. Where that has happened, say so and report the independent estimate.

Discordant-result resolution. This is where honest studies produce dishonest numbers. Record the pre-specified rule, not the rule that emerged once the discordances were visible.

Report point estimates with confidence intervals throughout. A sensitivity of 95% from forty samples and a sensitivity of 95% from two thousand are different claims, and only the interval shows it.

Sequencing: the limbs run in order

Article 58(5) makes this explicit for anyone planning to generate evidence in parallel. In the case of clinical performance studies, the analytical performance must already have been demonstrated. In the case of interventional clinical performance studies, the analytical performance and the scientific validity must both have been.

So the order is fixed: scientific validity and analytical performance before clinical performance, and both before an interventional study. A programme that plans all three at once has a sequencing problem the Regulation will surface at authorisation.

The literature search is three artefacts, not one

Annex XIII Section 1.2 sets one method across all three limbs: identify available data through a systematic scientific literature review — and identify any remaining unaddressed issues or gaps; appraise all relevant data for suitability; generate new data to address what is outstanding.

The gaps output is the half that gets dropped. A review concluding only that data exists has done half the job the Annex asks for.

And Section 1.3.2 requires the performance evaluation report to contain the literature search methodology, the literature search protocol and the literature search report — three separate things. “A literature review was performed” satisfies none of them.

It does not close at certification

Article 56(5) requires the performance evaluation and its documentation to be updated throughout the lifecycle with data from post-market performance follow-up and from the post-market surveillance plan.

Article 56(6) adds a hard cycle: the performance evaluation report for class C and class D devices shall be updated when necessary, but at least annually. A class C or D report bearing only its certification-date version is a finding on its face. The Article 29 summary of safety and performance must be updated as soon as possible where necessary — a faster standard than the annual cycle it sits beside.

The update trigger most often missed is not about the device at all. Clinical evidence is judged against the state of the art in medicine, and the state of the art moves without any action by the manufacturer. A competitor’s published improvement changes the benchmark your evidence is measured against.

PMPF, and where the justification lives

Post-market performance follow-up is a continuous process that updates the performance evaluation, addressed in the post-market surveillance plan. Annex XIII Part B Section 5.1 sets five aims and Section 5.2 sets eight required plan contents — including specific methods such as ring trials and other quality assurance activities, epidemiological studies, patient or disease registers, genetic databanks and post-market clinical performance studies.

Ring trials and external quality assessment schemes deserve more attention than they get. Where a device is used in laboratories participating in EQA, comparative performance data on it already exists — collected by someone else, at scale, under controlled conditions.

Where PMPF is not appropriate, the justification has to be written in two places and packs routinely record it in one. Annex XIII Part B Section 8 requires it within the performance evaluation report. Annex III Section 1(b) separately requires the post-market surveillance plan to cover a PMPF plan or a justification as to why a PMPF is not applicable. Two documents, one position — and an inconsistency between them is worse than either omission alone.

A checklist for the report

Annex XIII Section 1.3.2 lists what the performance evaluation report shall contain in particular. Work through it directly:

  • the justification for the approach taken to gather the clinical evidence;
  • the literature search methodology, protocol and report;
  • the technology the device is based on, the intended purpose, and any claims made about performance or safety;
  • the nature and extent of the scientific validity and analytical and clinical performance data evaluated;
  • the clinical evidence as the acceptable performances against the state of the art in medicine;
  • any new conclusions derived from PMPF.

The third item is quietly an audit of your own marketing. A performance claim that appears in a brochure but nowhere in the three limb reports is an unevidenced claim, and this is the document in which that becomes visible.

Our EU IVDR Toolkit ships this structure as nine documents — the procedure, the plan with all twelve Annex XIII Section 1.1 elements, the three limb reports, the report that assembles them, the PMPF plan and evaluation report, and a data appraisal log. For the wider picture see our guide to the EU IVDR, and for the transitional position, IVDR transition deadlines.

Frequently asked questions

What is IVDR performance evaluation?

The continuous process under Article 56 and Annex XIII of Regulation (EU) 2017/746 by which data are assessed and analysed to demonstrate scientific validity, analytical performance and clinical performance. The data and the conclusions drawn from all three together constitute the clinical evidence for the device.

Is a performance evaluation report the same as a clinical evaluation report?

No. A clinical evaluation report is a medical device concept under Regulation (EU) 2017/745. An IVDR performance evaluation report must contain three separate reports — scientific validity, analytical performance and clinical performance — plus an assessment of them, under Annex XIII Section 1.3.2.

Do I always need a clinical performance study?

Article 56(4) makes a study the default. Relying instead on peer-reviewed literature or published experience from routine diagnostic testing is permitted, but requires a documented justification. Separately, a study becomes mandatory for analytical reasons where trueness cannot be demonstrated because no certified reference material, reference measurement procedure or comparative approach is available.

How often must a performance evaluation report be updated?

For class C and class D devices, when necessary but at least annually under Article 56(6). For class A and B there is no fixed interval, but the report must still be updated throughout the lifecycle with PMPF and post-market surveillance data under Article 56(5).

What is scientific validity?

The association of an analyte or marker with a clinical condition or physiological state. It is demonstrated separately from the device’s own performance, and Annex XIII Section 1.2.1 permits it to rest on data from devices measuring the same analyte, peer-reviewed literature, consensus expert positions from professional associations, proof of concept studies or clinical performance studies.

Stay Compliance-Ready

Get compliance tips, new toolkit releases, and standard updates in your inbox.

We don’t spam! Read our privacy policy for more info.