What goes in an EU AI Act Annex IV evidence pack

Annex IV of the EU AI Act lists the technical documentation a provider of a high-risk AI system must draw up before placing the system on the market and keep up to date afterwards (Article 11). Most teams treat it as a checklist of documents. That is a mistake: what a market-surveillance authority, a notified body, or your own internal audit will actually test is whether the documentation is evidence, traceable to the real system, dated, versioned, and hard to quietly rewrite. This guide walks through what Annex IV asks for and how to assemble it as a single, defensible evidence pack.

Who needs this, and when

The obligation sits with providers of high-risk AI systems, meaning systems in the Annex III use-case list (recruitment, credit scoring, insurance pricing, essential services, and others) or safety components regulated under Annex I product legislation. Deployers do not draw up Annex IV documentation, but they will increasingly demand it from their vendors, and financial-sector deployers have parallel duties of their own (see our guide to the AI Act for banks and insurers).

On timing: under the Digital Omnibus (Regulation (EU) 2026/1744, in force since 27 July 2026), the high-risk obligations for stand-alone Annex III systems now apply from 2 December 2027, and from 2 August 2028 for AI embedded in Annex I regulated products. The deferral is breathing room, not a holiday. The prohibitions and AI-literacy duties have applied since February 2025, transparency obligations under Article 50 have applied since 2 August 2026, and supervisors in regulated sectors already expect documented AI governance today. Teams that wait until mid-2027 to start assembling documentation will discover that most of Annex IV cannot be reconstructed after the fact.

The nine things Annex IV asks for

1. A general description of the system

Intended purpose, the provider's name, the system version and how it relates to previous versions, how the system interacts with hardware, software, and other systems, the forms in which it is placed on the market, and instructions for use. For an agentic system, "interaction with other software" includes the tool and API surface the agent can reach. Keep this section short, but keep it versioned: the description must match the system that is actually running.

2. The elements and the development process

This is the largest section: design specifications and architecture, the key design choices and the trade-offs behind them, what the system is optimised for, the data requirements, provenance, labelling, cleaning, and the characteristics of training, validation, and testing sets, and the human oversight measures designed into the system (Article 14). For systems built on a third-party foundation model, document what you control (prompts, retrieval sources, tools, guardrails) and reference the model provider's documentation for what you do not.

3. Monitoring, functioning, and control

The system's capabilities and known limitations, foreseeable sources of risk, the circumstances that may affect performance, and the technical measures that keep a human able to interpret and oversee outputs. Honest limitation statements are protective here; a documentation file that claims no known failure modes reads as untested, not as safe.

4. Appropriateness of the performance metrics

Not just the metrics, but why those metrics are the right ones for the intended purpose and the affected population. If you evaluate a customer-facing agent only on answer accuracy and never on refusal behaviour or disclosure behaviour, expect that gap to be noticed.

5. The risk management system

A description of the Article 9 risk management system: how risks were identified, estimated, and evaluated, and which mitigations were adopted. The practical evidence is a living risk register with dates and owners, linked to the tests that verify each mitigation actually works.

6. Lifecycle changes

A description of relevant changes made through the system's lifecycle. Model swaps, prompt revisions, new tools granted to an agent, retrieval-corpus updates: each is a change that can invalidate earlier test results, so the record of changes and the record of re-testing belong together.

7. Harmonised standards applied

A list of the harmonised standards applied in full or in part, or a description of the alternative solutions used to meet the requirements where no standard was applied.

8. The EU declaration of conformity

A copy of the declaration of conformity itself (Article 47) once the conformity assessment is done.

9. The post-market monitoring plan

The Article 72 plan for how the provider will collect and review experience from real use, incidents, drift, complaints, and feed it back into the risk management system. For agents, this is where continuous behavioural re-testing lives.

From documents to a defensible pack

Four properties separate a folder of Word files from an evidence pack that survives scrutiny:

  • Traceability. Every claim links to its source: the intake declaration, the test run, the reviewer's decision. "Human oversight is ensured" is a sentence; an approval gate observed in a logged test run is evidence.
  • Immutability. Approval decisions and issued packs should be append-only. If your documentation can be edited in place after the fact, its value as evidence collapses.
  • Versioning. The pack should say which system version, which test plan, and which evaluation model produced each result, so a result can be reproduced or at least explained months later.
  • Citations to the regulation. Mapping each section to the obligation it answers saves the reader, and the authority, from doing the mapping themselves.

Common failure modes

The packs that fail review tend to fail the same way: documentation written once for launch and never updated after the model changed; test results that cannot be tied to a specific system version; oversight described in the design section but contradicted by the logs; and risk registers with no dates, no owners, and no link to any verification. All four are process problems, not writing problems, which is why a document-drafting sprint the month before a deadline does not fix them.

How Vidimus assembles this

Vidimus generates an evidence pack from the artefacts the platform already holds: the structured intake (purpose, data sources, tools, oversight design), the deterministic risk classification with its EU AI Act overlay, behavioural test results from adversarial probes run against the live agent with tool-calls observed on the wire, control-by-control checklist outcomes with reviewer overrides, the signed approval decision, and the append-only audit trail. Each probe carries the regulation passage it tests, and each pack is an immutable, numbered version. Agent-specific controls are covered in depth in our control framework for agentic AI.

If you want to see where one of your systems lands before committing to any of this, the free EU AI Act readiness check gives an indicative risk class and the obligations that follow from it, in the browser, with no account. And if you would rather walk through a real evidence pack, contact us. We onboard customers one at a time and sit with you through the first one.