On 27 July 2026 the AI Omnibus entered into force and moved the high-risk deadlines. Standalone systems in Annex III now apply from 2 December 2027 instead of this month. Systems embedded in regulated products move to 2 August 2028. Most coverage read that as sixteen months of breathing room.
It is not breathing room, and the reason is narrow enough to state in one sentence: the obligation that moved is the date you must be compliant, and the obligation that did not move is the requirement to hold records covering a period that has already started by then.
You cannot generate December 2027's evidence in November 2027. Whatever your agents did in the months before the deadline is either recorded or it is not, and a deferral changes only how much unrecorded history you will be asked about.
Why a deferral to December 2027 is not relief
The default reading is that a delayed obligation is a delayed project. Ship the feature, revisit the compliance work next year, and let the team that inherits it deal with the standard by the time it binds.
That reasoning holds for obligations satisfied by a document. It fails for obligations satisfied by a record, because a document can be written the week before an audit and a record cannot. Article 26(6) requires deployers of high-risk systems to keep the logs the system automatically generates, to the extent those logs are under their control, "for a period appropriate to the intended purpose of the high-risk AI system, of at least six months". Six months of retention means that on the day the obligation binds, you are expected to hold six months of history. That history is being generated now, by systems currently configured to discard it on whatever schedule the log drain defaults to.
The deferral therefore does one thing and not the other. It gives you longer to build the mechanism. It does not give you longer to have been running it.
There is a second, less obvious effect. Sixteen extra months is long enough for the agent estate to change substantially. The tools your agents can call, the credentials they hold, the repositories they write to, and the humans who approved any of it will all have turned over at least once. The record that matters is not only what the agent did but what it was permitted to do at the time, and permission is the part that decays fastest and reconstructs worst.
By the end you will have
- The verified timeline as it stands after the Omnibus, with what moved and what did not
- A test for whether your agents put you in scope at all, stated as deployment context rather than authorship
- The reason the three frameworks do not line up one to one, and what to do about the gaps
- The Agent Traceability Matrix, mapping obligation to framework function to the artefact in your repository that satisfies it
- A rule for what to build first, and an honest list of what none of this solves
What actually changed on 27 July 2026
The amending instrument is Regulation (EU) 2026/1744, which amends the AI Act along with two other regulations. It was proposed in November 2025, agreed politically on 7 May 2026, approved by Parliament in June, given a final green light by the Council on 29 June 2026, and entered into force on 27 July 2026.
| Milestone | Date | Status after the Omnibus |
|---|---|---|
| AI Act entered into force | 1 August 2024 | Unchanged |
| Article 50 transparency obligations apply | 2 August 2026 | Not delayed |
| Commission GPAI enforcement powers | 2 August 2026 | Active |
| Article 50(2) marking, systems on the market before 2 August 2026 | Six-month transitional period | Grandfathered |
| High-risk, standalone (Annex III) | 2 December 2027 | Deferred |
| High-risk, embedded in products (Annex I) | 2 August 2028 | Deferred |
| Deployer log retention, Article 26(6) | At least six months | Unchanged |
Three things about that table are worth stating explicitly, because each is a place where secondhand summaries go wrong.
Entered into force is not applies from. The Act has been in force since 1 August 2024. Its obligations become applicable on staggered dates, and conflating the two is the most common error in writing about this subject. The Omnibus has its own version of the same trap: it entered into force on 27 July 2026, and the dates it sets are in the future.
Article 50 is live right now, with one carve-out worth knowing. Transparency duties — disclosing that a user is interacting with an AI system, marking synthetic content, labelling deepfakes — applied from 2 August 2026. If you shipped a support agent or a content pipeline into the EU market, that obligation is not pending. It is days old at the time of writing and the Commission's enforcement powers over general-purpose models started the same day.
The carve-out is narrow and probably applies to you. Providers of generative AI systems already on the market or put into service before 2 August 2026 get a transitional period of six months for the marking and machine-readable detection obligations in Article 50(2) specifically — the requirement that generated or manipulated output be detectable as such. It is a grandfathering rule for systems that already existed, not a general delay of Article 50, and it does nothing for a system you launch today.
Check the exact conformity deadline against the final text of the amending regulation rather than against any summary, including this one. Secondary sources currently disagree about it, and the disagreement is arithmetic rather than interpretive: a six-month period running from either 2 August 2026 or the 27 July entry into force does not land on the date some of them print. When sources disagree, the disagreement is the finding, and this is a date you would rather get from the Official Journal than from a blog.
A simplification package added a prohibition. The Omnibus introduces a ban on AI systems that generate non-consensual sexually explicit imagery and child sexual abuse material. Coverage that frames the whole instrument as deregulation has told you something false about its direction, and noticing that is a reasonable test of whether a given summary read the thing it is summarising.
Who the deployer is when the agent writes the code
Scope first, because most teams reading this are not in it and should know that before spending a quarter on a traceability programme.
The high-risk classification is about deployment context. It follows from the Annex III list and from safety components of products regulated under Annex I. It does not follow from a model having written the code. An agent refactoring your marketing site is not a high-risk AI system, and no amount of autonomy makes it one.
The distinction that catches people is that the same agent stack serves both cases. The harness you use to refactor a marketing site is the harness someone else in the building uses inside a hiring pipeline, a credit decision, or software that ends up in a medical device. The agent did not change. The deployment context did, and the obligations attach to the context.
So the useful question is not "are my agents high-risk". It is:
- Does any workload this harness touches sit in an Annex III area, or inside a product regulated under Annex I?
- If yes, are the logs that system automatically generates under our control, in the Article 26(6) sense?
- If yes, do we currently hold six months of them, and can we tell which agent run produced which change?
A "no" at the first line is a legitimate exit, and it is worth writing down with a date so the answer gets revisited rather than assumed forever. A "yes" at the first and a "no" at the third is the ordinary state of a competent engineering organisation today, and it is the state this piece is about.
None of this is legal advice. Whether you are in scope is a legal determination about what your software does, made by someone qualified to make it. Take the engineering as a description of what a record has to look like to be worth anything, not as a scope opinion.
The role you occupy decides which obligations attach
Scope has a second axis that engineering teams routinely miss, and getting it wrong produces a matrix aimed at the wrong set of duties.
The Act attaches obligations to roles, not to companies. The same organisation can be a deployer of one system and a provider of another, and the duties differ substantially. Article 26 is a deployer article; the log-retention duty discussed above is a deployer duty. Provider obligations are a different and heavier set.
For an engineering organisation using coding agents, the ordinary position is deployer: you are using a system someone else placed on the market. Three things move you toward provider, and all three are common enough to check rather than assume.
You put your name on it. Placing a system on the market under your own name or trademark is the classic route. If the agent capability is part of what you sell, this is worth a determination rather than an assumption.
You substantially modify a high-risk system. Modification has a specific meaning here, and "we wrote a wrapper" is not automatically it. But a harness that materially changes what a system does, in a high-risk context, is exactly the fact pattern this is about.
You repurpose a system into a high-risk use. Taking something general and putting it to a use listed in Annex III can make you the provider of that high-risk system, even though you built none of the model.
The practical consequence for traceability is that provider duties reach further back into the lifecycle than deployer duties do. A deployer keeps logs of operation. A provider is answerable for the system's design, its documentation, and its behaviour across the market. If your determination lands on provider, the matrix below is a floor rather than a plan, and the correct next step is not more engineering — it is advice from someone qualified to scope it.
This is also the honest reason the scope section comes first and gets a date. Role is not a stable property of your organisation. It is a property of a workload at a point in time, and shipping an agent feature to customers can change it without anyone filing a ticket.
Mechanism dive: why the three frameworks do not line up one to one
There are three instruments in play and they are different kinds of object, which is why a naive mapping produces a matrix full of confident nonsense.
The AI Act is law. It states obligations and attaches them to roles — provider, deployer, importer — and to classifications. It tells you what must be true and says almost nothing about how.
The NIST AI RMF is a voluntary framework organised around four functions: GOVERN, MAP, MEASURE and MANAGE. GOVERN covers oversight and accountability, MAP contextualises risk, MEASURE applies metrics, MANAGE implements treatment. It is a way of organising activity, not a compliance checklist, and it is deliberately jurisdiction-neutral.
ISO/IEC 42001:2023 is a certifiable management system standard — an AI management system, in the same family as ISO 9001 and ISO 27001 — with 38 Annex A controls across nine control areas. It tells you what the management system must contain and leaves the technical implementation open.
NIST already publishes an AI RMF to ISO/IEC 42001 crosswalk, so two of the three legs are a solved problem and re-deriving them adds nothing. What is missing is the third leg, applied to the specific case where the deployer is an engineering organisation and the "logs under their control" are session transcripts, tool calls, and diffs.
The mismatch that matters is structural. The Act is written around a system with a lifecycle, a provider, and a defined intended purpose. A coding agent is not that shape. It is a harness assembled from a model, a set of tools, a permission boundary, and a repository, reconfigured continuously by the people using it. There is no release version of Tuesday's agent.
This produces three gaps that no crosswalk closes for you:
The unit of record. The Act assumes a system generates logs. A harness generates a session, and a session is not a log — it is a transcript, a set of tool invocations, an authority boundary, and a diff, produced by four components that do not share an identifier by default. Nothing joins the line in main to the run that wrote it unless you make it.
The authority boundary. Frameworks assume permissions are a configuration fact you can attest to. In an agent estate the boundary is a moving target, and the question an auditor asks is not what the agent could do today but what it was allowed to do on the day it did the thing. That is a historical fact about a config file that has changed since, and it is the single most commonly missing record.
The human in the loop. Oversight requirements assume a reviewable decision point. Agents produce hundreds of small decisions and one merge. Recording the merge tells you a human approved a diff. It does not tell you whether the human saw the reasoning, and the frameworks have no vocabulary for that distinction.
Everything below follows from those three gaps.
The traceability matrix
The mapping is only useful at the granularity of an artefact you can point at. "We have governance" is not an answer to anything. The row has to end at a file, a table, or a field.
| Obligation, in substance | RMF function | 42001 area | The artefact that satisfies it |
|---|---|---|---|
| Records exist for what the system did | MEASURE | Operational monitoring | A run record per session: id, start, end, model and version, repository, and the diff it produced |
| Records are attributable to a run | MAP | Tooling resources | A run id written into the commit trailer, so main joins to the session that wrote it |
| The permission state is knowable historically | GOVERN | Human and system resources | An authority snapshot per run: tools available, credentials in scope, approval mode, captured at start |
| Human oversight is evidenced | GOVERN | Human resources | The approval event, with who, when, and what they were shown |
| Records survive the retention window | MANAGE | Data resources | Retention configured to the obligation, not to the log drain's default |
| Records cannot be edited by the subject | GOVERN | System resources | The writer sits outside the agent's boundary |
| Risk of the deployment is assessed | MAP | Context | A scope determination per workload, dated, with the answer and who made it |
| Incidents are reportable | MANAGE | Operational monitoring | An incident path that can name the runs involved |
Read the last column as the deliverable. The first three columns exist to argue that the last one is required; only the last one is a thing you build.
Two rows carry most of the weight and deserve their reasons stated.
The run id in the commit trailer. This is the cheapest row and the one that makes every other row usable. Without a join between the repository and the session, every question becomes a manual archaeology exercise across two systems that share no key. With it, the question "what produced this line" is a lookup. It costs one trailer line per commit and it is the highest-leverage thing on this page.
The authority snapshot. This is the row people skip, because it feels redundant when the config is in version control. It is not redundant. Version control tells you what the file said; it does not tell you what the running process actually had — which tools were registered, which credentials resolved, whether approval mode was on. Capture it at run start, in the run record, as data. A record written today can be reconstructed from yesterday's logs. The boundary that was in force six weeks ago cannot be reconstructed from a config file that has changed nine times since.
The worked example: joining the repository to the session
A concrete version, because the matrix reads as bureaucracy until you see how little of it is actually work.
Start state: about forty repositories, agents with write access in a dozen of them, transcripts retained by the vendor for thirty days, and commits authored by a bot account with a generic message. Asked "which agent run produced this function", the honest answer was several hours of grepping two systems and a guess.
The first change was one line in the commit trailer. Every agent-authored commit gained Agent-Run-Id: followed by the session identifier the harness already had. Nothing else changed that week. The effect was immediate and larger than it sounds: the question stopped being archaeology and became a lookup, and every subsequent row on the matrix became something you could actually check rather than something you could only assert.
The second change was the run record, and this is where the estate's actual state became visible. Writing down model and version, exactly, surfaced that three different model versions were in use across teams and nobody could say which repository ran which. That is not a compliance finding, it is an engineering finding, and it was invisible while the only record was a transcript nobody joined to anything.
The third change was the authority snapshot, and it was the one that met resistance. The argument against it was reasonable: the permission config is in version control, so why duplicate it into every run record? The answer took an incident to land. A run had done something nobody expected, and reconstructing whether it had been allowed to meant checking out the config as of that date, which existed, and then discovering that the tool registry was assembled at process start from three sources, one of which was an environment variable that had since changed. Version control held the file. It did not hold what the process had. The snapshot is four fields and it closed a question that the file could not answer.
What none of that required was a governance programme. It was a commit trailer, a table, and a struct captured at run start. The retention number was a configuration change. The expensive part was not the engineering; it was noticing that the transcripts everyone assumed were evidence had a thirty-day life and no key.
What an evidence pack looks like
At some point a person who is not an engineer asks to see it. That person is an auditor, a customer's security reviewer, or your own counsel, and the artefact they need is not a dashboard.
An evidence pack is a document that answers a specific question with pointers into records that exist. It is assembled per question, not maintained as a standing document, and it has four parts:
The scope statement. What this workload does, whether it was determined to be in scope, when, and by whom. One paragraph. If the determination was that it is out of scope, this is the whole pack and it should say so plainly rather than implying more coverage than exists.
The record schema. What is captured per run, with field names, and the retention applied to it. A reviewer wants to know what could be known, before they ask what was. Publishing the schema is also the fastest way to discover that a field you assumed was captured is not.
Two worked traces. Pick one ordinary change and one that required approval, and show the full path: line in main, run id, run record, authority snapshot, approval event. Two traces demonstrate more than any amount of description, and picking them at random rather than curating them is the difference between evidence and marketing.
The limits. What is not covered, and why. Vendor-side data you do not control. The window before the records started. Workloads out of scope. A pack that claims total coverage invites the reviewer to find the gap themselves, and they will.
The reason to write this down before anyone asks is that assembling it is the test. If the pack takes a week to produce, the records are not usable, whatever the matrix says about them.
Magnet: Agent Traceability Matrix
Run this per workload, not per repository. One pass, and it produces either a scope exit or a build list.
## 0. Scope (do this first, and date it)
- [ ] Does this workload touch an Annex III area, or a product regulated under Annex I?
- [ ] If no: record the answer, the date, and who decided. Stop here. Revisit on material change.
- [ ] If yes or unsure: continue, and get a legal determination in parallel. Not instead.
## 1. The join
- [ ] Does every commit an agent produces carry a run id?
- [ ] From a line in `main`, can you reach the session that wrote it in one lookup?
- [ ] From a run, can you reach the diff it produced?
## 2. The run record
- [ ] id, start, end
- [ ] model and version, exactly, not "Claude" or "the API"
- [ ] repository and branch
- [ ] the prompt or task that initiated it
- [ ] the diff or artefact produced
## 3. The authority snapshot, captured at run start
- [ ] Which tools were registered and callable
- [ ] Which credentials resolved, by name, never by value
- [ ] Approval mode: what required a human, what did not
- [ ] Spend and blast-radius limits in force
## 4. Oversight
- [ ] Is there an approval event with who, when, and what was shown?
- [ ] Can you distinguish "a human merged it" from "a human reviewed it"?
## 5. Durability
- [ ] Retention set to the obligation, stated as a number, not inherited from a log drain default
- [ ] The writer sits outside the agent's boundary
- [ ] Records survive the agent, the repository, and the vendor
## 6. Reportability
- [ ] Given an incident, can you list the runs involved without reading a chat log?You should see: section 1 failing on the first pass, in almost every organisation, including ones with mature logging. That failure is the finding. Teams generally have transcripts somewhere and commits somewhere and no key joining them, which means every subsequent row is technically satisfied and practically unusable. If section 1 passes and sections 3 and 4 fail, you are in the ordinary good case and the authority snapshot is your next build. If every box passes on the first run, either the estate is genuinely mature or the checklist is being answered from memory — answer each box by pointing at the field, not at the intention.
Failure modes
| Smell | Result | Repair |
|---|---|---|
| Treating the deferral as sixteen free months | December 2027 arrives with no history behind it | Start the record now; the deadline moved, the evidence window did not |
| Debug logs offered as evidence | Retention set for cost, success path barely recorded, no join to the repository | Purpose-built run records, retention set to the obligation |
| Authority reconstructed from version control | You learn what the config said, not what the process had | Snapshot the boundary at run start, as data |
| The agent writes its own audit record | If the thing being audited can edit the audit, it is a log and not evidence | Writer outside the agent's boundary |
| Scope assumed once and never revisited | The estate changes, the answer does not | Date the determination, revisit on material change |
| Mapping produced at framework level | Rows that end in "governance" satisfy nothing | Every row ends at a file, a table, or a field |
| "The platform logs everything" | Discovered to be false during the incident | Name the retention number and the field |
| Summaries taken as the source for legal claims | You repeat an error that is now two weeks old | Primary sources for anything load-bearing |
What this does not solve
It does not make an agent's decisions good. It makes them reviewable, which is a different property and the only one you can engineer directly.
It does not establish whether you are in scope, or which role you occupy. Both are legal determinations, and the engineering above is worth doing at the point where you would rather not find out under time pressure.
It does not cover what you do not hold. Article 26(6) reaches logs "under their control", and a substantial part of what an agent does is recorded on a vendor's side under a retention policy you did not set. Name that boundary in the evidence pack rather than letting a reviewer discover it.
And it does not survive an agent with write access to its own record. If the thing being audited can edit the audit, you have a log and not evidence. Keep the writer outside the agent's boundary — the same rule as every other audit trail, and the one most often broken here because the agent is such a convenient place to put it.
When not to build this
- Nothing you run touches Annex III or a regulated product. Record the determination with a date and stop. Building an evidence programme against an obligation that does not attach to you is a real cost paid for nothing.
- You have no agents with write access. If humans author every change and agents only suggest, the traceability question is your existing code review process, and it is already answered.
- The estate is one person and one repository. The join still helps you, but the management-system apparatus does not. Take section 1 and skip the rest.
- You need a scope opinion. This produces engineering artefacts. It does not tell you whether you are in scope, and a matrix cannot substitute for a legal determination.
- You are certifying to ISO 42001 imminently. Then the management system is the deliverable and the artefacts are evidence inside it. Sequence it the other way round: auditor first, matrix second.
path
Make the record a property of the system
Agent OS Setup installs run records, authority snapshots, and retention as part of the harness, so the evidence exists whether or not anyone remembered to turn it on.
The obligation that moved was the date. The obligation to be able to answer the question did not move, because it was never really about a date — it was about whether the system was built to remember what it did.
Your next action: run the Agent Traceability Matrix against one workload, starting at section 0, and stop at the first box you cannot answer by pointing at a field. If that box is in section 1, add a run id to your commit trailer this week. It is one line, it is the join every other row depends on, and it is the only item here that gets harder the longer you leave it, because the commits it would have covered are already landing.
Related: Agent Governance and Traceability for a Regulated Engineering Org covers the record itself in more depth, and Why 10 Passing Tests Still Fail My Coding-Agent Merge Gate covers the artefact a reviewer actually reads.