Essay
The AI companies want this too
The story says creators and AI labs are on opposite sides of the provenance fight. The story is wrong. Every serious lab has the same problem the artists do, pointed the other direction. They need to prove what trained their models and what did not.
A senior counsel at a major model vendor sent a private message in March 2026 to the head of a provenance standards body. The message asked whether the body's working group would be open to adding a model-side counterpart specification: a manifest format that would allow a model vendor to publish, in a verifiable form, the cryptographic hashes of every dataset their model had been trained on. The standards body had been working on the creator side of the same problem. The counsel was asking whether they would accept the symmetric work on the other side.
The standards body said yes. The work is in draft. The lab that initiated the conversation has not put out a press release. The press release is not the interesting artifact. The legal-department behaviour is.
Why would a model vendor want this?
The vendor's legal exposure on copyright runs through one specific evidentiary primitive: what was in the training corpus, and when. The 2024 and 2025 copyright cases against major vendors (Getty v. Stability AI, the New York Times v. OpenAI, several class actions filed in California's Northern District) have all turned on what the vendor's discovery process can produce in response to a subpoena. The cost of producing that discovery has been substantial in every case, with one vendor reportedly spending in excess of $14 million on a single round of forensic dataset reconstruction.
A model vendor that has signed, dated manifests of its own training datasets at the moment of training does not face that reconstruction cost. The manifest is the discovery response. The vendor produces it, the opposing counsel verifies it against the registry, and the question of what was in the corpus is resolved on documentary evidence rather than on expensive litigation discovery. The infrastructure that lets a creator prove they were in the corpus is the infrastructure that lets a vendor prove a different creator was not.
This is not speculation. It is the same logic that drives every documentary record-keeping discipline in any industry that has been sued enough to learn the lesson. Banks did not invent transaction logs because they enjoyed the engineering work. They invented them because the alternative was reconstructing transactions from witnesses, which was both more expensive and less defensible.
Why is the public narrative still adversarial?
Two things keep the adversarial framing alive past its actual relevance. The first is the active litigation. A vendor that is currently a defendant in a specific case cannot publicly endorse the plaintiff's preferred infrastructure without affecting the litigation posture. Counsel will not allow it. The infrastructure work proceeds quietly, often through trade associations or standards bodies, until the active case is resolved.
The second is the political utility of the framing. Creator advocacy organizations have built their public narrative around opposition to model vendors. Vendor public-affairs teams have, until recently, treated the creator coalition as an adversarial constituency rather than a regulatory ally. Both groups have an incentive to maintain the adversarial frame in public, even as their respective legal departments are quietly converging on the same architectural answer.
The frame breaks when one lab does the math on its own exposure and decides the public alignment is worth the cost. The calculus is roughly: continued adversarial posture costs us roughly $40 million per year in discovery and settlements, with rising exposure as the active cases multiply. Public alignment with creator-side provenance plus the symmetric model-side specification costs us roughly $8 million per year in standards-body participation, engineering implementation, and reputational repositioning. The math is not subtle. The first lab to run it publicly is the lab that captures the regulatory dividend.
What does symmetric infrastructure look like in practice?
The creator-side artifact is the Pulse Signature or equivalent: a signed attestation that a specific human signed a specific work at a specific time, anchored in an append-only registry.
The vendor-side artifact is the training-set manifest: a signed attestation that a specific model checkpoint was trained on a specific dataset whose contents are enumerable by cryptographic hash, with the training date bound to the manifest. The vendor publishes the manifest at the moment of training. The registry accepts the manifest. When a creator subsequently asks whether their work was in the training set, the question is resolvable by hash lookup against the manifest. The answer is yes, no, or unknown, with the answer dated to the training event.
The two registries are symmetric. They share the same primitive (cryptographic hash, append-only record, registry-witnessed timestamp). They serve mirror-image queries. A creator's lookup answers "did anyone train on my work without my consent." A vendor's lookup answers "did I train on this work, and under what licence." The same infrastructure resolves both.
The cost of building both sides is non-trivial. The cost of building neither and continuing to litigate is significantly higher, for both parties, on every metric that matters to either side.
What is the timeline for the announcement?
The internal calendars at the major vendors suggest a window in the second half of 2026. The pressure points are specific: the resolution of one or two of the high-profile active cases, the operative dates on the EU AI Act provisions, the procurement cycles on enterprise contracts that begin requiring provenance attestations as a condition of purchase. Once any one of those events lands, the lab that is first to make a public alignment with creator-side provenance captures the resulting positioning.
The labs that move second and third do not capture the same positioning. They become signatories to a standard that someone else announced. The competitive value of being first is the kind of value that gets a specific quarter assigned to the announcement in an internal planning document.
The creator should not wait for the announcement. The creator should sign their work into the registry now, on the architecture that exists, with the understanding that the same architecture is the one the labs will adopt on their own side when the announcement is made. The creator who is in the registry on the day the vendor manifests start landing has a working evidentiary record. The creator who is not is back to discovery, which is the expensive path everyone is trying to leave.
The story is artists versus AI labs. The story is wrong. The architecture is the same on both sides, the legal departments know it, and the announcement is coming. The infrastructure that protects a working illustrator is the same infrastructure that defends the model vendor facing depositions. Once that becomes visible, the category reorganizes around the alignment, and the publication schedule for the announcement is already on someone's calendar.
Frequently asked questions
Why would an AI lab care about creator-side provenance infrastructure?
- Because the lab's own legal exposure runs through the same evidentiary primitive. A lab facing a copyright claim needs to demonstrate, with documentary evidence, that a specific work was not in their training corpus, or was used under a specific licence, or was used under a specific date-bounded permission. The infrastructure that lets a creator prove they signed a work is the infrastructure that lets a lab prove they did or did not ingest it. The lab's interest is in a comprehensive registry, not in a fragmented one.
Aren't AI labs in litigation with creators right now?
- They are, in specific named cases. The litigation is over what happened. The infrastructure conversation is over what would let the next dispute be resolved on documentary evidence rather than discovery. Those are different conversations. A lab can be the defendant in a 2025 case and simultaneously be quietly building the infrastructure that would have prevented the dispute, because the alternative (continued exposure on every future case) is structurally worse.
Has any lab publicly endorsed creator-side provenance?
- Several have endorsed individual standards (Adobe's leadership on C2PA, OpenAI's public statements on watermarking, Anthropic's responsible-scaling disclosures). No major lab had, as of the third quarter of 2026, made a comprehensive public statement aligning with creator-side biometric attestation specifically. The legal departments have been moving faster than the public communications departments. That gap is the predictable lag.