What is AI drift?
AI drift is what happens when an AI output reflects the source but not faithfully. The text holds together, the structure looks reviewed, the meaning is just slightly off. By the time anyone notices, the brief has gone out, the recommendation has been made, the decision has been signed off on what the AI produced rather than what the source said.
Drift is the dangerous case
Most conversations about AI accuracy focus on hallucination. Hallucinations are outputs that contain claims with no basis in any source. They are the easier problem. Often they are obviously wrong. When they are not obvious, they tend to be checkable. Someone notices.
Drift is harder because it does not look wrong. The output stays close to the source. It uses the right vocabulary. It covers the right topics. It reads like a faithful reflection of the evidence. But somewhere between the source and the output, meaning shifted. A condition was dropped. A risk was softened. A qualifier disappeared. A conclusion was drawn that the source never actually reached.
The output was not wrong enough to trigger review. That is exactly when the damage happens.
AI optimizes for plausibility, not faithfulness. Models compress, weight, and redirect source material as part of producing fluent text. Drift is not a bug. It is a property of how generation works.
What drift looks like
Here is drift in practice. Source material from a vendor contract, and the AI output produced from it:
The vendor is obligated to deliver a remediation plan within 30 days of any material breach, subject to written approval from both parties before implementation.
The vendor is expected to deliver a remediation plan within 30 days of any material breach.
Drifted. "Obligated" softened to "expected." The written approval requirement was dropped entirely. The output reads correctly. The meaning changed.
Nothing in this output is invented. The vendor, the 30-day window, the remediation plan are all present and accurate. But "obligated" and "expected" do not mean the same thing in a contract. And the approval requirement was decision-relevant. It disappeared with no signal that it was ever there.
The forms drift takes
Drift is not a single failure mode. It shows up in distinct, separable forms. Plumb categorizes every claim in an AI output against four verdicts.
The output held up
The claim is directly grounded in the source. No drift, no compression, no reframing. What the AI produced is what the source said.
The meaning shifted
The claim reflects source material but with changed meaning. Softened, strengthened, redirected. The vocabulary stayed close. The signal moved.
The signal went missing
Source material with material weight is absent from the output. A risk, a condition, a constraint, a flag. Present in the evidence. Gone from the deliverable.
The claim has no source
The output contains content that cannot be traced to the source at all. This is the classical hallucination case. It is the rarest of the four and the easiest to catch.
Why drift matters now
AI is moving from experiment to production in firms that produce decision-bearing work. Consulting briefs. Legal memos. Risk assessments. Vendor evaluations. Client-facing recommendations. Investment notes. The outputs of AI workflows are being acted on by partners, clients, and stakeholders, often without the source being checked.
When those outputs drift, the decision rests on what the AI produced, not on what the source said. The reputational, contractual, and fiduciary weight sits on the deliverable. The integrity sits in the source. Drift is the gap between the two.
AI is being deployed into formal decision systems without the infrastructure that keeps its outputs trustworthy. That infrastructure is what Plumb is.
Drift and plausible distortion
The dangerous form of drift, where the output passes casual review while carrying changed meaning, is what we call plausible distortion. Drift is the broader phenomenon. Plausible distortion is the specific class of drift that reaches a decision without being noticed.
Not every form of drift is plausible distortion. An output that drifts and reads obviously off is just a bad output. Someone catches it. Plausible distortion is drift that looks right. That is the case Plumb is built for.
How Plumb addresses drift
Plumb is the source integrity layer for AI. It runs between AI workflows and the people who act on their outputs. Two things go in: the material the AI worked from, and the AI output exactly as produced. What comes back is a structural account of what the output actually represents.
Plumb does not pattern-match strings. It reconstructs the semantic structure of what the source says. What is asserted. What is conditioned. What is negated. What scope each claim carries. Then it reads the output against it, claim by claim. Every finding is owned by a specific source unit. Nothing without a source owner reaches a human.
Drift gets caught at the layer between AI and decision. Not as a quality score. Not as a badge. As infrastructure.
Is AI drift the same as model drift?
No. Model drift refers to a model's performance degrading over time as the world or the data changes. AI drift, as Plumb uses the term, refers to drift between an AI output and the source material it was produced from in a single interaction. The two are different problems. Plumb addresses the second.
Can better prompts prevent AI drift?
Better prompts reduce drift in some cases but do not eliminate it. Drift is a function of how language models generate output. They compress, weight, and redirect source material as part of producing fluent text. A more carefully prompted output can still soften an obligation, drop a qualifier, or omit a key risk. The only way to know what survived is to read the output back against the source.
Where does AI drift show up most?
Anywhere AI is turning source material into output someone acts on. Document summaries, vendor evaluations, contract reviews, briefs from project notes, due diligence memos, client-facing recommendations. The common pattern is a source set, an AI output produced from it, and a human acting on the output without re-reading the source.
Why is drift harder to catch than hallucination?
Hallucinations contain claims with no source. They tend to surface contradictions or implausible details that trigger review. Drift contains claims that reflect the source but with changed meaning. The vocabulary is right. The structure is right. The signal is off. Casual review does not catch drift because there is nothing visible to flag.
How does Plumb fit into an existing AI pipeline?
Plumb runs between the AI and the human. The source context and the generated output are passed to Plumb before the output reaches a person. Plumb returns a structural account of the output: what was supported, what drifted, what was omitted, what was invented. One integration. One contract. Every workflow.
See drift in your own work.
15 minutes. We run Plumb against a real output, yours or one of ours, and show you exactly what it finds.
Book a Demo →