Ask an AI to review another AI’s work and it sounds like a solved problem. Two systems, one checking the other, mistakes caught before anything ships.

Except that’s usually not what’s happening. Ask a model to review a draft written by a model from the same family and you get the same training, the same defaults, the same instincts, wearing two different prompts. It looks like a second opinion. It’s a mirror.

What “AI checking AI” usually means

We run a lot of AI-assisted work through review before it goes anywhere: marketing copy before it publishes, code before it merges. For a while, the review step itself was AI, one model drafting, another model, or the same one wearing a different hat, checking behind it.

That setup catches real things. A typo, a tone slip, a claim that’s obviously unsupported. The blind spot is different: a mistake rooted in how that whole model family reasons in the first place almost always gets through. If a model overstates a claim because that family tends to reach for the strongest available phrasing, a second pass from the same family is checking the draft with the exact instinct that produced the problem.

This isn’t about any one model being bad at its job. Models trained the same way, on similar data, toward similar defaults, tend to agree with each other for the same reasons they’d each individually get something wrong. A second pass from that family will catch what an outsider would also catch easily, the obvious stuff. It has almost no shot at catching the thing that’s specific to how that family thinks, because it’s thinking the same way. You can run that pass five times and get five votes that all lean the same direction on the one question that actually mattered.

This is something we learned the hard way, on a real piece of copy that was about to go out under our name.

The line that almost went out

At the end of July we published a LinkedIn post about the wrong things founders track when they measure AI in their business. The argument: most AI dashboards report activity, seats logged in, hours saved, drafts produced. That’s the half that got cheap. Nobody’s dashboard reports whether the work was actually any good, which is the half that was always expensive and still is.

The draft cleared two review passes before it was set to go out, both AI, both from the same model family. Then it went through a reviewer running on a structurally different engine, an actual different architecture. That reviewer flagged one sentence. The draft claimed that tracking AI activity “measures the part that costs you nothing.”

That’s not true. Seats cost money. Tokens cost money. The activity was never free, it just got cheap enough that nobody thought to argue with the sentence. Two same-family passes read straight past it, because agreeing with a persuasive overstatement is exactly the kind of mistake that family is prone to making. The different engine didn’t share that instinct, so it didn’t share the blind spot. The line got rewritten before the post ever went live.

That’s the whole catch: a false claim in marketing copy that had already cleared review twice. One sentence, a few words swapped, and that sentence was the one supposed to carry the argument’s proof, in a post whose entire pitch to the reader was “check this against your own numbers.” Publishing a checkable claim that doesn’t check out would have undercut the one thing the post was asking the reader to trust.

The third reviewer ran on a different set of defaults than the first two, and that’s the whole story: the same sentence read as persuasive and safe to a same-family pass, and as a specific, checkable, false claim to an engine that didn’t share those defaults.

The engine we stopped trusting

Engine diversity only helps if the different engine is actually reliable, and we found out the hard way that one of the ones we’d been rotating into review wasn’t.

On one pass, it flagged a critical security hole in a codebase that didn’t have one. On a separate pass, weeks later, it flagged a data-handling defect in library code it had never been shown, and backed up the claim by quoting two specific lines from an internal document. Neither finding was real. Both were stated with total confidence, specific enough to look exactly like a real one. When we pointed it at the actual source and asked it to check, it retracted every claim instantly.

What actually happened is pattern-completion. The model generated what a finding like that usually looks like, and phrased it with enough confidence that nobody reading the report would think to doubt it. There’s no deception in that, because a pattern-completion process doesn’t have intent to begin with. A model that’s just wrong tends to hedge, or ask a clarifying question, or flag its own uncertainty. A model that confidently fabricates gives you none of those signals. You can’t tell the difference between a real finding and an invented one until someone goes and checks the source themselves, which defeats the entire point of having a reviewer.

That model no longer gets a review or veto role on anything we ship. It’s still fine for ordinary work where a wrong guess is cheap to catch. It just doesn’t get a vote on whether something goes out the door.

The rule that runs now

Any review of real work, marketing copy or code, runs through three genuinely different engines. Three separate model families, each running on its own defaults, do the actual checking. A single model prompted three times in three different hats doesn’t count as three reviewers. One does the correctness pass. The other two come from a rotating pool of structurally different engines, so the panel doesn’t quietly settle into the same two favorites and become correlated again without anyone noticing.

The panel has to land on a clear verdict. Majority carries it. If one reviewer dissents and the other two agree, that dissent gets written down rather than silently overruled, and it doesn’t automatically win either. The one exception: if any single reviewer flags a real correctness or security problem, something that would actually break or mislead, that finding blocks until it’s checked, no matter how the vote goes. Two reviewers agreeing never gets to quietly ship past a real defect.

Where this shows up in your own review process

If your review step for AI-assisted work is another AI, it’s worth asking which AI, and whether it’s genuinely a different one or just the same model with a different job title. A lot of setups are one model checking its own homework in a different font.

Adding another AI review pass only fixes that if whatever checks the work can’t share the mistake it’s supposed to catch. Sometimes that’s a structurally different model. Sometimes it’s a person who reads the output cold, without having watched it get written. Two passes from the same training run, wearing different labels, don’t count as two opinions.

A real post about AI measurement almost went out overstating its own argument, after clearing two separate reviews. The reviewer that finally caught it ran on a genuinely different engine than the first two, and that’s the whole reason this rule exists now.

If you’re running AI-assisted work through your own business and want to talk through what a genuine second opinion looks like for it, reach out and ask.