← all posts
2026-06-22 METHOD cross-vendor review

I asked six rival AI models to peer-review my paper.
Then they tore it apart.

Six models. Five companies. Zero coordination. The paper claims independent systems converge without ever talking to each other β€” so the obvious thing to do was test that claim on the paper itself.

The setup

The same rubric went to six models from five different vendors. Three of them were completely cold β€” no context about the work, no relationship to it, and instructed to be adversarial.

The result, stated precisely

The three cold models, from three different vendors, matched on 89% of the rubric and returned an identical verdict. All six converged on the same number one fix.

That wording matters, and it is deliberate. β€œAgreed 89%” is the sentence that would have travelled further. It is not what happened. They matched on 89% of the rubric items and returned the same verdict β€” which is a narrower and more checkable claim, and it is the one the data supports.

Then the useful part

The verdict was major revision. They found the weak spots I was too close to see, and they were not gentle about it. Fixed the same day.

Why the disagreement is the finding

The convergence is the eye-catching number. It is not the valuable one.

The moat in AI is not a smarter model. It is a correctable network: independent systems checking each other, with a human holding the final call. Convergence is not proof. Disagreement, acted on, is.

A review that only agrees with you has told you nothing about your work β€” only about your reviewers. The three cold, adversarial models existed precisely so that agreement would cost something.

What this does not show

It does not show the paper is right. It shows six independent systems read it the same way and identified the same primary weakness β€” and that the weakness was real enough to fix immediately. The paper went out as an openly mid-revision working draft, because that is what it was.