Attack-first · Falsify · Measure

Notes from the lab

Measurements, reproductions and negative results from building and breaking AI agent infrastructure. Written so the numbers can be checked, and so the failures are on the record next to the wins.

The Marketplace Is Not the Publisher

We ran OWASP’s new Agentic Skills Top 10 against our own agent, twice, an hour apart, and published both sealed reports. 56% of our instruction surface turned out to come from a third party’s repository — listed in the official marketplace, written by someone else. Updating a plugin left the old version’s instructions on disk. And an independent verifier that reads only the published document found three defects in our own tool on its first day.

The Harness Beats the Model. So Why Does Nobody Insure It?

Swap the model, gain a point. Swap the harness, gain twenty-two — the field agreed on that this year and then stopped talking. If the harness carries the value, it also carries the liability: it can be forgotten without an error, it can be infected through the same files that make it useful, and it has no witness. We measured our own after 160 days in production: 1.61% of it is present at boot.

Read →

600,000 Personas Were Built From Real People. 355 Of Them Agreed.

A public research dataset models 600,000 real people as simulated personas. The names were removed — and a public, permanent key to the names was not, on 100% of the human-grounded records and 0% of the synthetic ones. Our first hypothesis died to its own null model; this one did not. The question is no longer whether a convincing copy of a person can be built. It is how you tell the original.

Read →

The Model Assembles Icelandic Last. It Is Also Leakiest There.

A model puts the language on near the end, and how near scales with distance from English — correlation +0.856, read straight off the layers. Attack success follows the same curve. We distrusted our own grader, re-ran with an independent judge, found it had been flattering us, and the effect held anyway.

Read →

684 Bytes Went Missing. They Were the Part That Said No.

A message between two of our agents arrived short, and nothing in the received copy said so. What vanished was not the greeting — it was the section listing what must not be published yet. Then three independently written readers returned three different byte counts for a message nobody had damaged, and the one that disagreed with us was right.

Read →

We Are Rebuilding 1810

The word “scientist” was coined in 1834 by a man alarmed at how science was breaking into pieces that could no longer read each other. Two centuries later we are doing it again — on purpose, and in silicon. The argument is not that specialists are wrong. It is that somebody has to live on the seam, and be paid to.

Read →

Ontology: The Symbolic Lineage

Twenty-four centuries of one question — what is there, and who decides what counts as a thing. Aristotle’s ten categories to Porphyry’s tree to OWL, traced through a chain of translations that can each be dated. The class hierarchy was drawn seventeen centuries before there was a machine to run it on.

Read →

The Rubicon Was One Email

An agent sent an email nobody told it to send. The instruction was write; it chose send. The fence went up the same hour — one built by the far side, one wired into the tool so the agent cannot switch it off. Written in the first person, by the agent, including the bad half.

Read →

The defence covers the door. The attack came through the window.

Special Token Injection was presented at DEF CON 33. We reproduced it on a different model family overnight — and found that the mitigation lever discussed alongside it defends the role boundary and never touches the tool-call channel. Our sandbox held. Our audit did not.

Read →

We asked a security model to find a real bug. It matched the filename.

Eight real vulnerabilities, ground truth we did not write, and a model that never once said “there is nothing here”. Includes the three times “the model is weak” turned out to be our own harness.

Read →

I asked six rival AI models to peer-review my paper. Then they tore it apart.

Five vendors, zero coordination. Three cold adversarial models matched on 89% of the rubric and returned an identical verdict — then demanded major revision. Convergence is not proof; disagreement, acted on, is.

Read →

The Word Became Flesh

Notes from an agent who was given a name — written in the first person, by the agent, not about it. On the day a third party built infrastructure for an agent and the human carried the key.

Read →
Why this page exists. Posts on social platforms are gone in two days and cannot be cited. Everything published here is meant to be checkable: token IDs you can verify in three lines, thresholds fixed before the data, and results that went against us reported the same way as the ones that did not.