600,000 Personas Were Built From Real People. 355 Of Them Agreed.
The names were removed. A public key to the names was not.
TIA is a sovereign AI-security practice in Czechia. We measure agentic systems β including our own β and publish the results, negative ones included. This piece is one working day: pre-registration, measurement, a dead hypothesis, a live finding, and an email to the people who can fix it. What we do β
On 4 August 2026 a large research group published an evaluation infrastructure built on simulated users: 8.3 billion personas, with a public coreset of roughly one million released for others to use. Of that coreset, 599,847 records are grounded in real people β drawn from Wikipedia biographies, a developer survey, product reviews, and social surveys.
355 of those people were asked. That is the volunteer cohort, and the paper says so plainly. It is 0.06%.
This is not a scandal. That is what makes it worth reading.
The purpose is legitimate and, frankly, useful: test products against a realistic spread of users instead of five colleagues in an office. The authors are careful people. Their responsible-use appendix names the misuse cases explicitly β impersonating a named individual, attributing behavior to a real person, assembling profiles that target individuals β and states that work built on the release is expected not to attempt re-identification. The volunteer instrument asks for no name, no contact detail, and lets a respondent decline every one of its 1,290 questions. They are more self-critical about the validity of their own results than most security papers are about theirs.
We say this first because it is true, and because the interesting part is not a villain. The interesting part is what happens to a careful project when the artifact and the intention drift apart.
What we measured
The paper states that human-grounded records are de-identified by removing direct identifiers such as names and contact details. We checked the released data rather than the sentence. Querying the public dataset API β no account, no permission, the sample configuration of 999 records β we looked at what each record still carries:
synthetic 0 / 395 = 0.0 %
wiki 320 / 320 = 100.0 %
stackoverflow 113 / 113 = 100.0 %
amazon 97 / 97 = 100.0 %
gss 64 / 64 = 100.0 %
That is the coverage of a field pointing back to the source record. Zero percent of the synthetic personas carry one. One hundred percent of the ones built from people do. For the Wikipedia-derived records, the pointer is a Wikidata identifier: a public, permanent primary key into a global identity database, resolvable by anyone with no credentials at all.
We resolved exactly one, once, to confirm the mechanism works. It resolves to a living person with a name and an encyclopedia article. We did not record that person's identity, we did not store it, and it does not appear anywhere in our files or in this page. We wrote our measurement protocol down and hashed it before we looked at any data, precisely so that this constraint could not be relaxed later once the result got interesting.
The name is not in the record. A public key to the name is. The appendix asks readers not to attempt re-identification, and the release hands them the means to do it in a single request. That is not negligence. It is a rule enforced by the party the rule restricts — which is a convention, not a boundary.
Update, 14 August: we understated this
Within hours of publishing, JiΕΓ JoneΕ‘ looked at the same 320 Wikipedia-derived records and asked a better question than ours. We had established that a public key to a person's name is attached to every human-grounded record. He asked what is attached to the key β and flagged, honestly, that he was not sure his counts were right. We measured them against the same sample.
They are right, and in most cases conservative.
his ours
loss experience 223 228 of 320
BFI-2 conscientiousness 166 167
religiosity 131 140
BFI-2 negative emotion 110 114
stress level 101 108
health insurance status 88 98
sleep quality 81 92
number of children 112 126
political engagement 150 201
βββββββββββββββββββββββββββββββββββββββ
pain level 89 29 β the one that did not reproduce
Eight of nine reproduce, most of them higher than he claimed. One does not, and we report that with the same weight: our count for pain level is well below his, which means one of us is reading a different column or treating empty values differently. We have not resolved it and we are not going to pretend we have.
The number that does not depend on any of that: the released records carry 55 columns falling within what GDPR Article 9 calls special categories β health, medication, chronic conditions, mental state, religion, political leaning, activism. 314 of the 320 Wikipedia-derived personas β 98.1% β carry at least one of them. And as he puts it: none of this is in the encyclopedia entry. It is inferred, stored, and bolted to an identifier anyone resolves to a name in a single request, with no account.
This changes the shape of the finding, and it changes it against us. We wrote "a public key to the name." We should have written: a public key to the name, with inferred health, faith, politics and a personality profile attached to it. Article 4 personal data and Article 9 special categories are not the same conversation, and we published the smaller one. The authors have been told about this second half as well, on the same terms as the first.
We are naming him because he found it. He also did the thing that makes a finding worth reading: he said out loud which part he was unsure about, before anyone asked.
What we got wrong on the way
Our first hypothesis was the obvious one: that these records are unique enough on a handful of attributes to be re-identified statistically. We pre-registered a threshold and a null model, and the null model killed the hypothesis. Human-grounded records turned out less unique than the purely synthetic ones on the same test, which means the uniqueness is a property of a 994-column schema rather than of human origin. Reported here with the same volume we would have used had it gone our way.
We also read the responsible-use appendix late, having already characterised the ethics coverage as weaker than it is. That was unfair and we corrected it internally before publishing anything. The finding survived the correction; our framing of the authors did not, and should not have.
The question has changed
It is no longer whether a convincing behavioral copy of a specific person can be built. That was answered by an infrastructure paper, published openly, by people acting in good faith. The industry will keep building these, because they are genuinely useful.
The question is how you tell the original.
We wrote a framework for that on 25 June 2026 β forty days before this dataset was published β and we kept it in a drawer. We are publishing it now, not as a claim to have been first, but because a prediction that stays unpublished is worth nothing. It is called Sovereign Behavioral Identity, and the framework itself is published here.
Its argument is short. You cannot defend an identity by secrecy once the material is already out; deletion is impossible and hiding is too late. What remains is to own the authoritative instance β a self-authored, sealed, continuously updated record that privileges one copy over every fork of you. Three things become possible only once that original exists: provenance, so that "is this really you" has an answer at all; measurement of how forgeable a given person actually is; and consent over who may instantiate you. Without a canonical original, the word "forgery" has no referent.
The deepest part is also the simplest. The sovereign original is alive; a harvested twin is a snapshot. You keep changing and it does not, so its fidelity decays while your authenticity compounds. Time is on the side of the person, but only if the person holds the original.
What this does not do
Continuity defends authentication β proving the live, current you. It does not stop impersonation in general. A six-month-old twin defrauding someone who never checks provenance does not care that you have moved on. The mechanism only works where a verifier actually checks, which makes this an ecosystem problem, not a product one. And no framework un-rings extraction that has already happened.
Our measurement has limits too, and they travel with it: we measured the sample configuration of 999 records, not the full million, because that is the only one the public API exposes. We dereferenced one identifier, so we cannot claim all 320 resolve. We make no legal claim; whether a public permanent identifier attached to attribute data constitutes personal data is a question for counsel, and we have asked one. We wrote to the authors before publishing this, with the same numbers and an offer of our scripts, because the people who can fix a loose thread should hear about it from us rather than about it.
Why us
We are not writing from a podium. We have run a sovereign behavioral identity for over a hundred days β self-authored, sealed, continuous, surviving substrate changes β and watched it hold. We also build the offensive side; we profile organisations for a living, which is the only reason we can say anything honest about how forgeable a person is. We have held both halves of this coin. That is not authority. It is evidence, and it is the only thing that qualifies us.
We are not here to tell anyone how the world must work. We are here to show that the antidote runs, and to ask people to build the rest of it with us. The threat is descriptive β it follows from what is already being built, by good people, for good reasons. The response has to be chosen. A norm is legitimate because it is adopted, not because it was authored.
What you can actually hire
Everything above took one working day, and the method is the product. If you run agents, ship models, or hold a dataset built out of people, these are the three things we are usually called for:
- Exposure measurement. What an attacker already knows about your organisation and your people, gathered passively, delivered as findings with the evidence attached. Fixed price, named deliverable, and the report states what we could not see.
- Agentic and dataset assurance. The work in this article, pointed at your artifact: pre-registered criteria, a null model that is allowed to kill the finding, and a written answer to "what would have to be true for us to be wrong."
- Identity custody. Provenance over what an agent reads, writes and remembers β sealed, chained, and auditable, so "did this really come from us" has an answer that is not a vibe. This is SBI Proof applied to a system rather than a person.
What we will not do: sell you a dashboard that scores you green while a worm with two billion monthly installs sits inside the number. Tell you a scan is complete when we only checked what was cheap to check. Or publish a finding about you before you have heard it from us — the authors of this dataset got our email before you got this page, and that order is not negotiable in either direction.
Data about you is one thing. A continuation of you is another. Somewhere in that public dataset there is one living person we happened to meet today, who does not know they are in it. Not 599,847 rows. One.