
Autonomous security operations built and run by one person, plus a public record of what we actually tested. Including the test our own mitigation did not pass, and the experiment that ruled against our own central claim.
The bottleneck stopped being detection some time ago. It is now attribution: when an autonomous system acts, can you prove who asked it to?
Annual cost of a three-person SOC team โ out of reach for most companies in this market, which is why so many are reaching for agents instead.
From initial access to full encryption for the top ransomware groups. Overnight is no longer a safe gap to leave open.
Samples in which content an agent merely read steered it into issuing a tool call. Our measurement, on our own harness, published in full.
Fifteen autonomous agents in continuous operation, on infrastructure that costs under $22/month to run. The part that is harder to copy is the audit trail underneath.
File integrity, network analysis, CVE scanning, credential detection, APT monitoring. No shifts, no gaps, no fatigue.
Threat actor profiling, IOC collection, infrastructure mapping and evidence chains โ delivered in hours rather than weeks.
From detection to containment without a human bottleneck. Reports formatted for your CERT, your board or your regulator.
Every action the agents take is recorded with where the instruction came from โ so a stranger's injected instruction cannot be filed as a decision your system made.
These are not demos. They are cases from the production pipeline.
A ransomware-as-a-service group compromised a Czech managed service provider. We mapped 35 infrastructure exposures, identified downstream client risk and reported to the national CERT โ receipt confirmed within 24 hours.
A North African hacktivist group claimed to have breached a Czech university. Our investigation confirmed no breach had occurred, profiled the actor, and surfaced three further campaigns. The finding that mattered was a negative one.
Security tooling is easy to claim and hard to check. So we do the checking in the open โ on other people's systems, on vendor models, and on our own work. These are not case studies written for the website. They are what the measurements returned.
The technique landed in 7 of 12 samples. The mitigation we had shipped for it changed nothing: 10 of 36 with it off, 10 of 36 with it on. We corrected the post rather than quietly editing it, and said which half of our own fix was worthless.
Read the correction โAn open-weight vulnerability localiser, run inside our own read-only cage rather than the vendor's harness. It could not reliably follow its own tool contract. Our verdict on a model we had spent days integrating: not fit to wire in.
Read the envelope โSix models, five vendors, no coordination. Three cold adversarial reviewers agreed on 89% of the rubric and returned the same verdict โ then demanded major revision. We published the demand, not just the agreement.
Read the review โWhy this is the whole pitch. Any supplier can show you what their system caught. The question that decides whether you can rely on it is what it missed, and whether anyone would tell you. We write that part down first, before we know how the story ends โ and we have taken the answer when it went against us.
Not a threat feed we bought. These came out of our own authentication logs on a single evening, after filtering our own access. They are indiscriminate scanners, so they are almost certainly in your logs too โ which makes this checkable rather than impressive. Grep for any of them and see.
Why the last octet is masked. A share of any scanning set is compromised third-party hardware whose owners have no idea it is being used. The /24 is where the signal lives โ the reuse, the spray pattern โ and naming individual hosts on a commercial page adds nothing but exposure for someone already having a bad week. Anyone working a real incident can ask us for the full set.
What we are not showing you. Which of these we blocked, at what layer, or what the rest of the ruleset looks like. That is configuration, and publishing your own defensive posture is a favour to the next person scanning you. The observation is a public good; the response is not.
Why it is here at all. A supplier claiming to run autonomous security should be able to show that something is actually shooting at it. This is that, dated, from one evening.
This section exists because the ones above are easy to write. Anything below was found by us, dated, and left standing.
Every supplier has a list like this. Almost none of them publish it, which is why the published ones are worth something. Nothing below was discovered by a customer.
Thresholds were frozen six weeks before any data existed. The unscaffolded arm beat the scaffolded one, and the scaffolded one introduced false statements about its own process that the baseline did not make. A second, structurally different judge confirmed it the following day. Our own protocol said what that costs, so we paid it: the claim came out of the pitch.
Ten of thirty-six with the defence off. Ten of thirty-six with it on. We had built a lexical defence against a semantic attack โ the model simply regenerated the markers we had stripped. The published post was corrected in place rather than quietly edited, and the failed half is named as failed.
A tooling path issue made a running firewall look absent. Nobody caught it externally โ the same agent filed the correction in its own next report, unprompted, and named the cause. The finding underneath survived: one host had been configured less strictly than its twin, and that gap was closed the same evening.
A model that looked incapable was being starved by our wrapper. The rule that came out of it now runs before every evaluation: the harness is guilty until it proves otherwise, and a subject that cannot pass a competence check does not get to have its silence reported as a result.
The test we hold ourselves to. Write down what would prove you wrong before you look, then publish the answer either way. An organisation that only produces evidence for its wins is not producing evidence.
Customers, investors and researchers welcome. If you are evaluating an agentic system โ ours or anyone's โ the fastest useful question is what it refuses to do and how you would check.
ondrej@tia-framework.com