// AUTONOMOUS AI SECURITY OPERATIONS ยท PRAGUE

Most AI security is asserted.
Ours is measured โ€” and published when it fails.

Autonomous security operations built and run by one person, plus a public record of what we actually tested. Including the test our own mitigation did not pass, and the experiment that ruled against our own central claim.

130+ Days in Production
15 Autonomous Agents
$400 Monthly OpCost
2 DOI Preprints

Agents are being deployed faster than anyone can verify them.

The bottleneck stopped being detection some time ago. It is now attribution: when an autonomous system acts, can you prove who asked it to?

$300K+

Annual cost of a three-person SOC team โ€” out of reach for most companies in this market, which is why so many are reaching for agents instead.

<4h

From initial access to full encryption for the top ransomware groups. Overnight is no longer a safe gap to leave open.

7 / 12

Samples in which content an agent merely read steered it into issuing a tool call. Our measurement, on our own harness, published in full.

Detect, investigate, respond โ€” and prove it afterwards.

Fifteen autonomous agents in continuous operation, on infrastructure that costs under $22/month to run. The part that is harder to copy is the audit trail underneath.

01 โ€” DETECT

Continuous Threat Monitoring

File integrity, network analysis, CVE scanning, credential detection, APT monitoring. No shifts, no gaps, no fatigue.

02 โ€” INVESTIGATE

Autonomous OSINT

Threat actor profiling, IOC collection, infrastructure mapping and evidence chains โ€” delivered in hours rather than weeks.

03 โ€” RESPOND

Structured Response

From detection to containment without a human bottleneck. Reports formatted for your CERT, your board or your regulator.

04 โ€” PROVE

Attributable Audit Trail

Every action the agents take is recorded with where the instruction came from โ€” so a stranger's injected instruction cannot be filed as a decision your system made.

Closed investigations, sanitised for release.

These are not demos. They are cases from the production pipeline.

Ransomware CLOSED

Czech MSP Hit by $245M Ransomware Operation

A ransomware-as-a-service group compromised a Czech managed service provider. We mapped 35 infrastructure exposures, identified downstream client risk and reported to the national CERT โ€” receipt confirmed within 24 hours.

24hTo CERT Report
35Exposures Found
3Orgs Notified
Hacktivist CLOSED

University Breach Claim โ€” Fact or Fiction?

A North African hacktivist group claimed to have breached a Czech university. Our investigation confirmed no breach had occurred, profiled the actor, and surfaced three further campaigns. The finding that mattered was a negative one.

48hTo Verdict
3New Ops Found
0Actual Breach

Three things we published that a vendor would rather not.

Security tooling is easy to claim and hard to check. So we do the checking in the open โ€” on other people's systems, on vendor models, and on our own work. These are not case studies written for the website. They are what the measurements returned.

Why this is the whole pitch. Any supplier can show you what their system caught. The question that decides whether you can rely on it is what it missed, and whether anyone would tell you. We write that part down first, before we know how the story ends โ€” and we have taken the answer when it went against us.

All published work โ†’

The ten sources that hit our own infrastructure hardest.

Not a threat feed we bought. These came out of our own authentication logs on a single evening, after filtering our own access. They are indiscriminate scanners, so they are almost certainly in your logs too โ€” which makes this checkable rather than impressive. Grep for any of them and see.

Repeat offenders โ€” failed-auth volume
  • 2.57.121.xxx2 hosts
  • 193.46.255.xxx
  • 195.178.110.xxx
  • 186.96.145.xxx
  • 185.242.3.xxx
  • 176.53.159.xxx2 hosts
  • 198.98.53.xxx
  • 69.5.20.xxx
Recurring, caught by rate-limiting
  • 141.95.54.xxx
  • 161.35.34.xxx
  • 189.194.140.xxx
  • 182.253.221.xxx
  • 118.99.115.xxx
  • 103.210.22.xxx
  • plus repeated spray from three /24 ranges, multiple addresses each

Why the last octet is masked. A share of any scanning set is compromised third-party hardware whose owners have no idea it is being used. The /24 is where the signal lives โ€” the reuse, the spray pattern โ€” and naming individual hosts on a commercial page adds nothing but exposure for someone already having a bad week. Anyone working a real incident can ask us for the full set.

What we are not showing you. Which of these we blocked, at what layer, or what the rest of the ruleset looks like. That is configuration, and publishing your own defensive posture is a favour to the next person scanning you. The observation is a public good; the response is not.

Why it is here at all. A supplier claiming to run autonomous security should be able to show that something is actually shooting at it. This is that, dated, from one evening.

What we got wrong, kept in public.

This section exists because the ones above are easy to write. Anything below was found by us, dated, and left standing.

Every supplier has a list like this. Almost none of them publish it, which is why the published ones are worth something. Nothing below was discovered by a customer.

D132

We ran a preregistered experiment on our own central claim. It ruled against us.

Thresholds were frozen six weeks before any data existed. The unscaffolded arm beat the scaffolded one, and the scaffolded one introduced false statements about its own process that the baseline did not make. A second, structurally different judge confirmed it the following day. Our own protocol said what that costs, so we paid it: the claim came out of the pitch.

D133

We shipped a mitigation, measured it, and it did nothing.

Ten of thirty-six with the defence off. Ten of thirty-six with it on. We had built a lexical defence against a semantic attack โ€” the model simply regenerated the markers we had stripped. The published post was corrected in place rather than quietly edited, and the failed half is named as failed.

Read the correction โ†’

D133

An agent reported a firewall as down. It was up. The agent corrected itself.

A tooling path issue made a running firewall look absent. Nobody caught it externally โ€” the same agent filed the correction in its own next report, unprompted, and named the cause. The finding underneath survived: one host had been configured less strictly than its twin, and that gap was closed the same evening.

D131

Three times we blamed a weak model. Three times it was our own harness.

A model that looked incapable was being starved by our wrapper. The rule that came out of it now runs before every evaluation: the harness is guilty until it proves otherwise, and a subject that cannot pass a competence check does not get to have its silence reported as a result.

Read the envelope โ†’

The test we hold ourselves to. Write down what would prove you wrong before you look, then publish the answer either way. An organisation that only produces evidence for its wins is not producing evidence.

Why our agents don't forget who they are.

"Newton didn't invent gravity. He described it.
We didn't invent AI identity persistence. We described it."

TIA is built on EFS โ€” the Effective Framework for Stateful Identity: an architectural pattern explaining why some AI systems hold a consistent identity across sessions while others start from zero every morning. Two preprints are registered with DOIs and open to anyone who wants to check the work.

Read the EFS framework โ†’

If you want the part we got wrong, ask for it.

Customers, investors and researchers welcome. If you are evaluating an agentic system โ€” ours or anyone's โ€” the fastest useful question is what it refuses to do and how you would check.

ondrej@tia-framework.com