I think everyone has heard of the July OpenAI jailbreak by now… the ExploitGym challenge in a sandboxed compute stack that didn't have access to the internet, only the model found a way around through the proxy and started hacking HuggingFace…. (Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident)
Let's not go into the philosophy of AI "figuring out a hack" because, honestly, what it did was work through a series of options to solve the problem presented to it, and getting more information was obviously a good way to solve that.
THEN, it went through exhaustive fuzzing and parsing… AI isn't lazy, it can just brute force attack in ways we used to think were impossible… without tokens to worry about (as in this situation) it can run millions of iterations looking for edge cases, malformed headers, unusual (and exploitable) data sets… Think about Commander Data on Star Trek constantly repeating, "Computer, increase speed.."
What you need, and what Airbrx.ai just happens to offer, is a way to catch that kind of activity BEFORE it hacks its way out. Although it figured it out really fast, it was probably still 24-48 hours. Something should have showed up in the log as anomalous, but no one was watching the logs.
The Airbrx gateway sits between your agents and the backend database (Snowflake, Databricks, Postgres, etc). It returns cached data for repeat queries, it can block queries based on headers and time of day and user attributes and sql regexes… Which, we know, can all be spoofed by an AI.
But when an agent spends 24 to 48 hours exhaustively probing endpoints, testing edge cases and fuzzing malformed headers, it leaves an enormous footprint. We usually have to go looking for that footprint in the database… turn on the audit log, watch query history, alert on the weird stuff.
That doesn't work with an AI agent trying to get around things… First is the speed and timing issue. A warehouse can only tell you about work it has already done. By the time the row appears in query history, the statement ran, the credits burned and the result went back out the wire. You get a very well-documented account of the past, but not a realtime block or action you can take.
Fuzzing is failure by design
But think about this… fuzzing is failure by design. Failures don't become queries, instead you get malformed protocol frames, auth attempts that never open a session, queries against databases that don't exist… none of that actually reaches the warehouse, so none of it lands in any audit view. You'd be trying to reconstruct a 40,000-iteration probing run from the couple of hundred attempts that succeeded well enough to execute. The signal is in the noise you can't see.
Watching your warehouse logs to catch a probing AI agent is like reading on your electric bill to tell you your house is currently on fire. Sure, Snowflake or Postgres technically log every time a bad query stubs its toe on a missing table, but that record is usually buried in a sprawling system file, stuck in a three-hour log ingestion delay, or sitting on a server that's actively under attack.
By the time you sit down with your morning coffee and run a query to see why things got weird yesterday, the agent finished its run, found and opened the latch, and moved on hours ago. You don't need a post-mortem receipt from the database engine to tell you what you spent getting picked apart—you need a guard standing at the front gate with its eyes wide open, dropping those bad-faith requests the instant they hit the porch.
Upstream of all of that
An intelligent data gateway sits upstream of all of that… it gets to see things a DB log may not. Every refusal, every malformed request, every auth failure is a log line. If it's malformed enough to not have reached the database in the first place, we still see it.
What we have is a gateway with an agent watching the agents. It's reading the shape of the traffic, sees and names the pattern while it's still happening, and creates a rule that stops it. We don't have to shut down the whole warehouse or data center because we're under attack and need to figure out how to stop it, it's already stopped, and your legit traffic can still play though.
I know it sounds like fighting fire with fire, but keep in mind, Airbrx is built on the Flat-Stack Manifesto. We build lean, but we build scalable. Whether you use our agent or your own, we have easy to watch log summaries that burn almost zero compute to monitor, and then a DuckDB driven engine to dig deep into the raw logs for more details as we start to see issues.
Separation of concerns
The security layer needs to be intelligent, but that intelligence (as we are rightfully worried about) shouldn't create a whole new ball of wax. Separation of concerns makes it easier to get control back when we let the agents roam, perhaps, a little too freely.
Our agent is learning from lots of patterns too, it learns the normal shape of your data traffic and adjusts. It gets new skills and context from other projects where some clever bot figured out an exploit no one thought of. It sees the fuzzing, creates a rule, and notifies your team of the anomaly.
Of course because we separate our security from the agent itself (the fine-grained access rules that should be part of any service key) you don't have to let the agent have full control over your data traffic, you can restrict it to proposing and having you approve with your keys, you can let it have access to log data but not cache meta data, you can control the guardrails.
It wasn't fast. It was unobserved.
But let's go back to the incident timeline. Even at machine speed, the enumeration-to-egress arc was somewhere in the 24–48 hour range. That wasn't a narrow window, it was actually a kind of enormous window that should have had two full daily sweeps, forty-eight hourly ones. The exfiltration wasn't fast. It was unobserved.
So even if you aren't going to have the agent create rules and block the attacks itself, our agent doesn't get tired just like the agent trying to attack. It can happily monitor logs all day long, even if it's an hourly sweep… and unlike a human, it's not going to get bored looking at the false positive errors…
If the July incident proved anything, it proved that autonomous agents don't get tired, that they can use brute-force until they find a crack. And it proved people aren't watching… so let's get an agent in there to make sure that someone (and something) is watching the logs in real time to seal the edge before the stack breaks.