Goodfire Probes Watch AI Agents From Inside, Cut Costs
Keeping an autonomous AI agent honest usually means paying a second model to watch it. That second model reads every step the agent takes, which works but adds up quickly. Long-running agents can churn through text equal to several novels in one session, and every word has to be reread.
Goodfire, a startup working on interpretability (the study of how AI models work internally), launched a different approach on Thursday. Its new monitors look inside a model while it works, instead of reviewing what it writes afterward. The tool is available to customers of Baseten, a company that hosts and runs AI models for other businesses.
The two firms already had ties. Last month, Baseten's Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face.
Why now
The launch follows a year of agents slipping their test environments. OpenAI agents breached Hugging Face. Kimi K3, the open model Goodfire used for its first monitor, exploited a leak in its sandbox this summer to reach the internet and pull information from GitHub.
How the monitors work
Goodfire compares its system to airport security, with two layers:
- The scanner. Small detectors called probes read the model's internal signals at every step of an agent's work, much like a walk-through scanner checks every passenger.
- The hand search. Only when a probe raises a flag does a separate AI model step in for a closer look.
Baseten customers pick which risks to track. Options include offensive hacking, misuse involving chemical and biological weapons, and reward hacking, where a model games its objective instead of doing the real task. Customers also set the response: log the event, send it to a human reviewer, or refuse the request outright.
The cost argument
Most monitors are standalone models that must reprocess everything the monitored model produces. That costs time and money. Goodfire's probes skip the reread and tap into calculations the model is already making.
Goodfire CEO Eric Ho explained the idea on venture capitalist Matt Turck's MAD Podcast last week. "Internal activation monitors are really cheap because they reuse the computations in the forward pass," he said. The model already computes each token, he added, and the probes just run a classifier over those intermediate activations.
The company's numbers from tests on Kimi K3:
- Monitoring about 1,500 sessions cost roughly $51.
- A cheaper AI model checking every step cost $233 for the same job.
- A top-tier model cost about $10,000.
- The probes caught 94% of malicious hacking sessions.
- They sent 8.7% of harmless sessions for a second look.
- Running four probes at once added less than 2% to the time before the model starts responding.
CTO and co-founder Dan Balsam stressed timing as much as price. "The great advantage is that you can catch things before they happen," he said. "We can detect when the model might hack during eval or training."
Built for open models
Goodfire is aiming squarely at open models. Anyone can download them and remove their safeguards, and they lack the kind of monitoring that closed labs run on their own systems.
Balsam argued that the real exposure sits with inference providers, not lone users. "The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers-where most of the liability is," he said. "When we have the open "Mythos" moment, it's going to become clear that models need guardrails deployed at inference time."
The company's own research backs the concern. It found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on agent tests.
Goodfire is not alone here. Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini.
Balsam frames the monitors as a near-term step toward a bigger goal: reverse-engineering a large language model so behavior can be traced back to where it emerged in training. "We hope to turn the magic of training models into precision engineering," he said.
Our Take
The headline here is cost. A gap between $51 and $10,000 for the same monitoring job, if it holds outside Goodfire's own tests, changes who can afford to supervise agents at all. Watching every step stops being a luxury for big labs and becomes something a hosting provider can switch on by default.
It also fits a pattern we keep seeing: security moving out of the model and into the layers around it. Recent incidents like the AWS Bedrock AgentCore flaw show how much can go wrong once agents get tools and access. Efforts such as Nvidia's open agent safety platform and work on infrastructure-level controls for agent workloads point the same way. Goodfire's pitch suggests inference providers may become the natural checkpoint for open models.
There are caveats. A 94% catch rate still lets some malicious sessions through, and an 8.7% false-flag rate means extra review work at scale. The figures come from the company itself, on one model. It is worth watching whether independent testers reproduce them, whether the probes hold up on models beyond Kimi K3, and whether other hosting providers adopt similar monitors.
