Goodfire inside-out monitors
Goodfire launched probe-based monitors that read a model's internal activations to flag rogue agent behavior, available to Baseten customers, at a fraction of the cost of AI monitors.
Most agent monitors are a second model that rereads everything the agent does. Goodfire's small probes read activations the model already computes in its forward pass, and only flagged cases go to a separate AI model for a closer look. Customers choose the risks to watch, such as offensive hacking, chemical and biological misuse and reward hacking, and pick the response, which can be logging, human review or refusal.
- Date
- Thursday 8 October 2026
- Lab
- Goodfire
- Kind
- product
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Cost to monitor about 1,500 Kimi K3 sessions | about $51 compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier model | company |
| Malicious hacking sessions caught | 94% 8.7% of harmless sessions were sent for a second look | company |
| Added latency to first response with four probes | less than 2% time to start responding | company |
Offered through Baseten's platform; the first monitor was built around the open model Kimi K3. Figures come from Goodfire's own tests as reported by TechCrunch.
Sources
This record was checked and corrected against its sources on 8 October 2026. How we check