Goodfire inside-out monitors

Goodfire launched probe-based monitors that read a model's internal activations to flag rogue agent behavior, available to Baseten customers, at a fraction of the cost of AI monitors.

Most agent monitors are a second model that rereads everything the agent does. Goodfire's small probes read activations the model already computes in its forward pass, and only flagged cases go to a separate AI model for a closer look. Customers choose the risks to watch, such as offensive hacking, chemical and biological misuse and reward hacking, and pick the response, which can be logging, human review or refusal.

Date
Thursday 8 October 2026
Lab
Goodfire
Kind
product
Access
closed API

Figures

MeasureValueMeasured by
Cost to monitor about 1,500 Kimi K3 sessionsabout $51
compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier model
company
Malicious hacking sessions caught94%
8.7% of harmless sessions were sent for a second look
company
Added latency to first response with four probesless than 2%
time to start responding
company

Offered through Baseten's platform; the first monitor was built around the open model Kimi K3. Figures come from Goodfire's own tests as reported by TechCrunch.

Sources

  1. techcrunch.com/2026/10/08/goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-

This record was checked and corrected against its sources on 8 October 2026. How we check

Read the daily brief for 8 October 2026