The Day in AI

Thursday 8 October 2026

←

The day in brief

Google announced a universal agent in the Gemini Enterprise app on 8 October, in private preview for enterprise customers, that takes an objective, plans the work and delegates parts of it to subagents.

Until now, Gemini Enterprise was a place for companies to build and govern their own agents. The new agent takes a goal instead of step-by-step instructions, plans the work, and hands pieces to job-specific subagents that use custom skills and tools. It picks a model automatically, and users can choose another one, including Anthropic's Claude.

The agent has its own Workspace account with an email address, so its actions appear in the audit trail under the agent's name. It connects to Workspace, Microsoft 365, Slack, Jira, Snowflake, BigQuery and any server that speaks the Model Context Protocol (MCP), and users follow its progress in a tasks inbox. Sundar Pichai, Google's chief executive, said at the event that Gemini has over 1 billion monthly active users and that nearly 90% of Fortune 100 companies use Gemini Enterprise. Both figures are Google's own.

Goodfire, an interpretability startup, launched monitors that watch agents from inside the model, available to Baseten customers. Most agent monitors are a second model that rereads everything the agent does. Goodfire uses probes, which are small classifiers that read the activations the model already computes as it generates, and only flagged cases go to a separate AI model for a closer look. Customers choose which risks to watch, such as offensive hacking or reward hacking, and whether a flag triggers logging, human review or refusal.

Goodfire reports that monitoring about 1,500 Kimi K3 sessions cost about $51, against $233 for a cheaper model checking every step and about $10,000 for a top-tier model. It says the probes caught 94% of malicious hacking sessions and sent 8.7% of harmless sessions for a second look, and that four probes added less than 2% to the time before a response starts. These are company figures, and no outside evaluation has been published.

Anthropic launched a Critical Infrastructure Defense Program with 11 founding partners, among them CrowdStrike, Dragos, Palo Alto Networks and Rockwell Automation. It gives defenders of power grids, water utilities, factories and transport networks access to frontier Claude models, threat research and on-site engineers. The program works through security firms and equipment makers that operators already use, and many of these systems cannot be taken offline for patching.

Also in the news

  • StepFun released Step 5 Preview on OpenRouter, a sparse mixture-of-experts model with 600 billion total and 27 billion active parameters and a 1 million token context, at $1 per million input tokens and $2.70 per million output tokens.
  • OpenAI added an Ultrafast service tier for GPT-6.1 Sol in the Responses API that shortens the time between output tokens, open to all API users subject to rate limits.
  • Google released Google AI Edge Foresight, a free experimental Mac app that transcribes meetings and writes notes offline with on-device models, including EmbeddingGemma 2.
  • Microsoft said at its 7 October Windows keynote that DeepSeek V4 Flash, quantized to 1.6 bits, and a Nvidia Nemotron of more than 70 billion parameters will run locally on RTX Spark PCs, and that llama.cpp is coming to Windows ML.
  • Perplexity released pplx-embed-v2-late on 7 October, two open-weight multimodal retrieval models at 0.6 billion and 9 billion parameters under the MIT license, and reports 92.4% on MADQA for the larger one.
  • Architect Financial Technologies launched Liquid Inference, a router that auctions each request among providers serving the named model and charges the lowest offer that meets the buyer's limits.
  • Google Labs launched Playground on 7 October, an experimental US service that builds playable browser games from text prompts.
  • OpenAI reportedly withdrew three of the math papers it published on 6 October, according to a history file in its GitHub repository.
  • OpenAI told investors its annualized revenue was nearing $50 billion at the end of September, according to the Financial Times, well below the $70 billion reported earlier.
  • Arena, which runs the model leaderboard, reportedly raised $200 million at a $3.1 billion valuation and launched an Alignment Index.
  • DeepSeek is reportedly close to raising more than $11 billion.
  • Kuaishou has reportedly picked banks for a Hong Kong listing of Kling AI that targets at least $1 billion.
  • Anthropic updated its usage policy to ban sustained abusive behavior toward Claude, according to The Verge, and added Chinese display languages to Claude's web interface, while mainland China stays blocked.
  • Jasmine Wang, Tomek Korbak and Mikita Balesni, safety researchers OpenAI reportedly fired on 1 October over handling of confidential information, reportedly published an open letter on 8 October denying the claims.
  • Luo Fuli, who leads Xiaomi's MiMo model team, was reportedly promoted to vice president level, according to 36Kr.

Everything from this day

Launches and products

8 Oct 2026GoogleGemini Enterprise universal agent Google announced a universal Gemini agent inside the Gemini Enterprise app that plans work, uses subagents and connects to business systems, in private preview for enterprise customers.8 Oct 2026AnthropicCritical Infrastructure Defense Program Anthropic launched a program that gives critical infrastructure defenders its frontier Claude models, threat research and on-site engineers, with 11 founding partners including CrowdStrike.8 Oct 2026StepFunStep 5 Preview StepFun released Step 5 Preview, a 600B-total, 27B-active sparse MoE model for agentic work, with a 1M token context, on OpenRouter at $1 input and $2.70 output per million tokens.8 Oct 2026GoogleGoogle AI Edge Foresight Google released Google AI Edge Foresight, a free experimental Mac app that transcribes meetings and writes notes fully offline using on-device models.8 Oct 2026OpenAIUltrafast mode for GPT-6.1 Sol OpenAI added Ultrafast mode for GPT-6.1 Sol in the Responses API, a service tier that reduces the time between generated output tokens.8 Oct 2026GoodfireGoodfire inside-out monitors Goodfire launched probe-based monitors that read a model's internal activations to flag rogue agent behavior, available to Baseten customers, at a fraction of the cost of AI monitors.8 Oct 2026ArchitectLiquid Inference Architect Financial Technologies launched Liquid Inference, an LLM router that auctions each request across competing providers and charges the lowest offer that meets the buyer's rules.8 Oct 2026AnthropicClaude Chinese-language interface options Anthropic added simplified and traditional Chinese as on-screen language choices in Claude's web interface in supported markets, while mainland China users remain blocked.