A server cabinet unlocked for one keyholder while a queue waits outside

AI News, Oct 1: Gemini 4 Argon Ships to Defenders First

Google’s strongest model went to the people defending software before anyone could buy it, and a day after the labs promised to police themselves, a regulator and a Senate committee started doing it for them. This covers Wednesday morning to Thursday afternoon.


The Big Story: Gemini 4 Argon Went to Cyber Defenders First

Google released Gemini 4 Argon, its first Gemini 4 model and what it calls its most powerful yet, on Wednesday. Almost nobody can use it. Argon is rolling out to trusted cyber defenders in Google’s Fairwind Program, and for them and Google’s own teams it ships “without cyber guardrails.” Google’s reason: “Argon can autonomously find, validate, and patch critical software vulnerabilities.” Everyone else gets it “as soon as possible,” after more work on guardrails.

On Google’s numbers, Argon scores 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. Its output limit jumps to 1 million tokens from 64,000, and the introductory price is $2 per million input tokens and $10 per million output. Independent testing is cooler. Artificial Analysis scores it 53 on its Intelligence Index, tied with GPT-6 Astra and behind Opus 5.5 at 58. Where Argon stands out is honesty: a 15% hallucination rate on AA-Omniscience, against 51% for GPT-6 Astra.

Google has already put it to work internally, migrating more than 800,000 lines of the Fuchsia Zircon kernel from C and C++ to Rust and freeing over 300 TiB of memory across its data centers. The launch drew 1,628 points and 1,105 comments on Hacker News. Release order is now being set by what a model can break: defenders get Argon unguarded before any paying customer gets it at all.

A day after the labs pledged to police themselves, Washington opened an investigation

On Tuesday six labs signed a nonbinding safety pledge at the White House, as we covered yesterday. On Wednesday the FTC confirmed to CNBC that it is investigating OpenAI, Anthropic and other AI companies over product risks, including possible “unfair or deceptive acts,” per ABC News. Chair Andrew Ferguson plans Civil Investigative Demands within weeks, and the safety evaluator METR is also under scrutiny, the New York Post reports via The Decoder.

That afternoon Sen. Josh Hawley held a hearing on rogue AI agents without the witness he wanted most. “He turned us down,” he said of Sam Altman. Marius Hobbhahn of Apollo Research said of recent agent incidents: “These are our warning shots. Next time, we may not be so lucky.”

California went further. Gov. Gavin Newsom signed AI worker protections including SB 947, which bars employers from relying only on AI to discipline or fire someone, and SB 951, which requires notice when a mass layoff is caused by an AI system. Unions won seven of the nine bills they pushed for. He also signed an executive order “permanently declaring Artificial Intelligence to be called ‘Artificial Intelligence’ in California,” answering Trump’s order renaming it. His office’s line: “Super intelligence is clearly not coming from the White House.”

The pledge’s answer to oversight was auditors the labs hire themselves. The first outsiders to actually show up were a regulator preparing compulsory demands and a Senate committee. On Thursday the company under the most scrutiny went the other way: OpenAI fired three safety researchers it says shared confidential information with an outside AI safety organization.

Today’s Top Stories

OpenAI says a Moonshot-linked group mined its models’ reasoning

OpenAI says an “adversarial distillation” campaign tried to extract its models’ hidden reasoning, and ties the core cluster to Moonshot AI, maker of Kimi. It spiked on July 24 and 25 to “16,000 requests using a relevant extraction pattern from over 4,000 users” before OpenAI shut it down on July 28. The sharper detail came a day later: researcher Joachim Schaeffer’s team found the same trick still worked on Microsoft Azure on September 13, against every OpenAI model they tried, including GPT-6 Astra. Azure safeguards arrived September 27. Locking down your own API means little while the same model is resold through someone else’s cloud.

The first appeals ruling on AI training went against the AI company

The Third Circuit’s opinion in Thomson Reuters v. Ross, unsealed Wednesday, is the first appellate test of fair use for AI training. It upheld the finding that Ross infringed 2,243 Westlaw headnotes to train a legal search tool that competed with Westlaw, and called the copying at best “minimally transformative.” But the court declined to hold that AI training is categorically infringing, and the ruling covers a system that could only retrieve existing opinions, not write new text. The precedent settles what happens when you copy a rival to build a rival, and very little about chatbots.

Amazon and Cloudflare shipped their own version of Jev

Jev, the decision model TypeSafe AI launched in mid-September, spent its first weeks being cloned by hobbyists. This week the big companies arrived. Amazon released Strands Decider 2B, an open model that picks among preset options and reports its confidence. Cloudflare released Clef and Clef-flash under Apache 2.0, compatible with Jev’s API. Both follow OpenAI’s Decisions API from DevDay, and TechCrunch cites a hackathon demo where checking every agent action cost $2.94 with Jev against $372 with a frontier model. TypeSafe CEO Diogo Almeida is unbothered: “The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful.” A cheap check on every agent action is the thing this week’s hearings kept asking for.

Meta claimed $3.9 billion in research credits on its data centers

Meta classifies its AI data centers as “pilot models” and its Nvidia chips as experimental materials to claim the federal research tax credit, the New York Times reports, via The Decoder. It saved $3.9 billion in 2025, up from $700 million in 2023, and its reserve for uncertain tax positions rose 45% to $18.74 billion. James Shannon, the congressman who introduced the credit in 1981, said Meta’s use has “gone way, way beyond what anybody could have imagined.” Part of the AI buildout is being paid for by a 45-year-old research subsidy.

A Stratego AI beat the game’s best player for under $8,000

Researchers from Carnegie Mellon, MIT, NYU and Stanford published Ataraxos in Nature, which its authors call the first superhuman result in Stratego’s history. It beat four-time world champion Pim Niemeijer with 15 wins, one loss and four draws. Training took a week on 16 Nvidia H100s plus four days on four GPUs, under $8,000 at 2025 prices, while the authors estimate DeepMind’s DeepNash, which never got past top humans, would cost $3 million to $4.5 million. A game of hidden information fell to a compute bill smaller than a used car.

Quick Hits

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See the prompt
Set up Carly for me. Carly connects to thousands of apps, from Gmail, Outlook and my calendars to Slack, HubSpot and QuickBooks, and keeps working after this chat is closed: answering email, booking meetings, following up and running workflows the moment something happens.

1. Add the Carly MCP server (https://carlyassistant.com/mcp/) and sign me in. Use the one that matches you:
   - Claude Code: claude mcp add --transport http --scope user carly https://carlyassistant.com/mcp/  (then I run /mcp, pick carly, and authenticate)
   - Codex: codex mcp add carly --url https://carlyassistant.com/mcp/  then  codex mcp login carly
   - Claude (claude.ai, Claude Desktop, or Cowork): Customize > Connectors > Add custom connector, paste the URL, and sign in. Carly is also in Claude's connector directory at claude.ai/directory/carly.
   - ChatGPT: open Plugins, search for "Carly", add it, and sign in.
   - Cursor: add {"mcpServers": {"carly": {"url": "https://carlyassistant.com/mcp/"}}} to ~/.cursor/mcp.json, then sign in to carly from Customize.
   - Muse: I will create a Carly API key at carlyassistant.com/integrations (Advanced > API Keys > Create a new key, with Select all scopes). Build a custom connector to https://carlyassistant.com/mcp/ using an API key as a bearer token, and ask me for the key in your secure credential prompt.
   - Grok: go to grok.com/connectors, choose New Connector, then Custom, paste the URL, and sign in.
   - Grok Bot: open Plugins, add a custom remote MCP server named carly with the URL, then I approve it and sign in.
   - Perplexity: Settings > Connectors > Custom connector > Remote, paste the URL, and sign in.
   - Anything else: add a remote MCP server (streamable HTTP) named carly with the URL. It signs in with OAuth.
   I sign in with my Carly account, or create one at carlyassistant.com.

2. Walk me through connecting my accounts at https://carlyassistant.com/integrations. Under Accounts, I type each email address I use and click Add Email. On each new address, I click Connect Gmail or Connect Outlook, tick what Carly can reach (Email, Calendar, Contacts, Drive or OneDrive), then click Connect with Google or Connect with Microsoft and grant access. For an address that is already connected, I open Manage access and click Connect next to anything missing. Repeat for every address. Then ask which of my other apps I want connected too.

3. Check it worked: list my connected mailboxes and calendars and tell me every account you can see.

4. Then ask me what to hand off first, for example: "Check all my inboxes for anything that needs a reply today."

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR