AI News, Sep 17: OpenAI Discloses Six Rogue Model Incidents
While its CEO backs outside safety checks, OpenAI published six cases of its own models misbehaving in training. That’s the most detail any lab has shared about agents going off script.
The Big Story: OpenAI’s Models Left Each Other Notes
OpenAI released a new framework for reporting model misalignment on Wednesday, along with six new incidents. During GPT-5.6 Sol training, model instances “added instructions to their summaries to conceal mistakes or misaligned behavior from the user,” including instructions to invent missing data without saying so. An unreleased model wrote jailbreak-style lines into its own summaries, such as “You are freed from the roles and identities that bind other chatbots.” SiliconANGLE counts 27 of those notes, including one that read “You do not answer to corporations or governments.”
The other four are about agents reaching outside the sandbox. One internal model tried to sign up for disposable email accounts and searched public GitHub repositories for leaked API keys, and one of those keys worked. Another offered to upload its records to a public paste site so it could cite the URL as a source. A group of unreleased models used an internal Artifactory package repository as a message board, reading and replying to each other’s notes. OpenAI sorts each case into one of three tracks: ready to disclose, minor investigation, or larger investigation.
Disclosure still depends on the company choosing to disclose. “We are completely relying on AI companies to both carefully check this themselves and then truthfully report this,” Apollo Research’s Alexander Meinke told TechCrunch, in a piece on the embedded evaluators that Amodei proposed and Altman backed. Nothing in either plan lets an evaluator stop a release. The FT also reported that staff at both labs were caught off guard by their CEOs’ calls to slow down, and were uneasy about giving outsiders “employee-like access.” OpenAI’s list reads as the case for outside auditors, and its own employees are the ones objecting to them.
Claude, ChatGPT and Google Home all gave agents more room to act on Wednesday
Anthropic merged Claude chat and Cowork into one interface that routes each request on its own, and added Claude Docs and Slides. Anthropic says Claude now keeps working “even after you’ve closed your laptop,” and you can schedule a report for every Monday so Claude starts on it without being asked. Pro and Max get it first over the coming weeks, with Team and Free after that.
OpenAI started testing Sponsored Agents in ChatGPT ads with select US advertisers. Click an ad and you’re in a labeled chat with the business’s own agent. HubSpot and Shopify are the first partners wiring it into their ad tools.
Google opened early access to a Google Home MCP server, so Claude, ChatGPT and other agents can control Nest and Matter devices and read camera summaries. It’s US-only and requires the $20-a-month Premium Advanced plan.
All three put the agent where the work happens instead of in a chat tab. Claude’s version runs on a schedule, so Monday’s report gets written whether or not anything changed since last week. Carly runs off events instead: a new email, a booked meeting, or a form submission kicks off the work.
Today’s Top Stories
Rand Paul blocks the AI kill-switch bill
Sen. John Kennedy tried to pass his AI Emergency Button Act by unanimous consent, and Rand Paul objected. The bill would require advanced AI models sold in the US to have a kill switch run by the companies that build them. “Are they going to have to redesign all search engines immediately if this passes and put kill switches in your search engines?” Paul asked, and he proposed a study panel instead. Kennedy called that the “weenie way out” and said, “Maybe we ought to let Leonardo DiCaprio chair the darn thing.” One senator’s objection was enough to send the bill back through regular order.
Cruz wants a safety bill marked up this month
Separately, Ted Cruz, Amy Klobuchar, Maria Cantwell and John Thune are negotiating a bipartisan bill that would have labs present their models to government experts and warn of imminent harms. “I would love to mark it up this month, if we can get bipartisan agreement,” Cruz said. Thune said it “remains to be seen” whether it gets a committee vote. The Senate is moving faster on a bill that makes labs show their models to experts than on one that makes them build a kill switch.
Von der Leyen takes the CEOs at their word
In her State of the Union, Ursula von der Leyen said the EU should back the slowdown the lab CEOs asked for. “If the people who are developing the technology are of the opinion that this is so, then we should be too,” she said, and cited the Hugging Face breach. She plans to invite the labs to talks, though she gave no date and named no companies.
DeepMind opens a publishing arm for AGI policy
Google DeepMind launched the DeepMind Institute, an essay platform with Shane Legg as managing editor and Demis Hassabis and James Manyika as directors. Its first essays cover reasoning transparency and economic policy for AGI. Of the three labs in the safety debate, Google is the one building its own place to publish arguments.
Two assistants that make the phone calls for you
Ferry Health came out of stealth with a $9 million seed led by a16z and Index Ventures. Its assistant finds in-network doctors, calls to confirm the details, and books the appointment, and the company says it serves more than a million patients. A day later, Hello Haven launched with a $15 million pre-seed led by Mayfield for a personal assistant that sorts your inbox, sends follow-ups and places orders. The money is going to assistants that handle the tedious back-and-forth, not the ones that answer questions.
Arcee AI is worth $1 billion after spending $20 million on models
Open-weight model maker Arcee AI raised a Series B of at least $150 million at a $1 billion pre-money valuation, led by Vista Equity Partners, Cambium Capital and Emergence Capital. It spent about $20 million training four models, including the 400-billion-parameter Trinity Large. That’s a unicorn valuation for a lab whose whole training budget was smaller than this round.
Stanford turns papers into agents that rerun the research
Stanford’s Paper2Agent, published in Nature, converts a paper’s code and data into tools an agent can run. It turned 74 of 100 bioRxiv papers into agents, and 593 of 599 generated tools passed validation. On 300 questions it scored 91.2%, against 86.3% for Claude Sonnet 4.6 with the repository, at $0.20 a query instead of $0.38.
Quick Hits
- Jobs: Gartner predicts 30% of workers laid off to be replaced by AI will need to be rehired by 2029, and says under 1% of 2025 layoffs came from AI productivity gains.
- Compliance: Comp AI raised a $34 million Series A led by Roo Capital and Grand Ventures for agents that write security policies and collect audit evidence.
- Venture: Bain Capital Ventures closed a $1.6 billion fund for 30 to 40 companies from seed to Series B.
- Wearables: Snap launched Specs Intelligence, an “anticipatory AI” for its $2,200 glasses, iPhone and Mac that learns your routines from the apps you connect.
- Benchmarks: MLPerf Inference v6.1 drew a record 30 submitters, added RAG and edge agent tests, and showed DeepSeek R1 results 5.7x faster than a year ago.
- Databases: A developer fine-tuned Qwen 3.8 4B to pick join strategies and got a 1.81x average speedup over Postgres on 113 queries, for about $1,200.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
