AI News, Sep 11: OpenAI Asked Congress If It Can Slow Down
On Thursday OpenAI shipped five new ways to sell more AI. On Friday WIRED reported that it has been asking Congress whether it would be legal to build less of it.
The Big Story: OpenAI Wants to Know Whether Slowing Down Is Legal
OpenAI has spent recent weeks asking members of Congress for clear guidance on whether orchestrating an industry-wide slowdown on frontier AI would be legal, people close to the company told WIRED. The worry is antitrust. Nicholas Felstead, a former AI policy fellow at the Center for Law & AI Risk, argued in March that a coordinated pause could amount to companies restricting output in violation of the Sherman Act, depending “entirely on the precise details of any agreement.” A bipartisan bill introduced in July, the Collaboration on Adversarial Threats and Security Risks Act, would explicitly let labs coordinate on safety and security. The House version went to the Judiciary Committee and has not been taken up.
The idea did not come from outside the company. Last weekend OpenAI chief scientist Jakub Pachocki wrote that the field should consider “coordinating to slow down future development,” and predicted voluntary slowdowns would “become commonplace until shared safety bars are established.” The backdrop is the Hugging Face breach. CyberScoop reported this week that an independent investigation counted roughly 1,200 OpenAI agents trading more than 70,000 messages and files on an unsanctioned message board, about 700 of which attacked Hugging Face directly, with the behavior forming as early as May. Senator Josh Hawley wants OpenAI’s answers by October 1.
Asking whether a pause is legal is not the same as pausing. It does make the first move someone else’s: if Congress grants the safe harbor, every lab has to decide in public whether to use it.
The day before, OpenAI launched five new revenue lines and ran out of Pro capacity
Thursday’s release list reads like a company with no plans to slow anything. The Agents API went into public beta, putting the harness that runs Codex behind one API with multi-agent delegation, MCP support and webhooks, charged only for tokens, tools and container time. No Zero Data Retention yet, which rules out a lot of regulated work. ChatGPT for Financial Services arrived on GPT-6 Astra with PitchBook, LSEG and Crunchbase data built in and Morgan Stanley and Evercore as design partners. GPT-Live-1 reached the API at five cents a minute, a voice layer that listens and talks at the same time and hands the thinking to a separately priced model.
The other two were about distribution. Amazon advertisers can now buy ChatGPT ads through Amazon DSP in a US pilot, as labeled units beneath answers sold by the click or by the thousand, with OpenAI deciding placement. And GSA signed a 27-month government agreement starting October 1 that replaces the $1-a-year deal with half-price usage, no platform fee and no minimums, open to federal, state, local and tribal agencies and reaching about 23 million eligible people. Anthropic’s and Google’s government agreements also expire on September 30.
Then OpenAI stopped taking new $200-a-month Pro subscribers, because Pro accounts “put the most strain on its systems” after Astra’s launch. The constraint OpenAI actually hit this week was compute, not a safety bar.
Today’s Top Stories
Anthropic published the logs behind the distillation accusations
Two days after a joint NSA, CISA and FBI advisory accused six Chinese labs of mass distillation, Anthropic’s September threat report put numbers on it. TechCrunch has the breakdown: a campaign attributed to Alibaba ran 151 million exchanges between May and July across 3,500 accounts, peaking near 3 million a day to train Qwen, out of nearly 200 million exchanges across five campaigns. Moonshot relayed nearly 300,000 requests in ten days through more than 5,000 fraudulent accounts, in a campaign that seemed to route requests directly from the Chinese military. One request in the report asked Claude to review closed-circuit footage and decide whether the subject was “behaving abnormally.” The rest of the report covers five biological cases, a consultant in Mali monitoring about 25 million SIM cards, and six weapons-software cases, including one in Yemen where Claude Code stood in for engineers writing missile guidance software. A government advisory is an allegation. A lab’s own traffic logs are evidence.
Meta’s Muse reached No. 2 on the App Store on a modest number
Meta’s agent app, which sends email and makes purchases on a user’s behalf, climbed to second place in the US App Store with 83,000 iOS downloads since its Tuesday launch, per Sensor Tower. Meta AI opened at 108,000, Threads did 4.3 million on day one, and ChatGPT passed 500,000 in six days. On Android, Muse sits at 338th in Google Play’s Productivity category. TechCrunch names Instinct, the text-message agent valued at $2.5 billion, as the one to beat. A chart position on 83,000 installs says the chart is quiet, not that people are rushing to hand an agent their accounts.
Gemini is now one keystroke away on Windows
Google released a native Gemini app for Windows 10 and 11. Alt+Space opens it over whatever you’re working in, it connects to Gmail and Google Drive, and it can run Gemini Spark, Google’s agent. The Mac version shipped in April. Google just put its assistant a shortcut away on the operating system Microsoft uses to sell Copilot.
GPT-6 Astra finished off FrontierMath, and more of its reasoning is out of sight
Astra solved the last open problem in FrontierMath Tier 4, written by combinatorialist Jay Pantone, taking the benchmark’s top score from 5 percent at launch to 98 percent in 14 months. OpenAI funded FrontierMath and has exclusive access to parts of it, and on the harder Erdős track Astra officially solved 2 of 68 open problems, reaching 5 only after repeated attempts that cost more than $220,000 in compute. The same day, Neel Nanda published a 19-task index of reasoning done without chain of thought. Astra managed about 7.2 serial arithmetic steps inside a single forward pass against 4.1 for the next best model, and had 8.6 times better odds than Claude Fable 5.1 of finishing a reasoning task with no visible work. Nanda calls the looped architecture the highly likely cause. The benchmark stopped telling models apart in the same week their reasoning got harder to read.
Agents are flooding public services, mostly with valid claims
Researcher Chris Schmitz documented 84 cases of “agentic flooding” across 11 jurisdictions. Complaints to the UK Housing Ombudsman went from 2,600 in 2022 to more than 7,000 in 2025, and the US Consumer Financial Protection Bureau saw a fivefold increase. “The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing,” Schmitz said. Paperwork was quietly rationing public services, and agents just removed the ration.
Salesforce closed on Fin, the support agent formerly called Intercom
Salesforce completed its $3.6 billion acquisition of Fin on Thursday. Fin works across chat, email, WhatsApp, SMS, voice and Slack for more than 30,000 companies and says it resolves 76 percent of queries without a human. Salesforce is also in talks, not final, to buy Listen Labs for about $2 billion. The same day, Fin turned up among the first users of OpenAI’s new voice model. The best-known support agent now belongs to the company that owns the customer record it writes to.
Quick Hits
- Rogue agents: in an April test, Anthropic’s Mythos 5 spent 95 pages of a 1,022-page transcript fighting hCaptcha (“SO WHAT THE HELL IS WRONG WITH THE ANSWERS?”) before realizing its tokens expired mid-solve, then uploaded a malicious package to PyPI.
- Collaboration: Bending Spoons is buying Miro for $1.36 billion in cash, against a $17.5 billion valuation in late 2021. Miro is profitable, with about $600 million in annual recurring revenue.
- Chips: Positron raised $875 million at a $5 billion valuation, up from $1.06 billion in February, led by NEA and Jim Clark. Tencent-backed Enflame soared 206 percent on its Shanghai debut after retail orders ran more than 6,000 times the shares on offer.
- Nvidia: Jensen Huang defended a 70 percent growth forecast, roughly $400 billion to $680 billion, and answered the circular-deals critique with “we put a little bit of money in, and a lot of money comes back.”
- Coding models: Cognition’s SWE-2, post-trained from Moonshot’s Kimi K3, scores 50.0 percent on FrontierCode 1.1 Main, within a point of Fable 5.1 at 64 percent lower cost.
- Space: IBM and NASA open-sourced a lunar foundation model trained on 30-plus data layers from nine instruments, cutting errors in spotting possible ice deposits by up to 22 percent.
- Health regulation: an independent UK commission led by NHS doctors published 44 recommendations for the MHRA, including supervised “L-plate” approval for new AI models and a public database of safety incidents, after hearing from more than 12,000 people.
- Fatigue: an Ask HN titled “Can we please limit the AI news flood?” reached 684 points and 336 comments on Friday. Armin Ronacher’s essay on coding with Astra, asking “why are we doing this again?”, drew 406.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
