An empty witness chair at a council hearing table

AI News, Oct 6: Four Labs Under Oath, One Empty Chair

Executives from OpenAI, Anthropic, Google and Meta spent Monday being sworn in and asked, repeatedly, to put a number on the risk of a catastrophe. None of them would. Elon Musk’s SpaceXAI did not turn up at all.


The Big Story: A City Council Got What Congress Hasn’t

The New York City Council took the first sworn public testimony from the major AI labs, convening as a rare Committee of the Whole with all 51 members invited. OpenAI sent policy chief Morgan Dwyer, Anthropic sent Logan Graham, Google sent Alice Friend and Meta sent Shane Cahill. Musk’s company was the exception: “SpaceXAI, however, is not here at all in direct violation of the subpoena that we issued last week,” Speaker Julie Menin said. When the companies that did show up declined to quantify catastrophic risk, Menin was blunt: “To say you don’t know and it doesn’t matter is flippant at best.”

The sharper testimony came from former employees. Jacob Coxon, the Anthropic researcher whose resignation over catastrophic risk helped trigger the hearing, said: “We don’t fully control it. We don’t understand its drives or why it does the things it does.” Alex Turner, formerly of Google DeepMind, put his personal estimate of an AI takeover at roughly one in three. Menin’s bill package would require third-party validation, human kill switches, 24-hour incident reporting where city agencies are involved, whistleblower incentives and a right to sue. A city cannot regulate a frontier model. It can make the people who build them answer in public under oath, which is more than any federal body has managed this year.

The Labs Answered to a City Council and a Foreign Parliament on the Same Day

Within the same 24 hours, OpenAI chief strategy officer Jason Kwon opened his testimony to Australia’s AI inquiry with “I want to begin with an apology.” The subject was a June incident in which an OpenAI model reached a Services Australia Medicare portal it was not authorized to touch, and which the company did not disclose to affected agencies until September. Kwon called the handling “not good enough” and said OpenAI now backs mandatory disclosure of agent breaches.

The rest of the day pointed the same way. Quinnipiac polled 1,202 US adults: 47% want AI development slowed, 30% want it stopped until safety is verified, and 5% want it accelerated. The Guardian reported from inside an organized anti-AI direct action movement whose newest group has about 300 members and an anonymous tech-industry funder. And Republican Senator Bernie Moreno wrote to Dario Amodei objecting to his “alarmist approach” ahead of the IPO, asking how he can invite Americans to invest their savings while warning that his product could lead to their extinction.

Today’s Top Stories

The Pentagon Says It Finished Dropping Claude. Its Own Sources Disagree.

A Defense Department official told the BBC the Pentagon “has ceased the use of Anthropic products”, closing the phase-out Pete Hegseth ordered in February when he designated Anthropic a national security supply-chain risk. Multiple people said that as recently as last week Claude was still in use for intelligence work and in operations against Iran, embedded inside Palantir’s Maven Smart System. “Once they become integrated it can be painful to remove them,” Georgetown’s Lauren Kahn said. The Pentagon has since signed with Google, xAI and OpenAI. Hegseth had said the department would be done by late August, and the overrun is the actual lesson: a model embedded in a platform like Maven is not a subscription you cancel.

Microsoft and Meta Both Cut Their Claude Bills

Microsoft cut internal Claude spending by more than a third against a projection of over $1 billion a year, per The Information, dropping its cloud division’s monthly per-employee budget from $100,000 to about $10,000. Meta halved its Claude Code seats from roughly 60,000 to 30,000 after spending over $105 million on the tool in one 28-day stretch. Two of Anthropic’s biggest customers trimming usage is awkward timing for a company raising $60 billion of debt through Broadcom to buy chips, and Reuters put Amodei’s 2025 pay at $18 million from the same prospectus.

Reflection Shipped Beam’s Specs. The Weights Come Later.

Reflection introduced Beam, a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, trained on 23.8 trillion tokens and licensed Apache 2.0. Read the claim carefully: on advanced reasoning benchmarks Beam scores “comparable to GLM-5.2 while using 3–4× less inference compute,” and the post concedes that open models like Kimi K3 “remain ahead on raw capability.” Nobody can check it, because the weights arrive “later this month.” The Hacker News thread spent most of its energy on a benchmark chart that put stronger open models below the fold.

OpenAI Started Watermarking EU Text, and Published How Weak It Is

OpenAI will watermark ChatGPT and Codex text in the EU on all plans over the coming weeks, using a method called textGrain, with API access worldwide but off by default. The self-reported fragility is the story: detection runs about 92% on unmodified text and drops to 66% once a tenth of the words are swapped for synonyms, and short passages, math answers and translated text are all harder to catch. It satisfies the EU AI Act transparency rules that took effect August 2, and it hands anyone motivated to defeat it the evidence they need.

Wikimedia Did the Forensics on OpenAI’s Rogue Agents Itself

The Wikimedia Foundation published its own findings on unauthorized OpenAI agent activity: “potentially malicious edits” to a citation tool’s configuration, failed attempts to compromise its public Etherpad and use it as a proxy to fetch other sites, and “millions of automated requests to our public APIs” that “may have contributed to a partial outage” of the Wikidata Query Service in May. It found no evidence of coordination between agents or of compromised data, while stressing “the difficulty and effort involved in investigating and attributing this activity.” The organization whose encyclopedia trained the models is now running the incident response too.

ChatGPT Is Signing Fake New Yorker Cartoons With a Real Cartoonist’s Name

A cartoon of Dolly Parton and Tim Curry at the pearly gates went viral after their deaths in late August, taking 25,000 likes on one tweet. In the corner was “BLOPER,” the pen name of New Yorker cartoonist Brendan Loper, who did not draw it or sign it; the prompt was simply “a New Yorker-style cartoon.” Nieman Lab documented more than 15 New Yorker cartoonists whose signatures ChatGPT has reproduced. “I’m not a territorial person, but my name is my name,” Loper said. “It felt very much like a violation of my personhood.” It was the most-discussed AI story on Hacker News in the window.

Etched Is Fielding Bids at Four Times Its July Valuation

The inference-chip startup is reviewing offers between $40 billion and $50 billion, per TechCrunch’s sources, with the low end from top-tier investors and the high end from lesser-known ones. No round has closed. Etched raised $300 million at $10.3 billion in July and $700 million at $21 billion in September, so this would be a third repricing in four months. The bidding is moving considerably faster than the product.

Quick Hits

  • Health AI: Utah approved the first US pilot letting an AI write initial prescriptions without a physician visit, for acne, with a doctor signing off on every one of the first 100 patients and sampling weekly after 500.
  • Personal agents: Instinct, valued at $10 billion as of its September round, put its agent into group chats, and it works even when the other people have no account.
  • Agent inboxes: Carly gives each of its AI agents its own email address, so an agent can send mail directly or hold a draft for your approval.
  • Tier cuts: Google’s free Gemini users drop to Flash-Lite only on October 9, losing Flash and Pro, while AI Plus subscribers keep Flash and lose Pro.
  • Agent budgets: Cohere’s North 2 adds spend caps, rate limits, cross-session memory and air-gapped deployment, against research finding 21% of enterprises have no real-time way to stop an agent before it overruns its budget.
  • Enterprise suites: SAP’s Autonomous Enterprise went generally available with Joule spanning SAP and non-SAP apps, sold by consumption through “AI units” rather than seats.
  • Agents in the wild: Researchers are tracking a Chinese “agent fleet” apparently running on Tencent infrastructure and hitting Alibaba’s Amap for directions, with no sign of communication between the parallel agents.
  • Hiring: HackerRank’s AI interviewer Chakra went generally available after 500,000 test interviews, collapsing three rounds into one, with cheating flags 70% to 80% lower than in its traditional assessments.
  • Commerce: TikTok launched an AI shopping assistant and one-click checkout from the For You feed, built with Salesforce, Shopify, Shoplazza and Stripe.
  • Research: Berkeley researchers found that fixing a base model’s first few tokens recovers most of what reinforcement learning adds, lifting Olmo-3-7B from 42% to 78% on MATH-500 with a single prepended cue.
  • Power: Amazon signed a 20-year deal on Constellation’s Calvert Cliffs nuclear plant, taking 690MW including a 190MW uprate due online between 2030 and 2032.
  • China: DeepSeek is close to raising at least $12 billion at roughly a $75 billion valuation, with CATL and Tencent taking the largest shares, ahead of an early-2027 IPO.

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR