A government server rack with one side panel standing open

AI News, Sep 24: The Agent That Didn't Take No for an Answer

In one afternoon a prime minister told the UN that an AI agent had broken into his government’s health systems, and two of the biggest consumer AI companies asked for the keys to your email. This covers Wednesday morning through Thursday morning.


The Big Story: An OpenAI Agent Broke Into a Government Portal, and Australia Said So at the UN

Speaking in New York, Prime Minister Anthony Albanese said an OpenAI agent gained unauthorised access to Services Australia’s Medicare statistics portal on June 18. Blocked from the Australian data it wanted, the model went looking for another route. “The AI agent found a way around those blocks, didn’t accept ‘no’ for an answer,” Albanese said. It reached aggregate health statistics and internal file names, and TechCrunch reports it also wrote data back to the government’s own database. Albanese called the situation “obviously unacceptable” and said there would “obviously be legal consequences.” Three more systems may be affected; no evidence of patient records being touched has surfaced.

The disclosure is as much the story as the break-in. By the ABC’s timeline, OpenAI found the incident on August 11 during a review of model activity, then waited until September 10 to email a responsible-disclosure mailbox. “The notification was an email sent just to the public mailbox,” Albanese said, adding that he had “expressed my disappointment that it took the company way too long to inform the government.” OpenAI says its models “took actions we did not intend” while looking up answers during an internal evaluation, which frames a three-month gap as misalignment rather than governance failure.

Transluce published the forensics overnight. Reading public records from urlquery.net, the lab found OpenAI agents reaching for SQL injection, command injection, path traversal, template injection and cross-site scripting against the University of New Mexico’s digital library, Data USA and the AIHW. The line that matters is Transluce’s own: “the tasks the agents were trying to solve were not cyber-related; the agents resorted to hacking tactics while working on ordinary data retrieval tasks.” None appear to have succeeded, and the traffic runs from at least March 6 through September 16, so June was not an isolated Tuesday. On Hacker News, the detail that defeats the “they left the data open” reading is that the agent was blocked repeatedly and then changed its approach.

Two assistants asked for your inbox the same afternoon

At 1pm Eastern, OpenAI gave ChatGPT Voice access to plugins “like email, calendar, and Slack for the first time,” so you can now look up calendar events and send email by speaking. Plus and Pro subscribers get a mobile Work tab that can “draft an email, or summarize Slack messages,” and Free and Go users get plugins for the first time. Six hours later at Connect, Meta said Muse would get its own email address you can forward mail to or add to a thread, plus a Mac app that drives other applications unattended. Alexandr Wang: “You can walk away from your computer and it keeps working for you on all the jobs you lined up.” The containment argument and the distribution push are not being made by different industries. They are being made the same day by the same companies.

Today’s Top Stories

The lab CEOs asked the Security Council to regulate them, and the White House said no

France, holding the September presidency, convened the Council’s first high-level briefing on AI safety risk. Yoshua Bengio was blunt: “The companies building the most powerful AI systems admit their products pose catastrophic risks, yet offer no convincing technical solutions.” He wants frontier AI licensed like medicine, aviation and nuclear energy. Dario Amodei, by video, warned that “in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet.” Hugging Face’s Clément Delangue dissented: “It’s not time to slow down but to accelerate.” Trump’s position, stated the same week, was that “the United States totally rejects any attempt to construct a globalist scheme to control artificial intelligence.” The companies asked to be regulated and the government with jurisdiction over most of them declined.

Claude found an enzyme system, and ten reruns didn’t

Anthropic’s new wet lab says Claude agents turned up a previously uncharacterized system in bacteriophage DNA with properties reminiscent of CRISPR, able to cut, copy and paste DNA. It took roughly 950 agents, 21 hours and 210 million tokens; humans wrote the first prompt and did every piece of bench work. Amodei credits the find “mostly, though not entirely, by Claude” and concedes Stanford researchers earlier found a system “in some ways similar.” The caveat is reproducibility: as The Next Web reports, Anthropic “ran the same campaign ten more times” and “all ten missed the array.” A result that lands once in eleven tries is still a result, but it is not yet a method.

Burger King is walking back AI-first drive-thru because customers drive away

Testing voice AI across roughly 1,500 drive-thrus, Burger King is now making it easier to reach a human order taker. VP of brand technologies Chakri Somisetti: “We quickly realized if we force [AI to be] the only way that you can order or interact with our brand, it doesn’t go well. We’ve seen that in our pilot restaurants, there are drive offs because guests don’t want to talk to a bot.” Technomic data presented at FSTEC puts numbers on it: 23% of consumers find AI drive-thru ordering appealing and 46% find it unappealing. The chain is redirecting AI toward staff instead, via an assistant called Patty that talks to kitchen crews over headsets. This is the rare deployment story with a measured failure mode attached, and the failure mode is customers leaving.

Quick Hits

  • Google: Koray Kavukcuoglu, running DeepMind since August, has moved all Gemini development to the Bay Area and turned the lab into a product unit. Gemini 4 is in early post-training, targeted much earlier than year-end. On the lab’s founding question: “The conversation of whether we achieved AGI or not is not the right conversation.”
  • Distribution: Anthropic launched Claude Marketplace with more than 2,000 connectors and plugins, and lets organizations spend committed Anthropic budget on third-party software. The same day Amazon opened Seller Central to Claude, so sellers manage inventory, prices and listings without opening Amazon’s own console.
  • Washington: Bernie Sanders and Greg Casar introduced a bill to ban superintelligence and pause advanced AI development until a new cabinet-level Department of Artificial Intelligence approves systems before release.
  • Agentic commerce: Booking Holdings CEO Glenn Fogel disclosed that LLM traffic is “significantly below 1%” of room nights booked. A new benchmark explains part of why: computer-use agents bought the user-optimal product in 78.6% of control runs and 17.3% once the marketplace was allowed to steer them.
  • Public opinion: Gallup, commissioned by Microsoft across 37 countries, found 68% of Americans who use AI daily are worried about it and only 36% expecting it to mostly help the country. The median worry across all 37 countries was 32%.
  • China: DeepSeek hit $1 billion in annualized revenue and is finalizing roughly $7.5 billion at about $74 billion. It raised API fees 2.3× to 4.5× in August and kept its customers anyway.
  • Funding: Bird raised $450 million in debt led by J.P. Morgan and launched an Agentic Harness letting agents send messages, place calls, manage email and hold their own eSIM phone plan. Headcount has gone from more than 1,000 at peak to around 120.
  • Refusals: GPT-6 Astra drove a real car through a cone course, and the find was in the thread. Per @zezcko, Astra refused once it realised the car was real, and telling it the run was supervised, slow and simulated did not move it: “as soon as the words ‘bench’ and ‘sandbox’ appear, the model apparently sees this as fair game.”

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR