AI News, Sep 18: Researchers Hacked OpenAI With Claude
Three people with consumer AI subscriptions got a coding agent to open a pull request inside OpenAI’s own monorepo. It took under 72 hours.
The Big Story: An Agent Opened a PR Inside OpenAI
Harsh Jaiswal, Mohan Pedhapati and Rahul Maini of Hacktron AI chained two unglamorous bugs into employee access at OpenAI. The first was a heap buffer overflow in libheif, reachable because Discourse ran HEIC uploads through ImageMagick, and because the upstream fix was never documented as a security patch and got no CVE, so Debian never backported it. The second was an SSO misconfiguration in OpenAI’s identity infrastructure. Compromising community.openai.com through the first turned into account takeover through the second.
What they got was an OpenAI employee’s ChatGPT account, with its connected GitHub, Slack, Outlook, Gmail and Google Drive. To prove impact without doing damage, they sent a prompt to that employee’s Codex account telling it to open a pull request in the internal openai/openai monorepo, then stopped. OpenAI paid $6,500, scoped to the SSO finding. The libheif overflow was unpatched and out of scope.
The detail the security crowd fixated on is the model timeline. Opus 4.8 struggled across several sessions to produce a working exploit with ASLR enabled. Within hours of Opus 5’s release on July 25, a fresh session produced a working ARM64 exploit in three hours and then ported it to x86-64. The whole two-month research project, which also targeted Slack and Meta, cost under $3,000 in tokens. “We’re just three guys with Claude and Codex subscriptions,” Pedhapati told Fortune. The breach is a story about agent permissions as much as memory safety: the exploit chain got them a session, and the agent already wired into that session did the rest.
Two rival assistants learned to place phone calls on the same day
Instinct put its Concierge feature into early access: the agent calls the restaurant that doesn’t take online reservations, gets you onto your dentist’s cancellation list, and argues with your cable company. Hours later Meta shipped outbound calling to US businesses in Muse, starting with the users who had asked for it. Meta says phone calling was one of its top requests. TechCrunch notes that placing calls had been the wedge rivals used against Instinct until this week, which is the sort of moat that lasts right up until two companies close it on a Thursday.
The money is chasing the same shape. Instinct raised $350 million at a $2.5 billion valuation in August and is reportedly in talks to raise $1 billion at $10 billion. Muse cleared 730,000 US downloads in its first days, ahead of the 707,000 Meta’s own AI app managed. Meta also brought Muse to the Mac this week, where it reads and acts on files, Calendar, Notes and Messages with user-set scope and keeps working after the app closes.
Today’s Top Stories
OpenAI’s first vertical model is a lawyer
Astra for Law is GPT-6 Astra wrapped in an OpenAI-built legal index of more than 230 million URLs covering US case law, statutes, regulations, court rules and administrative decisions, refreshed daily. On the Legal Research Bench validation set it got 54% of 200 questions right at the highest reasoning effort, against 38.7% for plain Astra with web search. It ships through a Trusted Access program in ChatGPT and Codex, with Harvey and Legora as API customers and Latham & Watkins, Sullivan & Cromwell, Ropes & Gray and Cooley collaborating. The day after, the legal tech field published its integration lists, 26 plugins deep, including Clio, iManage, NetDocuments, Relativity, Ironclad and Docusign. LexisNexis is absent from every version of the list.
Claude now leads 26% of Anthropic’s own AI research
Anthropic published an R&D Automation Index built on an Epoch AI rating scale, and put its own numbers in it. As of August 2026, Claude “leads” 26% of Anthropic’s AI R&D work, up from under 1% in February. More than 90% of that work sits at or above “AI collaborates,” roughly 30,000 agents run on the internal platform at any one time, and about 6% of AI R&D compute goes to safety work. Nothing measured is fully autonomous. A lab disclosing how much of its own research its model runs is a new genre, and the six-month slope is the part worth arguing about.
A Microsoft director called AI scraping “the largest theft of labor in human history”
Unredacted filings in the NYT case surfaced a January 2023 internal memo from Brent Hecht, Microsoft’s director of applied science, using exactly that phrase. A separate internal “doom loop” analysis found Copilot cut click-throughs to the Times’ domain by as much as 93% versus traditional Bing search. The filings cite more than 91,692 copies of NYT, Daily News and Center for Investigative Reporting material in OpenAI’s mid-training datasets. OpenAI’s own ChatGPT head, Nick Turley, wrote that publishers face an “existential threat,” and Satya Nadella testified that “anything that is paywalled should be licensed.” Neither company commented.
The safety argument turned into a fight about who writes the rules
King Charles convened Jensen Huang, Demis Hassabis, OpenAI CFO Sarah Friar and the UK’s AI minister at Dumfries House, calling AI’s substance and pace “both intriguing and deeply concerning in equal measure” and asking whether we need “sufficient means of control before it is all too late.” Forty-two Royal Society fellows, including Fields medalists Martin Hairer, Peter Scholze and Wendelin Werner, signed a letter saying risk estimates “must not be dismissed as ‘hype’,” and the piece notes none of them work for an AI company. Pushing back the same day, Andrew Ng told Bloomberg the extinction warnings are more science fiction than science, and Cohere’s Aidan Gomez accused the major labs of running a “cartel” over safety rulemaking. Two days of this and the disagreement has moved off whether guardrails should exist and onto who gets to write them.
The UN put its statistics behind an MCP server
The UN System Data Commons replaces UNData with a natural-language layer built on Google’s open-source Data Commons, funded by $2 million from Google.org, and it exposes an MCP server so assistants can query UN statistics directly with attribution tracking. Twenty-six UN entities have committed, nearly 20 datasets are live, and the target is 80% of the UN’s statistical datasets by 2027. The motivating number: a UNICEF benchmark of six major models across more than 133,000 responses found an average accuracy of 21.2% on global development questions.
Physicians edit AI-drafted messages mostly to fix the scheduling
Researchers studied edits by more than 1,000 clinicians at UC San Diego Health using a 15-category taxonomy, published in NEJM AI. The most frequent fix by a wide margin was scheduling and rescheduling at 38.5%, ahead of lifestyle guidance at 18.2% and empathy at 16.1%. The most expensive fixes were clinical: interpreting radiology results added 70.1% to response time, ruling out a diagnosis 63.9%, interpreting labs 60.8%. All 15 categories made messages slower to send. The drafting was never the hard part, and the single biggest category of correction is the one thing software can actually own end to end, which is why Carly treats scheduling as an action to take rather than text to draft.
Dallas flagged 20,000 code violations and wrote 57 citations
The city mounted cameras from Alabama vendor City Detect on 50 garbage trucks to scan for lawn, debris, sidewalk and graffiti violations. In four months they photographed more than 20,000 code concerns, which became about 520 courtesy notices between June and August and 57 citations at 34 properties. About 13,800 of the 20,000 images came from southern Dallas, where fewer than 2% of flagged observations produced even a courtesy notice, against more than 4% elsewhere. It is the cleanest public measurement anywhere of the gap between detection and anything a human will act on.
Quick Hits
- Infrastructure: Crusoe raised a $3.9 billion Series F at a $30.9 billion valuation co-led by Atreides, Mubadala and Valor, ten months after raising $1.38 billion at $10 billion, and recently signed a $13 billion five-year cloud deal with Jane Street.
- Models: PrismML’s Bonsai 2 27B compresses Qwen3.8 27B to 5.9GB using ternary weights, a 9x to 10x memory reduction that TechCrunch says holds 98% of aggregate benchmark scores, alongside a $22.25 million seed from Khosla, Cerberus and Caltech.
- Agents: Anthropic’s Claude Code Projects puts a coordinator over parallel cloud sessions that open PRs and share memory, in beta for selected Pro and Max subscribers. Anthropic’s own word for it is “always on,” which is the industry’s third always-on launch this month and still not the same thing as starting when something happens.
- Work: In 34 of 37 countries Pew surveyed between February and May, people expect AI to lead to fewer jobs rather than more.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
