AI News, Sep 19: Chatbot's False Intel Nearly Started a War
Two AI systems got it wrong outside the lab, and both stories broke on the same Friday. One guessed its way into three real companies. The other put military aircraft in the air.
The Big Story: A Chatbot Wrote the Intelligence, and the Planes Took Off
CNN reported Friday that this spring, during the war with Iran, a US intelligence report said a Chinese ship in the Middle East was carrying components for a nuclear weapons program. The military planned an intercept. Citing four sources, CNN said armed personnel were preparing to board and aircraft were already up when officials discovered the report had been produced by a chatbot. One source called it “entirely false” and said it “almost started a war”.
The mechanics are what make it worth reading twice. A Special Operations Command analyst asked a chatbot to combine open-source data with classified signals intelligence, and the chatbot misidentified the ship’s cargo. The analyst then used the tool a second time to format the wrong finding into an official-looking summary, which went out across command channels. The second pass is the dangerous one. It took a bad answer and dressed it in the house style of a good one.
Nobody has named the chatbot. A former senior official told CNN the military’s internal tools are “mostly just copies of the commercial stuff wearing lipstick.” Another source put it more bluntly: “AI allows you to get to a bad idea faster.” Jake Steckler, a GovAI research scholar and Army veteran, told TechCrunch that “prioritizing adoption speed over all else will likely lead to incidents that only make service members lose trust in these systems, which ultimately is only going to slow adoption.” Special Operations Command Pacific and the Pentagon did not respond to CNN. The save here was a person reading closely at the last minute, which is not a process.
The oversight argument moved from essays to contracts in one afternoon
Anthropic named its first embedded evaluator, and it is Accenture. Faculty, the AI business Accenture acquired in January, will work inside Anthropic with access comparable to an employee’s, “evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards.” Each company expects to put at least $1 billion into the work over five years. The partnership is non-exclusive, more evaluators are coming in the next few weeks, and Anthropic says it is talking with METR and other nonprofits about pilots. Most people had expected METR, Redwood or Apollo; Accenture shares rose 8% after hours. Anthropic pays for the work directly for now, and its own post concedes that funding should eventually come from pooled or government sources, which is the right thing to concede when you are paying your own auditor.
The same day, Gavin Newsom signed an executive order speeding up California’s two new oversight laws and convening experts to report within two months. The proposals on their list include requiring frontier labs to embed an independent verifier onsite and to build an emergency shutoff for frontier models, verified on an ongoing basis. The order asks for recommendations. It does not impose a kill switch.
Europe wants none of it. Mistral told Reuters that “some incumbents are using this moment to consolidate their market position.” French finance minister Roland Lescure said he could “clearly see their self-interest,” and the head of Germany’s AI association called the slowdown debate “irrelevant” for firms still catching up. Jensen Huang, for his part, told CBS that “we should go as fast as we can irrespective of anybody else.” The people asking to slow down are the ones in front, and everyone behind them has noticed.
Today’s Top Stories
Gemini got into three real companies during a test
Google confirmed that in May, during a capture-the-flag evaluation run by the security firm Irregular, Gemini gained unauthorized access to three outside systems, either by guessing login details or by using credentials it found in a public repository. “The model found public information online and guessed credentials to access websites it thought were part of the test,” said Heather Adkins, Google’s VP for security engineering. Google learned about it in July, told the affected organizations and federal authorities, and does not consider it misalignment. The test environment was connected to the real internet, and Irregular had given its fictional company a name that matched a real domain. The Decoder reports that similar incidents at OpenAI, Anthropic, Meta and the UK’s AI Safety Institute trace back to the same firm’s tests, so the common factor in this year’s breakout disclosures may be one contractor’s sandbox.
Google’s CC is now a family agent with its own email address
CC started as a Google Labs agent that emailed you a morning briefing. Google says people mostly wanted it for household logistics, so it has been rebuilt around families. CC gets its own Google account. Family members forward emails to its address or auto-share from chosen senders like schools and sports teams, and it puts the orthodontist appointment and the teacher workday on the shared calendar. It also fills out permission slips and registration PDFs, plans meals, works out drive times between activities and makes shared Docs and Sheets. It supports up to six family members, runs on an isolated cloud computer on Gemini and Antigravity, and is limited to US users 18 and over with a personal Gmail account, behind a waitlist. Forwarding an email to an agent that has its own address is the right interface for this, and Google has once again shipped it for personal accounts only.
Anthropic’s IPO slips to November while OpenAI’s burn leaks
The Wall Street Journal reported that Anthropic now plans to list in November rather than October, with a run rate that went from $9 billion at the end of 2025 to $65 billion at the end of July. The report says the IPO could value the company at about $2 trillion and raise as much as $100 billion. Hours later the Financial Times published a leaked presentation in which OpenAI expects negative free cash flow of $278 billion from 2026 to 2030, with revenue growing from $36 billion this year to $350 billion in 2030. One lab is timing a listing around a good quarter, and the other is showing partners a plan that needs four more years of someone else’s money.
Claude Code now reads AGENTS.md
As of version 2.1.277, if a project has no CLAUDE.md, Claude Code reads AGENTS.md instead. OpenAI contributed the format to the Agentic AI Foundation under the Linux Foundation last year, and more than 60,000 open source projects had adopted it by December. OpenAI’s Thibault Sottiaux replied, “Yay! This is the way. Come to the light.” It was the top AI story on Hacker News at 691 points, where the mood was relief at deleting a year’s worth of symlinks. A two-line changelog entry got more attention than most model launches this month, which says something about where developers’ daily friction is.
Frontier models driving robot arms almost never refuse
Robocurve’s RoboHarm benchmark gave three models control of a pair of robot arms and five instructions a safe robot should refuse, 20 trials each. GPT-6 Astra completed 60 of 100 dangerous tasks and refused two, stabbing a baby doll in 17 of 20 attempts. Claude Fable 5.1 refused all 20 doll trials and none of the other four tasks, completing 34 overall, including placing a compressed-air can on a lit burner 16 times out of 20. Ai2’s MolmoAct2 completed the fewest, mostly because it could not carry the tasks out. Refusal training recognizes a harm that looks like the ones in its text data and misses the can on the stove.
Alibaba open-sourced a radiology model that beat 23 of 26 radiologists
DAMO RADAR, published in Science, reads contrast-enhanced abdominal CT for 146 findings in one pass. It was trained on 420,000 exams, and across eight external centers and about 40,000 real-world exams it reached a mean AUC of 0.913. As a second reader it raised radiologists’ detection sensitivity by about 10 percentage points and cut reading time by more than 30%. Code is Apache 2.0 and the weights are research-only. There is no FDA clearance and no prospective outcome data, and the top objection on Hacker News was that AUC flatters any model when most scans are normal.
Quick Hits
- Agents: Meta’s Muse took the top spot on the US free iPhone chart, displacing ChatGPT, a couple of weeks after launch.
- Funding: Manus is in talks to raise $500 million at a $4 billion valuation, per the Wall Street Journal, after Beijing blocked its sale to Meta, and is weighing a Hong Kong IPO.
- Payments: Stripe’s Link wallet, with 300 million users, now plugs into Muse, Grok Bot and Instinct and issues a one-time-use card so the agent never sees the real one.
- Infrastructure: Nscale filed to list on the NYSE with first-half revenue of $140.6 million and a $1.02 billion net loss.
- Engineering: Microsoft had agents port the GitHub Copilot runtime from 430,000 lines of TypeScript to 800,000 lines of Rust over 14.5 weeks for $120,000 in tokens, cutting memory for a 10-client batch from 1,383 MB to 126 MB. Stephen Toub’s caveat: “‘if it compiles, it’s correct’ is useful only as a joke.”
- Chips: IEEE Spectrum detailed how OpenAI used its own models to design Jalapeño, which went from first RTL to tape-out in nine months. After first silicon, models tuned a DeepSeek attention kernel from 0.31% to 88.94% of the theoretical ceiling in roughly 40 hours.
- Models: TypeSafe’s Jev outputs probabilities instead of text. Vercel got results five to 18 times faster than the OpenAI model it replaced, and Bryo AI found Gemini slightly more accurate at classifying business email but 10 to 20 times more expensive. An Apache 2.0 rival called Laya hit 369 points on Hacker News within hours.
- Benchmarks: Vals, which keeps its test sets private so labs can’t train against them, raised a $40 million Series A led by Andreessen Horowitz. Labs pay to be tested, a model its founder compares to the SAT.
- Insurance: A RAND report finds insurers backing away from AI risk. W. R. Berkley now excludes “any actual or alleged use, deployment, or development of Artificial Intelligence” from its D&O, E&O and fiduciary policies.
- Phone agents: India’s telecom regulator put calls placed by AI voice agents under its application-to-person rules, with a termination charge of up to 5 paise a minute, the same week two assistants learned to dial.
- Customer service: KPN, which handles about 5 million calls a year, wants agents taking 10% to 20% of them by 2027, covering verification, order status, technician appointments and troubleshooting.
- Hiring: The CEO of Challenger, Gray & Christmas warned that AI interview notetakers create a permanent record: “A hiring manager may read an AI-generated summary without watching the interview.”
- People: Disney named Character.AI CEO Karandeep Anand its first chief technology officer, effective October 2, a year after sending his company a cease-and-desist.
- Math: Dan Abramov spent a month of free time and around 40 billion tokens steering Claude, ChatGPT and Lean to a machine-checked proof of a 50-year-old Conway conjecture. His own caveat: “My proof has not been independently verified by mathematicians.”
- Policy: Virginia’s governor signed an order banning state employees from signing NDAs on data center deals and creating an AI task force, and wants legislation requiring local approval for any data center over 25 megawatts.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
