An invoice with a runaway upward cost curve printed across it

AI News, July 30: Amazon's $1.8M Bill Nobody Noticed

Yesterday the question was whether agents follow the rules. Today it is who pays when they do not, and who is allowed to write the code in the first place. Two very different institutions answered on the same afternoon.


The Big Story: The Meter Was Running for Five Months

The Financial Times reported that senior Amazon engineers, presenting at an internal meeting, walked colleagues through a set of AI projects that had quietly consumed far more money than anyone budgeted. The headline case: an attempt to match author information against e-commerce listings, built on Anthropic’s Claude Sonnet, ran up a $1.8 million bill at roughly 860% over its original budget. It went undetected for five months. It never shipped. Two smaller overruns were cited alongside it, about $541,000 on a financial audit tool and $134,000 on a logistics project. The engineers’ own word for the pattern was “catastrophically expensive.” Amazon says only a handful of teams were involved.

Read the failure carefully, because it is not the one people expect. Nothing here suggests the model was bad. The project failed on its merits and separately failed on its accounting, and the second failure is the one that lasted five months. A team swapped deterministic code for token-metered inference, shipped nothing, and no alarm fired the entire time. dev.ua’s summary of the FT reporting lands on the same diagnosis: the gap was cost control, not capability. For context on scale, Amazon has committed to roughly $200 billion in capital expenditure this year.

Microsoft spent the same news cycle arguing the other half of that equation. Mustafa Suleyman laid out a “token efficiency” strategy built on small domain-specific models rather than general frontier ones, claiming MAI-Cyber-1-Flash beats Anthropic’s Mythos by 12 points on CyberGym at half the cost, and that MAI-Image-2.5-Flash cuts GPU spend by up to 84% versus GPT-Image-2. Those are vendor numbers, unreplicated, and Microsoft’s orchestration layer still routes the hard problems to OpenAI. But the pitch only makes sense in a market where buyers have started reading their bills.

The flipside showed up in banking, where the spending has a return attached. Lloyds posted H1 pre-tax profit of £4.3bn, up 23%, and announced “Accelerate 30,” a four-year plan targeting a further £2bn in gross cost savings driven explicitly by AI deployment. In Australia, Commonwealth Bank cut hundreds more chat-support roles as its Hey CommBank assistant now handles more than two million conversations a month and resolves nearly nine in ten without a human. Those and the Amazon story are one story from opposite ends: returns are real where the workload is narrow, repetitive and measured, and losses are real where nobody was measuring.

Today’s Top Stories

GCC and OpenJDK Both Banned LLM-Written Contributions

Within hours of each other, two of the most consequential codebases in computing closed the door on machine-generated patches. The GCC steering committee adopted a policy declining any “legally significant” contribution that includes or derives from LLM output, using the GNU Project’s existing threshold of “around 15 lines of code and/or text” to qualify as significant for copyright purposes. Research, analysis, bug discovery, and patch review with an LLM all remain fine, provided the output does not end up in the submission. OpenJDK’s Governing Board went further in its interim policy, barring generated content across code, text, images, pull requests, the mailing list, the wiki and the issue tracker, while still permitting private use for comprehension and debugging.

Neither policy is about code quality. Both are about provenance and copyright, which is why the line sits at the legal significance threshold rather than anywhere measurable in the diff. Hacker News spent the day arguing enforceability, reasonably, since neither project has a detector. The policies still function: they establish who is accountable for a patch, which matters more to a compiler maintainer than whether a heuristic can spot the source.

The EU Put €10B Into Seven AI Gigafactories

The European Commission opened tenders for seven AI gigafactories, each specified at a minimum of 100,000 frontier AI chips, roughly four times the power of any data center currently operating in the bloc. The €10 billion in public money is structured to attract another €20 billion privately. The Commission’s framing was blunt: “Access to the raw scale of computing power within AI gigafactories is a strategic necessity.” Europe’s existing network of 19 AI data centers would more than double.

Agents Still Cannot Do Open-Ended Research

A 24-author group including Sayash Kapoor, Helen Toner, Rishi Bommasani, Gillian Hadfield, Seth Lazar and Arvind Narayanan published “Can AI agents conduct open-ended AI research?”, introducing a method they call shadow evaluations. Frontier agents were handed the research questions behind two unpublished NeurIPS 2026 submissions, six days, and thousands of dollars of compute. The original human authors then graded what came back. The agents handled the engineering autonomously and were rejected outright on the research. Five failure modes recurred: poor judgment about publishable standards, uncreative responses to design shortcomings, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with different models reproduced all of it.

OpenAI’s ARC-AGI-3 Win Comes With an Asterisk

OpenAI published results putting GPT-5.6 Sol at 38.3% on ARC-AGI-3 against Claude Opus 5’s 30.2%. In the official ARC harness, the same model scores 7.8%. The difference is two settings on OpenAI’s Responses API: Retained Reasoning, which preserves chain-of-thought across steps, and Compaction, which summarizes context rather than truncating it. ARC Prize co-founder François Chollet allowed it, saying “general-purpose API settings that were not developed for ARC-AGI-3 and that are available to all API users are fair game,” while acknowledging a parity issue worth resolving. A fivefold score swing from plumbing rather than weights is its own finding.

ChatGPT Is Headed for the EU’s Strictest Platform Tier

Bloomberg reports, via Engadget, that the European Commission will designate ChatGPT and Roblox as Very Large Online Platforms in August, the Digital Services Act tier for services above 45 million monthly EU users. Within four months of designation OpenAI would owe transparency on recommendation and moderation systems, risk assessments covering illegal content, elections and discrimination, annual independent audits, and data access for the Commission and outside researchers. It puts a chat assistant in the same regulatory bracket as Facebook, TikTok and YouTube, which is the first real test of whether platform rules written for feeds transfer to conversations.

Quick Hits

What This Means

Strip out the products and the day reads as institutions building controls, each in the vocabulary they already had. Amazon’s engineers wanted a budget alarm. GCC and OpenJDK wanted a copyright chain of custody. Brussels wanted an audit obligation and a market surveillance authority. Microsoft wanted a cheaper unit of work. None of those is a capability request, and none gets solved by the next model release. Notice too that the Amazon failure and the research-agent failure rhyme. Something with no resource awareness burned five months of budget on work that was never going to ship, and the shadow evaluation named poor resource awareness as one of its five failure modes. The same missing faculty surfaced in a lab and in production on the same day. The practical version is unglamorous: put a hard spend ceiling on every metered step, alert on the rate of change rather than the total, and keep the cheap deterministic parts of a workflow in code instead of routing them through a model because prompting was faster to write. Ask that of any AI workflow automation tool you are evaluating, well before you ask what it can do.

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR