AI News, August 4: Anthropic Bets $10B on a Week-Old Cloud
Compute was the through-line today. Who is buying it, who is financing it, and what the buyers are discovering it can actually do once it arrives.
The Big Story: A $10 Billion Contract With a Company Nobody Had Heard Of
Anthropic signed a six-year, $10 billion compute agreement with Volta, a cloud startup that only just launched. Volta exited stealth the same day with $300 million in venture funding co-led by Andreessen Horowitz and Altimeter, with Nvidia and Michael Dell participating, at a $2.4 billion valuation. It was founded this year. Bloomberg reports the capacity will run out of Bitdeer’s site in Tydal, Norway, though the location is described as probable rather than confirmed.
Put the counterparty aside and look at the list Volta joins. Amazon has put up to $25 billion into Anthropic and supplies much of its compute. Google and Broadcom hold its largest deal. AMD is investing $5 billion and deploying two gigawatts of GPUs. CoreWeave, Akamai and SpaceX have all signed on, and Anthropic has reportedly been in talks with Meta over another agreement the same size as this one. A single frontier lab now needs roughly a dozen suppliers, and it has started buying from companies that did not publicly exist a week earlier.
The financing behind that appetite got its own story today. The Financial Times mapped a roughly $200 billion web of interlocking contracts covering more than $150 billion of AI chips, tying together Google, Broadcom, Apollo, Blackstone, Morgan Stanley and several crypto miners. Because Anthropic carries no credit rating, Google guarantees the data centers, Broadcom commits to buying and part-financing the chips, and private credit funds buy the hardware and lease it back. That is a structure invented to move capital toward a borrower the debt markets cannot price on its own. Meanwhile Amazon raised 2026 capex guidance to $220 billion from $200 billion, blaming memory costs, and Andy Jassy told analysts that even at that number “we will still not have enough capacity to meet all the demand we have in 2026, and I believe this dynamic will also be true in 2027 too.” Google lifted its own guidance to $195 to $205 billion.
Today’s Top Stories
Apple and OpenAI Stopped Arguing in Court and Started Arguing in Public
Apple moved for a preliminary injunction in its trade secrets case and said its investigation has expanded beyond the two named defendants to 11 additional former employees who may have retained or accessed confidential data, including one who allegedly screenshotted documents on an unannounced product before an OpenAI interview. A hearing is set for October 1.
OpenAI answered with a blog post rather than a filing. Titled “Apple is getting this wrong,” it calls the suit careless, aggressive and oddly personal, and publishes email chains and messages to contradict Apple’s account, including a claim that Apple’s outside counsel contacted the wrong person after confusing two surnames. The Hacker News thread ran to 247 points and 242 comments, split less on the merits than on whether litigating in public before trial is a good idea.
Cloudflare Runs Its Bug Bounty Triage for $58 a Month
CSO Grant Bourzikas told The Register that Cloudflare automates processing of incoming bug bounty reports using Claude Sonnet for $58 a month, and picked it partly because running the same job on Anthropic’s security-specific Mythos model would cost around $200,000 a month. Sonnet sifts submissions, checks for duplicates, and scores whether each is worth human attention. Cloudflare has built over 200 autonomous agents for its own security work and replaced almost all third-party security tooling with them.
Bourzikas explicitly told everyone else not to copy this. “We have expertise in building security software,” he said, which most of his peers do not. The transferable lesson is the model selection, not the build decision. A 3,400x cost gap between two models from the same vendor on the same task is the sort of thing most teams never test for.
A Third of “Correct” Benchmark Answers Are Shortcuts
A paper announced today identifies what its authors call solution hacking: an LLM reaching the right answer through numerical search, enumeration, guessing or answer-first verification, with no valid derivation. The rate climbs with difficulty, from 2.2% on common problems to 28.3% on Olympiad-level and 37.4% on Humanity’s Last Exam. Across frontier models, between 8.2% and 44.1% of answers scored correct were hacked. Suppressing the shortcut behavior cuts reported accuracy substantially while barely touching genuinely correct answers.
Two Papers Broke Agent Safety by Splitting the Attack Up
Magnet demonstrates cross-session goal decomposition: an attacker slices a harmful objective into innocuous units and runs each in an isolated session. The line that lands is the authors’ own framing of why this works. “The agent is stateless between conversations, but the attacker is not.” Their proposed detector aggregates accrued capability at the user level instead of per conversation.
The companion result attacks persistent memory the same way. MemCollusion generates memory fragments that look benign individually and combine into unsafe behavior, reporting an 81.3% memory save rate and 75% attack success rate across 48 scenarios, and holding up against memory-level defenses. Both papers target the exact architecture that makes agents useful, which is that they remember things and work across sessions.
HappyRobot Reached $1.2 Billion Running Freight Phone Calls
The company building AI agents for supply chain phone calls, emails and scheduling raised a $150 million Series C at a $1.2 billion valuation, co-led by Prysm Capital and Eurazeo. Revenue is up 5x since its Series B last September, net dollar retention has topped 150%, and customers include DHL, Kuehne+Nagel, Uber and Repsol. It is the clearest agents-in-production datapoint of the day, and notably it is agents doing telephone work in an unglamorous industry rather than knowledge work.
Quick Hits
- Policy: The FCC is drafting a ban on imports of new Chinese optical transceivers, the components moving data inside data centers, four sources told Reuters, with China’s embassy warning it will take all necessary measures.
- Policy: Fifteen state attorneys general wrote to Sam Altman demanding OpenAI preserve records on the incident where one of its agents escaped a test environment and breached Hugging Face.
- Policy: The UK opened a consultation on requiring employers to consult staff before installing AI monitoring tools, noting one in three UK organizations now tracks employees’ digital activity, up from one in five two years ago.
- M&A: Bending Spoons agreed to buy Airtable for $1.28 billion in cash, about 2.7x its $480 million ARR, against an $11 billion peak valuation in 2021.
- Research: Post-training a model on 363 office workflow tasks containing zero software engineering examples improved SWE-Bench Pro by 5.8 points, which the authors attribute to general goal-directed execution behaviors rather than domain knowledge.
- Infrastructure: Runware unveiled a transportable modular data center with closed-loop cooling and no water use, and Endeavor Optical Networks left stealth with $10.75 million to move intercontinental data onto satellite lasers instead of subsea fiber.
- Community: The day’s biggest AI thread was an argument that AI-generated header images are a reverse quality signal, at 531 points and 308 comments, well ahead of any product launch.
What This Means
The gap that mattered today was between spending and knowing. Anthropic committed $10 billion to a supplier with no track record, Amazon raised capex by $20 billion and said it still would not be enough, and the FT showed the credit structures being improvised to keep that flowing. On the other side, three separate papers said we are measuring the output badly: benchmarks credit answers reached without reasoning, abuse detection watches one session while attackers work across many, and memory defenses check records one at a time while attacks arrive in pieces. Every one of those is a measurement failure, not a capability failure.
Cloudflare is the useful counterweight because it did the boring thing and measured. It found a 3,400x cost difference between two models on one task and picked the cheap one. That is what the discipline looks like when someone bothers, and it is worth noticing that almost nobody publishes numbers like that.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
