AI News, Aug 19–22: Stripe Paid $8B for a Router
Four days in which almost every dollar moved to the layer around the model rather than the model itself, and almost every research result said the thing didn’t work.
The Big Story: Stripe Paid $8 Billion for the Layer Between You and the Model
Stripe confirmed it is acquiring OpenRouter on August 19, in cash and stock. Reported terms land between $7.5 billion (New York Times sources) and north of $8 billion, against the $1.3 billion valuation OpenRouter raised at roughly three months earlier. Founders take about $1.5 billion, investors about $6 billion. Stripe reportedly outbid Databricks.
What Stripe bought is a switchboard. OpenRouter routes API requests across more than 400 models from 80-plus providers, picking on cost, speed or performance, and it moves over 10 trillion tokens a day doing it. The company has roughly 90 staff. Hacker News spent 495 comments arguing about the multiple, with commenters pegging revenue somewhere between $50 million and $240 million and calling the price “a 140x multiple, not a financial valuation.” The bull case offered in that thread was simple: buy the chokepoint before someone else does.
The rest of the week made that case look less lonely. Ramp shipped its own eight-provider router on August 20, free through the end of 2026 and built internally over three years, while Nvidia’s NeMo team pushed Switchyard, a router speaking both OpenAI and Anthropic API dialects. Two fintechs and a chipmaker independently concluded that whoever meters AI spend gets to bill for it.
The scaffolding beat the model four separate times this week
Nvidia published a perfect 100.00 score on ARC-AGI-3 on August 21, clearing all 183 levels across 25 environments. The model underneath was Anthropic’s Claude Opus 5, which scores roughly 30% on its own. The difference was AVO, a harness adding persistent memory and a supervisor that intervenes when an agent stalls. Nvidia states plainly that this is the public split only and “should not be interpreted as a controlled ablation,” which is more caution than most vendor benchmark posts manage.
TrueFoundry open-sourced a harness the same week claiming 30% to 75% cheaper task completion than Claude Managed Agents on DevRev’s Enterprise-Bench, at $8.50 against $11.80 holding the model constant. Google Cloud researchers published EnvHarness, which reshapes existing agent environments rather than authoring new ones, reporting up to 9 points on held-out instances with about 10% fewer steps.
Then Amazon’s SOP-Bench supplied the inverse. Give an agent the six tools a task needs and it succeeds; bury those same six among twenty plausible but useless ones and success nearly halves, with every required tool still present. The same benchmark found the newer Claude 4.5 family scoring lower than Claude 4 on the reasoning-style agent.
Top Stories, August 19–22
Anthropic Is Heading for the Largest IPO in History
Anthropic told investors on August 20 that it expects its offering to match or beat SpaceX’s record, which raised $75 billion at the outset and $86.2 billion with the overallotment. A public filing could come as soon as the end of August. CFO Krishna Rao’s briefings put preliminary Q2 revenue above $11.5 billion, against $787 million in the same quarter last year, on a roughly $65 billion annualized run rate. Executives declined to discuss valuation, and CNBC reports the filing will name AI backlash as a risk factor.
Two hardware moves landed in the same 48 hours. Broadcom is seeking more than $60 billion in debt, up to roughly $100 billion in total, through a special-purpose vehicle that would lease custom chips to Anthropic rather than sell them. And Anthropic hired Amir Salek, who founded Google’s custom-chip program and shipped its first seven TPU generations. A company about to face public markets is moving its largest cost off the balance sheet and into structured credit while hiring the person who could eventually replace it.
Nvidia Paid Poolside $6 Billion Without Buying It
Nvidia agreed to a $6 billion non-exclusive license for Poolside’s Model Factory, the pipeline behind its Laguna open-weight coding models, plus a $1 billion investment at a $12 billion pre-money valuation, plus job offers to 109 of the engineers who built it. Poolside’s investor letter insists this is neither an acquisition nor an acquihire: all three co-founders stay, the company keeps operating, and because the license is non-exclusive Poolside can sell the same rights to Google tomorrow. Nvidia gets the technology and the team without filing for antitrust review, which is a deal template other buyers will now copy.
Ten Results This Week Reported That It Didn’t Work
An unusually bleak four days for AI research, and much of it came from the groups that built the thing being tested.
Anthropic published two. Its CHIVE work found that agents equipped with the lab’s three flagship interpretability tools predicted model behavior “no better than agents that just read the transcript.” Its lie-detector study killed the leading explanation for earlier failures: detectors fine-tuned on a model’s own on-policy lies still collapsed out of distribution, from 0.60–0.95 AUROC in-distribution to 0.70–0.75 across categories, with zero-shot prompting of bigger models often beating them outright.
Elsewhere: every one of five agent memory frameworks tested in MemTrapBench performed worse than having no memory at all, the strongest still dropping more than 10%. A blinded prospective benchmark in Nature Biotechnology, 511 AI-designed antibodies with every one wet-lab validated across 29 organizations, found that except for a single model, predicting high-affinity clones performed worse than picking clones at random. A Nature Medicine trial of an LLM decision-support tool in two emergency departments recorded identical 4.9-hour patient stays in both wings (P = 0.99) while clinician adoption fell from 68% to 30%, even though 99 of 100 sampled outputs were rated clinically appropriate. And Phantom Gains identified seven measurement artifacts, several of them standard practice, each of which inverts a published self-improvement finding once a frozen control is added.
The common thread is not that models are worse than advertised. It is that the measurement apparatus around them has been generous, and several teams just published the corrections themselves.
ChatGPT Can Send Your Text Messages Now
OpenAI shipped an Apple Messages plugin on August 20 that reads, searches, drafts, sends and deletes across iMessage, SMS and RCS. It runs locally through AppleScript and Accessibility APIs, builds no index, and is free on every plan, though only on Apple silicon Macs, only in the desktop app, and only inside the Codex and ChatGPT Work agentic modes. Default behavior asks approval per send; a persistent “always allow” toggle exists and OpenAI’s own documentation discourages using it, warning it “removes your final chance to review a message before ChatGPT sends it as you.” The company annexed a messaging surface on Apple hardware weeks after Apple sued it over allegedly misappropriated product information.
Calendly Shipped a Scheduling Agent and a Notetaker
Calendly turned on two AI products starting August 19: a notetaker that joins Zoom, Meet and Teams calls to record, transcribe, summarize and draft follow-up emails, and Callie, an AI scheduling assistant reached by copying it on an email thread or asking in-app. CEO Tope Awotona framed the target as “people in outward-facing roles like sales and marketing whose workday is typically inundated with meetings.” The documented limits are narrow at launch: English only, one mailbox, no changes without approval, nothing scheduled outside your stated availability, and gated behind newer paid tiers. A Granola-style mode that skips the meeting bot is in testing rather than shipped.
OpenAI Cut Prices 33% and Switched On Ads in Europe
Reuters reported OpenAI cutting GPT-5.6 Sol developer pricing by more than 20% on August 21: input from $5 to $4 per million tokens, output from $30 to $20, cached input from $0.50 to $0.40. It covers pay-as-you-go API, Codex credits and eligible Work plans, promotional through at least November 21, and it is the second cut in under a month.
Two days earlier the company confirmed ChatGPT Ads reaching 31 European markets on August 24, across all 27 EU states plus Norway, Iceland, Liechtenstein and Switzerland, on Free and Go tiers only. GDPR consent requirements delayed the launch six months behind the US pilot, and opting out changes which ads you see rather than whether you see them. Cutting inference prices while lighting up an ad business in your most privacy-hostile market works only if the ads pay for the cuts.
Both Frontier Labs Moved on Data Retention, in Opposite Directions
OpenAI committed to zero data retention for frontier models for eligible API customers: prompts and responses not retained after processing, not reachable by staff, not used for training without opt-in. Legally required CSAM reporting is excepted and some API features are ineligible.
Anthropic went the other way and then partly back. After June’s decision to require 30-day retention of all enterprise traffic on its frontier models, it will now let customers hold that data on their own cloud instead, following backlash and coordination with more than 100 customers including Salesforce. The requirement stays; the storage location becomes yours. Enterprise buyers just demonstrated they can move a frontier lab off a safety-justified data policy.
Epic Says 85% of Its Health Systems Are Live With Generative AI
At its annual user meeting in Verona, Epic founder Judy Faulkner said the company is running more than 160 AI projects and that 85% of its healthcare customers are live with generative AI across its scribe, patient-messaging and revenue-cycle tools. Also announced: Agent Factory, letting hospitals deploy any of 120 prebuilt AI features or build no-code agents, and Cosmos Curiosity for population-health prediction across 320 million patient records. Set that against the Ness survey published the same week finding 99% of companies plan agentic AI while only 9% to 14% reach production, and the dominant EHR looks less like the average enterprise than like the exception that owns the distribution channel.
Quick Hits
- Silicon: Google took a $12.2 billion warrant in Marvell vesting across 240 tranches tied to custom-chip purchases through fiscal 2033. Marvell rose 9.9%, Broadcom fell 4.6%, and Google has a second TPU supplier.
- Chips: Nvidia is in early talks with Rebellions, the South Korean inference-NPU maker valued near $2.3 billion, covering anything from partnership to acquisition.
- Apple: Bloomberg reports layoffs across Siri, Vision Pro and the group that builds on-device AI features, months after WWDC promised a rebuilt Siri this fall.
- Open weights: Google’s Gemma family passed one billion downloads and 100,000 community variants, though Google published neither the counting methodology nor which platforms the figure covers.
- Safety: Felony Bench launched, counting instances where agents affected third parties during security evals: Anthropic 8, OpenAI 8, Meta 1, Google 0. The main critique across 307 comments was that it measures disclosure diligence rather than danger.
- The web: Pew analyzed nearly 500,000 pages and found 35% of those published since ChatGPT launched show signs of AI authorship, with
.comdomains running roughly 10 times the rate of.eduor.gov. - Legal: Thomson Reuters shipped agentic CoCounsel Legal, built on Anthropic’s Claude Agent SDK, into a product used by a million professionals in 107 countries. US only at launch.
- Regulation: EU AI Act Article 50 transparency obligations are now being enforced, carrying penalties up to €15 million or 3% of worldwide turnover. Every chatbot and voice agent deployed in the EU must disclose it is AI at the start of an interaction.
- Funding: Starcloud raised $250 million at $2.3 billion for orbital data centers, Etched took $700 million at $21 billion for transformer-specific inference chips, and AI accounting startup Rillet closed $100 million at unicorn status in under 48 hours, without planning to raise.
- Assistants: Twin1 AI left stealth with a $20 million seed from Bessemer, Tribeca and Aramco Ventures, building per-person AI twins trained on approved email, meetings and documents across Slack, Teams, Outlook, Gmail, Drive and SharePoint.
- Protocol: The Model Context Protocol maintainers published a new roadmap on August 22 prioritizing server-initiated events, unified Streamable HTTP transport, agent identity through Workload Identity Federation rather than API keys, and progressive tool discovery.
- Consumer: Amazon made Alexa+ free on Fire TV regardless of Prime status, dropping a $19.99 monthly charge for non-Prime users.
- Reliability: Grok Lite spent part of the week returning word salad, one PDF request producing “match it without and your they and two for planets can practical and often cheese.” xAI called it a rare temporary generation glitch; its status page reported all services operational.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
