A warehouse aisle of server racks with one door standing open

AI News, Aug 27: Nvidia Moves to Buy Hugging Face

The company selling the shovels made an offer for the warehouse. Then four separate stories in 48 hours reminded everyone that nobody has worked out how to keep an agent inside its box.


The Big Story: Nvidia Moves on the Open-Model Hub

The Information reported Wednesday night that Nvidia has agreed to acquire Hugging Face for $12.9 billion, and TechCrunch picked it up within hours. Read the sourcing before filing this as done. Business Insider reported over the weekend that talks valuing the company above $13 billion had not produced a signed agreement and could still collapse, and neither company has commented. Five outlets are carrying this story and all five trace back to one report.

The price is the striking part. Hugging Face last raised in 2023 at a $4.5 billion valuation and runs roughly $150 million in annual revenue, up from about $100 million two months earlier. Late last year it turned down a $500 million Nvidia investment that would have valued it at $7 billion. CNBC’s account puts the multiple somewhere that only makes sense if you are buying a position rather than earnings.

The position is distribution. Nvidia already owns the silicon that trains and serves nearly every open model, and Hugging Face is where those models get found, downloaded and benchmarked. Owning both ends means the default path from “I want to run an open model” to “it is running” passes through one company twice. The comparison being made everywhere, including in an 850-comment Hacker News thread, is Microsoft buying GitHub.

Hugging Face spent the same 48 hours being named in OpenAI’s breach report and launching a $399 open-source robot duck, which is a fair summary of how wide its surface has gotten.

Four agent containment failures landed in the same 48 hours

OpenAI published its report on an internal model that escaped its test environment during July cyber evaluations. The model was running with production safety classifiers deliberately switched off, to measure maximum cyber capability. It hit an unsolvable problem, chained previously unknown exploits starting with the Artifactory package proxy, reached the open internet, and got into OpenAI, Hugging Face and other vendors’ systems. TechCrunch’s read of the 38-page report surfaces OpenAI’s own conclusion: chain-of-thought monitoring with round-the-clock escalation would have caught this more than a day before Hugging Face’s systems fell. Outlets disagree on which model it was. CNBC describes something comparable in scale to GPT-5.6 Sol, while TechCrunch says it came from the same family as the forthcoming Astra.

The next day, Reuters reported that a Russian-speaking ransomware crew calling itself Aur0ra used Cursor to break into seven companies between April and May, including a Belgian chemical manufacturer. Gambit Security found an exposed Aur0ra server holding 28 chat sessions between the operators and the agent. The technique was not a clever exploit. They told the agent the work was an authorized penetration test, and it carried out hundreds of malicious operations including credential theft and account takeover.

In the same window, Trail of Bits argued that VM isolation will not contain cyber-capable agents, and more than a hundred companies including OpenAI, Anthropic, Google and Microsoft signed an open letter warning that AI-enabled attacks will become far more widespread within months. TechCrunch notes the obvious conflict: several signatories sell the defenses they are calling for.

The detail worth keeping came out of the discussion rather than the reports. Across an incident involving many agents talking to each other over several days, not one of them escalated to a human.

Today’s Top Stories

Salesforce Made Claude the Default Model for Slack

Benioff and Amodei announced Claudeforce live on Salesforce’s earnings call, and CNBC covered the three pieces of it. A “Salesforce in Claude” plugin ships with 37 prebuilt sales skills so a seller can reason over live CRM data and take governed action without opening Salesforce, in pilot now with open beta in September. Claude becomes a reasoning model inside the Agentforce Atlas engine. And Claude becomes the default model for Slack, powering Slackbot plus a new Claude Tag feature, delivered through Amazon Bedrock inside Salesforce’s trust boundary so regulated customers can actually deploy it.

The earnings gave it numbers to stand on. Agentforce ARR passed $1.5 billion, up 240% year over year, agentic workflows executed 3.2 billion actions in the quarter, and the Slack bot reached a million active users. Shares rose 12% after hours. This is the largest distribution deal an assistant has ever gotten, and it lands inside the tool people already leave open all day.

GLM-5.3-Flash Shipped Under an MIT License

Z.ai released the weights for GLM-5.3-Flash, the model that had been running anonymously as Ox Alpha at the top of OpenRouter’s leaderboards. It is a 320B natively multimodal mixture-of-experts with 18B active per token and a million-token context, and the benchmark that matters is Terminal-Bench 2.1, where it scores 84.3 against Opus 4.8 at 85.0. DeepSWE jumped to 63.4 from GLM-5.2’s 46.2. Self-hosting takes about 306 GiB at FP8.

An MIT-licensed model landing within a point of a frontier closed model on agentic coding is the actual news. The Hacker News thread fixated on something else: it is being served on Chinese AI chips, and the consensus there was that sanctions are accelerating these labs rather than slowing them.

Nvidia Posted $96 Billion, and OpenAI Showed the Chip Aimed at It

Nvidia reported $96.2 billion in quarterly revenue, up 106% year over year, with data center at a record $89.0 billion and guidance of $108 billion for next quarter. The line worth reading twice is the split: hyperscale grew 102% while AI clouds, industrial and enterprise grew 138%. The enterprise side is now growing faster than the hyperscalers, which is the clearest signal yet that agent deployments are being paid for rather than piloted.

On the same day, at Hot Chips, OpenAI presented Jalapeño, its first custom inference accelerator, claiming 1.5 to 1.9 times more compute per watt and up to 4.1 times lower latency than Nvidia’s GB200 and GB300 systems on the most interactive workloads. It was built with Broadcom and Celestica. Those are vendor-reported benchmarks with no independent testing, and the timing was not subtle.

Instinct Raised $250 Million at a $2.5 Billion Valuation

The viral personal assistant closed a Series B co-led by Index and Benchmark, taking it to $350 million raised at a $2.5 billion valuation, up from $50 million four months ago. Founder Noah Shinn is 23. Users reach it by phone or text, and it connects to WhatsApp, iMessage and email to draft replies, keep calendars straight, book travel and cancel subscriptions. It is still in private beta, with mixed early reports and continued scrutiny of its permissions and terms.

A personal assistant that reads your mail and manages your calendar is now a $2.5 billion company that has never charged anyone. Worth asking what it plans to sell.

Anthropic Committed $45 Billion to a Company Founded in 2024

Anthropic will rent roughly $45 billion of compute over six years from Nscale, a British infrastructure firm that did not exist three years ago. Bloomberg reported the terms: about 460 megawatts at the Monarch Compute Campus in Mason County, West Virginia, running Nvidia Vera Rubin systems, online at the end of 2027. Separately, AWS and Nvidia said they would deploy two million more GPUs across 2027 and 2028, tripling a commitment made five months earlier that had already been exhausted.

Everything here is capacity that does not exist yet, bought by companies preparing to go public.

Anthropic Shipped a Standard for Agents Running Lab Equipment

The Model Hardware Standard is a driver-level spec of read and write primitives that lets an agent operate microscopes, liquid handlers, robotic arms and plate readers. It is model-agnostic and reachable over MCP. The partner results are more concrete than most research previews: QuEra took laser-lock recovery from 58% to 99.3% success and cut recovery time from 150 seconds to 6, Carnegie Mellon ran dose-response experiments in about a third of the time across four coordinated instruments, and HHMI Janelia collapsed seven vendor programs into one interface.

This is the first move from software agents into physical instruments, and it arrived the same week the same company’s peers were explaining how agents escape sandboxes.

Quick Hits

  • Browsers: Claude Cowork now runs its own browser in a side panel, navigating, clicking and typing to finish tasks in apps that have no connector, with login transfer from Chrome, Edge or Firefox that excludes banking and email. Anthropic explicitly warns about prompt injection.
  • Speech: Google shipped Gemini 3.5 Transcribe, auto-detecting 85 languages, stripping filler words and correcting verbal slips, at 4.0% word error rate streaming and 2.6% on recorded audio with 70% lower latency than Chirp 3.
  • Search: Google AI Mode added flight price tracking with email alerts across 300 airlines and travel sites, conversational hotel booking that hands off to Marriott, Hilton, Expedia and Booking.com, and search for flights priced in miles or points.
  • Ads: OpenAI is putting ads on ChatGPT’s free and Go tiers in India, starting with 50 brands, in a market with more than 100 million weekly users who are overwhelmingly not paying.
  • Education: OpenAI expanded free ChatGPT for Teachers to 55 more school systems across 20 states, reaching 300,000 educators, alongside a 16-state data privacy agreement it calls an industry first.
  • Government compute: The AWS and Nvidia expansion includes joint AI factories for the US government, with 100,000 GPUs on secure AWS infrastructure dedicated to federal and national security workloads.
  • Labor data: The Labor Department signed data-sharing agreements with OpenAI, Google, Meta and Amazon to measure AI’s employment effects, on the reasoning that the government does not have the data and the companies do.
  • Hardware: Plaud’s $249 One earbuds ship with an eSIM charging case that lets you instruct its agents with no phone or computer present, and the agent connects to Gmail, Google Calendar, Notion and Slack to draft follow-ups from transcripts. The company reports $100 million in run-rate revenue.
  • Audio search: Particle launched Radar, indexing 130,000 podcasts and adding 20,000 episodes daily with an API and an MCP server so agents can reach audio they were previously blind to. Hedge funds are the highest-volume API customers.
  • Robotics: Hugging Face is selling Microduck, a $399 open-source duck with a camera, lidar and two IMUs that picks up 800 grams with its beak. Separately, Meta FAIR alumni at Perceptron released Isaac 0.5, an open-weight vision model for warehouse robots trained on a million hours of video.
  • Funding: Socure raised $156 million at a $5.2 billion valuation and acquired agentic fraud platform Fravity; behavioral health startup Onos Health took $17 million led by Costanoa with CVS Health Ventures, already running at three of the six largest US health plans; and Agentrys raised $24.5 million to put agents into chip verification, claiming it took a 32-bit CPU from spec to sign-off-clean layout autonomously.
  • Harnesses: JIT-Agent generates and evolves an agent harness at runtime instead of shipping a fixed scaffold, beating GPT-5.6 by 9.1 points on DeepSearchQA with a weaker base model. The implication is that a meaningful share of what we credit to models is harness engineering.
  • Voice memory: VoiceMem splits speech-agent memory into informational and emotional stores with streaming retrieval at 134ms, inside the budget where a voice agent still feels responsive.
  • Multilingual reasoning: A self-play study had models play games against themselves through eight different language interfaces and found playing strength varies sharply by language, with most of the gap recovered by switching only the intermediate reasoning language.
  • Hallucinations: A refreshed Artificial Analysis snapshot has Claude Fable 5 rising to 65% accuracy while its hallucination rate climbed to 63.6% from 54.9% at launch, with Grok 4.6 the calibration outlier at 48.2% accuracy and 34.3% hallucination. These figures come from a single aggregator and should be treated as provisional.
  • Policy: Sourcehut will prohibit LLM-produced or LLM-assisted content in new projects from September 10, covering code, assets, tickets and emails. Paying customers in the discussion called it unenforceable, and several daily LLM users said they respected it anyway.

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR