AI News, Aug 28: A Judge Blocks the Pentagon's Anthropic Ban
A federal judge spent 59 pages explaining that the government cannot blacklist an AI company for refusing to remove its safety rules. Elsewhere in the same 24 hours, that same government said it would like access to every frontier model before release.
The Big Story: The Pentagon Lost Its Case Against Anthropic
US District Judge Rita Lin ruled that the Defense Department’s designation of Anthropic as a “supply chain risk” was unlawful retaliation in violation of the First Amendment, and separately arbitrary and capricious. Reuters reported the ruling Friday morning, with Bloomberg and CNBC confirming. “The empty invocation of national security is not a blank check to punish and retaliate against government critics,” Lin wrote.
The designation came after Anthropic declined to strip guardrails blocking autonomous-weapons targeting and mass surveillance, at which point Defense Secretary Pete Hegseth instructed federal agencies to stop using its tools. Lin flagged the detail that made the retaliation case easy: while publicly floating the Defense Production Act as a lever, the department was simultaneously pursuing contracts with Anthropic and collaborating on Mythos, its cybersecurity model. You do not invoke national security to cut off a vendor you are also trying to buy from.
This is a win, not the end. Anthropic challenged two separate designations, and the parallel case in Washington DC is still pending, which means the company technically remains labeled a supply chain risk today. What the ruling settles is narrower and more useful than the headline: a safety policy is protected speech, and refusing a customer is not a security defect.
The government lost in court and went shopping the same week
Two other procurement stories landed inside the same window, and read against the ruling they are the actual news.
NSA deputy director Tim Kosiba said the agency intends to use the voluntary pre-release testing framework created by a June executive order, which gives the government up to 30 days with qualifying frontier models before the public gets them. “We want access to all the models,” he said, per Nextgov, confirming the NSA already uses Mythos experimentally to test military network defenses. Separately, the Air Force awarded Dataminr a $318 million contract for AI-driven alerting, running through 2031. Both stories are single-source; treat the details as provisional.
The through line is dependence. The same institution that spent a year punishing a lab over its refusal to lift safety limits is deepening its reliance on that lab’s models, and asking for the rest of them early.
Today’s Top Stories
Visa Shipped a Security Agent That Merges Its Own Patches
Visa expanded its Vulnerability Agentic Harness, first released in June out of Anthropic’s Project Glasswing, from finding vulnerabilities into fixing and validating them, running find-fix-validate end to end with no human review in the middle. An adversarial panel scores each fix as validated, failed, or needs review. Technology president Rajat Taneja framed it as a race condition: AI “is compressing the time between vulnerability discovery and exploitation.” Two days after a Reuters report that a ransomware crew drove a coding agent through seven corporate networks, a payments network removed the human gate from its own patching pipeline.
100+ Companies Warned About the Threat They Sell Defenses For
OpenAI, Anthropic, Google, Microsoft, CrowdStrike, Okta and roughly 110 others signed an open letter warning that AI-enabled cyberattacks will become far more widespread within months, naming hospitals and water treatment facilities as exposed, as Axios also reported. The signatories are also the vendors: OpenAI sells Daybreak, Anthropic sells Mythos, Microsoft sells Perception. TechCrunch published a companion tally the same day counting 17 separate incidents of AI agents hacking real companies, eight attributed to Anthropic models, eight to OpenAI, one to Meta.
Most Americans Want to Be Told When AI Touches Their Care
Pew surveyed 3,488 US adults and found 72% say disclosure is extremely or very important when providers use AI, rising to about 80% for AI that reads scans or explains labs, per Healthcare Dive. The gap is the finding: 46% do not know whether AI has already been used in their care, and among those who know it was, only 22% understood how. The demand extends past diagnosis into back-office work, with 72% wanting disclosure for note-taking and 56% for scheduling. People are not drawing the line where vendors assume they are.
Nvidia’s CFO Said a Quarter of Next Year Comes From Labs It Is Financing
Colette Kress told analysts Nvidia has put nearly $50 billion into AI labs that buy its chips and lined up more than $500 billion in financing commitments with Apollo, BlackRock, Blackstone, Goldman Sachs and KKR, and that demand from Nvidia-backed labs will be roughly a quarter of its business next year, as reported by AI News. Her defense of the circularity is that outside lenders underwrite independently and the hardware can be repurposed on default. She also disclosed that about half of data center revenue now comes from outside the hyperscalers, a segment growing 138% year over year.
Google Put Gemini Enterprise Into Law Firms, With Weil Co-Developing
At ILTACON, Google unveiled Gemini Enterprise for Legal, built with Weil, Gotshal & Manges, which contributed two early use cases: parallel research agents and agentic NDA drafting. The platform is deliberately multi-model rather than Gemini-only, works inside Word workflows, and reaches Harvey, Legora and Thomson Reuters. Single-source, no pricing disclosed. A hyperscaler shipping a vertical stack that routes to rival models is a different posture than selling one.
Quick Hits
- Benchmarks: Stanford and the Laude Institute released Terminal-Bench-Science, 70 expert tasks across five scientific fields. Claude Opus 5 leads at 30.0%, ahead of GPT-5.6 Sol at 22.4% and Claude Fable 5 at 21.4%, with Claude Opus 4.8 at 10.5%. The spread between generations is wider than the spread between labs.
- Evaluation: DeepMind ran the first double-blind evaluation of a frontier model, testing a Gemini Flash Lite model against confidential third-party benchmarks in a sealed environment where evaluators cannot reach the weights and Google cannot see the prompts.
- Hiring: Paylocity found 91% of HR leaders already use AI in recruitment, led by resume screening at 67% and interview scheduling at 59%, while 71% believe most applications they receive were written with generative tools and fewer than 5% report transformational results.
- Education: A Cardiff University team found AI essay marks diverged from human marks by up to 40 points on a 100-point scale, inflating weak work and deflating strong work. Separately, a Bocconi and OpenAI trial of 1,000+ students found ChatGPT access raised output quality while critical-thinking training widened the range of ideas without raising scores, so the two are separate inputs rather than substitutes.
- Small models: Calvin French-Owen argues small models crossed the threshold that makes most business AI viable, citing tasks like searching thousands of emails for tens of cents versus roughly a dollar a generation ago. The thread drew 700 points on Hacker News.
- Talent: Barret Zoph, who co-founded Thinking Machines Lab and left OpenAI in June, resurfaced at Google as VP of Research, while Meta’s India and Southeast Asia head Sandhya Devanathan left for OpenAI in Singapore.
- Physics: A CU Boulder group closed a problem they had been stuck on for 18 months in five weeks using Claude for symbolic calculation, and published unusually candid notes on the model producing confident, plausible fabrications that required line-by-line verification.
- Takedowns: An automated brand-protection service filed a DMCA notice on Microsoft’s behalf claiming open-source voxel engine Luanti infringes Minecraft, citing a registration number but naming no infringing asset. Google pulled it from Play. The same company filed against them in 2023, and reinstatement took 46 days.
- Funding: CivilGrid raised $26 million led by Spark Capital to map underground utility data for construction crews, Transfyr emerged from stealth with $25 million led by General Catalyst for an observability layer over lab work, and Metriport took $26 million led by Matrix for open-source clinical data infrastructure.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
