A padlocked gate in front of a row of server racks

AI News, Oct 7: Europe Ships a Trillion Parameters

Mistral put out a trillion-parameter model and said the weights arrive at the end of the month. Separately, a group of retailers and payment firms started writing the rules for which AI agents get to use their websites, and the two biggest labs are not in the room.


The Big Story: Mistral Shipped a Trillion Parameters and Hedged Every Claim But One

Mistral launched a public preview of Mistral Large 4, “unofficially ML4, very officially: le Chonk.” It is a one-trillion-parameter natively multimodal model with 49 billion active parameters, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European data centers. The weights come “by the end of the month,” which makes today an announcement rather than a release.

Read the launch page closely and almost every superlative is fenced to its own category: state of the art on SciCode-Verified “among open-weight models,” cybersecurity that “ranks among the top five models globally and leads open-weight models developed outside China.” Exactly one claim arrives unqualified: on a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, “the highest of any model.” Mistral explains why in the very next sentence. Claude Opus 5.5 and GPT-6 Astra “score near zero on the same test because they refuse to perform the task.”

That refusal gap is the product, and Mistral is explicit about where it leads: until the weights drop it is red-teaming “with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.” Anthropic announced something structurally identical the same day, splitting its Cyber Verification Program into three tiers with fewer blocks on security work at each step up, the top tier reserved for organizations cleared to test power grids and telecom networks. Two labs on two continents decided on the same day that the way to compete in security is to turn the guardrails down for customers who pass a background check.

The Web Spent the Day Building Locks for Agents, and OpenAI Wasn’t in the Room

Meta, Walmart, Stripe and Sierra launched the Personal Agent Protocol, a standard for authenticating a customer’s personal agent to a business so a site can tell it apart from a scraper. It is led by Sierra co-founder Bret Taylor, who is also OpenAI’s chairman, and neither OpenAI nor Anthropic is participating; Taylor says he expects OpenAI to back it eventually.

They are standardizing because the blocking already started. TechCrunch documented how websites are shutting agents out, with Amazon cutting Meta’s Muse off from its retail catalog outright and users reporting the same from Yelp, eBay, Zillow, Pizza Hut and Adidas. Insurers are pricing the other end of it: Aon has reviewed more than 300 AI-related legal cases, and Verisk’s head of underwriting says the buck stops with the CEO, with directors-and-officers policies potentially covering suits against Sam Altman personally.

The funding followed the same logic. Ampersand raised a $15 million Series A led by Bessemer to let agentic applications read and write inside legacy systems of record, and Monid raised $7.7 million for one prepaid account where an agent can find, test and pay for third-party APIs by the call. Nobody funded a smarter model this window. They funded the doors.

Today’s Top Stories

OpenAI Published 722 Machine-Written Proofs, Citing a Committee That Asked It to Stop

OpenAI released 722 manuscripts organized into 372 families produced by what it calls an internal frontier model, the average result using “the equivalent compute of roughly three hours of ChatGPT Pro thinking.” The repository is candid about its state: “Not all have accompanying Lean formalizations,” and “Some of the unformalized results could have issues.”

The governance story is sharper than the math. OpenAI’s announcement says it has been consulting the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and has “drawn on their advice and public recommendations to inform how we release these results.” Those recommendations open by saying the opposite: “we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.” They also specify repositories that “should not be controlled by any AI lab.” OpenAI published to one it owns.

The field answered within hours. Thomas Bloom, who runs erdosproblems.com, froze problem comments and removed every problem status on a site hosting 1,221 problems, citing “a wave of AI-produced solutions provided with no explanation” which are “displacing and discouraging those who are actually interested in the mathematics.” The sharpest objection to the 722 was never that a particular result is wrong. It is that nobody can plausibly have checked them.

Lambda Is Raising $4 Billion, and One Customer Is $35 Billion of Its Backlog

The GPU cloud is raising up to $4 billion at a $14.5 billion pre-money valuation, co-led by Coatue and Blackstone, in what is likely its last private round before a planned 2027 IPO. An investor letter shows its backlog growing from $15 billion in June to $50 billion in September, and $35 billion of that is one commitment from Anthropic signed in late August. An IPO prospectus resting on a single lab’s purchasing plans is a new kind of risk disclosure.

Finland Told Google to Stop Digging

The Finnish Licensing and Supervision Agency ordered Google’s local subsidiary to suspend building work at Muhos and Kajaani until environmental impact assessments are complete, with a halt deadline of October 23. It lands a month after Google committed $15 billion to Finnish AI infrastructure. The company’s response was unusually contrite: “We understand the concerns and have fallen short of our own high standards in this instance.” Permitting, not chips or power, is the constraint a regulator could reach this week.

Three Decision Models Landed in a Single Day

OpenAI put its Decisions API into public beta: one model (gpt-6-luna) returning a typed answer instead of generated text, meaning a probability, a choice from a fixed set or a score against a rubric, “about 10x faster than the Responses API” at $0.10 per million input tokens with no output-token charge. In the same window Musubi released PolicyLM-1.7B with open weights, built to apply a plain-English content policy in under 50 milliseconds, and Strands published an open-weight Decider 2B that swaps the text-generation head for a pointer head. A category nobody had a name for in the spring now has three entrants and a price war.

Quick Hits

  • Open weights: The window’s only frontier-lab release with downloadable weights was EmbeddingGemma 2, a 740-million-parameter multimodal embedding model under Apache 2.0 that loads just 270M parameters for text-only work.
  • Education: Maryland’s own randomized trial of a ChatGPT-based course tutor across 2,379 undergraduates found that only about 15% of students with access used it, and that access “marginally reduced final grades and LMS participation”. The provost is holding off on campuswide deployment.
  • Productivity: Anthropic put Claude inside Google Docs, Sheets and Slides in public beta on all paid plans, defaulting to an “Ask before edits” mode.
  • Agent integrations: Carly connects to thousands of apps and starts its workflows from events like a new email or a booked meeting, rather than waiting to be asked.
  • Banking: At least seven South Korean banks and lenders were breached through loan-recruiter portals since late September, with evidence pointing to an open-source pen-testing tool driven by multiple AI agents, though the sector’s cyber agency stressed the AI “did not act independently without human involvement.”
  • Layoffs: HubSpot cut about 660 jobs, 7% of staff, with CEO Yamini Rangan describing a shift to “delivering outcomes for them with AI” while noting the decision “is not driven by AI-related efficiencies.”
  • Energy: The Department of Energy made a $4.2 billion conditional loan commitment to Vistra covering 433 MW of nuclear uprates, tracing back to Meta’s 2.6 GW power purchase agreement from January.
  • Health AI: Utah approved three more AI health pilots, including one that renews prescriptions for chronic conditions, though each needs further written approval before reducing human oversight.
  • Funding: North American venture funding fell 35% quarter over quarter to $92 billion while rising 50% year over year, with roughly two-thirds going to AI companies. Crunchbase blames the absence of new OpenAI and Anthropic megarounds, not a weakening climate.
  • Economics: Microsoft published a Nobel laureate’s bearish AI forecast on its own blog, in which Daron Acemoglu argues AI will add about 1.5% to GDP over a decade and replace at most 5% of jobs, because “what’s missing are easy-to-deploy apps that change how things get made.”
  • Office software: LibreOffice is marketing “no AI” as a shipped feature, arguing that for anyone handling privileged data “the only assurance that survives an audit is that it does not leave the machine.”

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR