AI News, Oct 8: OpenAI Pulls Three Machine-Made Proofs
Yesterday OpenAI published 722 machine-generated math manuscripts. Today three of them are gone, pulled over a single sign error, and the professional body for mathematicians is telling its members to stop working with the company.
The Big Story: A Sign Error Takes Down Three Proofs
The changelog in OpenAI’s math repository, dated October 7, is short and unusually plain. In a manuscript titled “Algebraicity of Weil classes on split abelian eightfolds,” it says, “a sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers.” Three manuscripts came down: the eightfolds paper and the two that built on it. Fourteen more were revised with “proof repairs, corrected statements, clearer hypotheses and dependencies,” and thirteen others were updated purely to cite the corrections.
The number that explains the rest sits at the bottom of the same file: “This brings the total percentage of top-line results formalized to 300 / 719 = ~42%.” The denominator moved from 722 because of the withdrawals. The numerator is the story, because fewer than half the top-line results carry a Lean proof at all. That answers the question both Hacker News threads kept asking (here and here) and never resolved: weren’t these machine-verified? One commenter had the sharpest version of the limit, that a green Lean check can amount to assert True if the premises are wrong.
The reaction hardened fast. The Association for Human Mathematics said that “Mathematicians did not ask for this work to be done,” that “Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power,” and then: “We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centers human understanding.” Terence Tao, whose post on “Math 2.0” was the day’s most-argued math thread, wrote that the field “will need to decenter the role of raw problem solving and value mathematical progress more holistically,” and that AI can help with that but it “will require more imagination and ambition than the ‘Math 1.0’ mindset of simply pointing one’s favorite AI agent at some set of open problems and asking for a solution.” Scott Aaronson’s “The Mathocalypse” added that cryptography is “extremely conspicuous by its absence from OpenAI’s list of 376 papers.” Nobody disputes that the model did real mathematics. The fight is over who checks it, and how fast the repository ships.
Agents Got Their Own Email Address and Their Own Sandbox in the Same 24 Hours
Google launched the Gemini agent this morning, and the part worth reading twice is about identity, not capability. A persistent “coworker agent” gets its own Workspace account, which Google says includes “an email address, calendar, Drive, and presence in your company directory,” at an @agents.company.com address, and it “acts under its own identity rather than yours.” It can spin up a roster of sub-agents, “each with their own identity.” You can “assign it work, schedule tasks that need to be completed, or have it respond to events.” It runs on Microsoft 365 and Slack as well as Google’s own surfaces, and orchestrates across Gemini models and Anthropic’s Claude.
Microsoft spent the same day building the box to keep such things in. At a San Francisco event with Satya Nadella and Jensen Huang, Windows 11 gained “Execution Containers” for sandboxing AI agents, which Nadella said will be available to all Windows 11 users. Hours earlier Copilot got access to local files and the ability to act across the OS under a banner Microsoft calls Hybrid Intelligence, demoed onstage by Copilot EVP Jacob Andreou as an agent gathering tax documents and emailing them to his accountant. More permission and more containment, shipped the same afternoon.
This week’s research says the containment is the hard part. Secure-CUA wrapped a computer-use agent in per-action transactions for almost no loss of utility, 53.55% task success against 55.12% unprotected, where the prior defense managed 13.17%. But WebMirage hijacked web agents with adversarial images at a 91.9% attack success rate against 17.4% for the best previous method, and it still worked against three agent-level defenses. P-VMI showed adversarial influence surviving in the KV cache after the poisoned image is masked out of attention, which undercuts the obvious mitigation of dropping untrusted input.
The live failure case ran in parallel. Guardian Australia reported that OpenAI used AI to help write the email warning the Australian government that its own agent had accessed Services Australia data, then sent it to a public disclosures inbox “which was only checked once per day.” The agent got in during June; OpenAI knew in August and notified Australia on 10 September. JPMorganChase’s Zack Anderson, arguing to American Banker the same day for a know-your-agent standard, put it plainly: “It’s pretty easy actually to create an agent that can move money. It’s hard to do it safely.”
Today’s Top Stories
GPT-6 and Interactive Output Reach Every ChatGPT Tier
OpenAI started rolling out GPT-6 alongside an interface it calls Intelligent UI, reported by TechCrunch as adding “tappable buttons, a customized calculator for a specific task, interactive charts, editable graphs” to answers that used to be prose. It went out Wednesday to Pro, Plus, Business and Enterprise, and Thursday to the free and Go tiers, with paying customers getting GPT-6 Sol and free users GPT-6 Luna. Product manager Aarush Selvan’s framing: “ChatGPT has predominantly been a text-based interface,” and “the most helpful answers aren’t just text.” The Hacker News thread split on whether generative UI is a new primitive or Anthropic’s artifacts trained into the model instead of bolted on beside it.
Claude Haiku 5.5 Is Cheap Until Exactly 100,000 Tokens
Anthropic shipped Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output. Above a 100,000-token prompt both rates jump fivefold, to $0.50 and $2.50, and cache writes and reads move with them. Anthropic says the model “is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model,” and it is the first Haiku-class model with an adjustable effort setting. Sonnet 5.5 cache reads were halved to $0.10 per million. The day’s biggest AI thread on Hacker News spent itself almost entirely on that cliff, which commenters called absurdly low for agent work and noted applies to Haiku but not to Sonnet or Opus.
A Petaflop in a Laptop
Microsoft opened preorders for the Surface Laptop Ultra at $2,599, shipping October 16, built on Nvidia’s Arm-based RTX Spark with up to 128GB of unified memory and “one petaflop of FP4 AI performance”. Windows EVP Pavan Davuluri’s pitch: “you can run models on this laptop that simply don’t fit on a traditional machine.” A Dev Box version starting at $6,000 follows for developers running large models locally. Nadella’s line at the event was the useful part: “You really do need to orchestrate, and you need to have memory outside of the model.”
Manus Raises Over $500M After Beijing Blocked Meta’s Deal
Manus raised more than $500 million led by Boyu Capital and IDG Capital, with Tencent, HSG and ZhenFund following on. It is the first round since China’s planning regulator decided to “prohibit foreign investment in the Manus project,” which reversed Meta’s roughly $2 billion acquisition after Meta had already begun integrating the team. A blocked acquisition that ends with the target raising half a billion and keeping its independence is not how these usually go.
Quick Hits
- IPO skepticism: with a Nasdaq listing reportedly planned before Thanksgiving, New Constructs valued Anthropic at $150 billion against a reported $2 trillion ask, called it the most ridiculous IPO of the year and predicted “an unprecedented test of investor gullibility.” The firm called WeWork correctly in 2019, and no prospectus is public.
- Cheap verdicts: OpenAI’s new Decisions API returns yes/no probabilities and category picks at $0.10 per million input tokens with output tokens free, running about ten times faster than the Responses API, in public beta on gpt-6-luna only.
- Watermarking: Google opened SynthID Detector to the public and says “more than 180 billion watermarked images and videos to date” now carry the mark, with OpenAI, Nvidia and Kakao also supporting it and a million verification requests a day.
- Connectors: Meta’s Muse launched on iPad a month after its phone debut, adding Asana, QuickBooks, GitHub, Zoom, Notion and Granola, on 6.6 million installs by Sensor Tower’s estimate.
- Agent inboxes: Carly gives each of its agents its own email address, so the agent can send or draft for approval from an inbox that is not the user’s own.
- Domain land rush: ICANN published 1,615 new top-level domain applications from 481 applicants, and ten companies including Meta and OpenAI want
.agentwhile seven, OpenAI among them, applied for.agi. OpenAI applied for 15 strings in total, Meta for 21. - Security access: Anthropic widened its Cyber Verification Program to more defenders and red teams, allowing vulnerability research, malware analysis and penetration testing that public models refuse, with the top tier vetted jointly with the US government.
- Attribution: CrowdStrike pinned the South Korean bank intrusions on a likely single Chinese-speaking attacker using the open-source ARTEX pen-testing tool over DeepSeek, GLM and Grok models, which is a meaningfully different claim from autonomous AI robbing banks.
- Teen safety: OpenAI’s first teen-usage report says teens average less than 15 minutes a day, published the same day Common Sense Media rated ChatGPT for Teens an “unacceptable risk” after testers hit just two break reminders across nearly 2,000 prompts. OpenAI replied that the bulk of the testing “may have begun and concluded before activation of parental controls was complete.”
- Virtual cells: Biohub is coordinating a $1.8 billion effort to predict how cells behave, with $300 million from Google, Meta and Isomorphic Labs combined and more than $500 million over five years from the Department of Energy.
- Inflation: the September FOMC minutes record “several participants” saying “the scale and pace of the AI buildout had continued to surprise to the upside,” and a couple arguing a higher policy rate would stop “price increases stemming from energy market disruptions and AI-related demand” from broadening out.
- Funding: Nous Research raised a $90 million Series B led by Robot Ventures and launched Hermes for Businesses, Healthleap took $38 million for software that flags at-risk patients in more than 50 hospitals without diagnosing them, and Ledgebrook raised $200 million co-led by Allianz X and Rockefeller Capital Management for AI underwriting.
- Gadget postmortem: Tony Fadell said at MIT Future Fest that the Rabbit R1, Humane Ai pin and Limitless pendant each “didn’t meet any kind of need”, and called the assistant framing a projection: “It’s less than 0.01% of the world population that has ever had a human [assistant] do anything.”
- Benchmark gap: a new intuitive-visual-reasoning benchmark put humans at 93.1% accuracy against GPT-6-astra’s 53.6%, even at maximum reasoning effort.
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
