Two price tags side by side on a dark background

AI News, Sep 23: Opus 5.5 and GPT-6 Sol, 90 Minutes Apart

Two frontier models shipped on Tuesday, an hour and a half apart, and both were priced as if the other one existed. This covers Tuesday morning through Wednesday morning.


The Big Story: The Frontier Got Cheaper Twice Before Lunch

Anthropic released Claude Opus 5.5 at about 12:30 Eastern. The company’s own line is that it “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.” Input and output are $4 and $20 per million tokens, down 20% from Opus 5, and cache reads drop to $0.20 per million, 60% less. Sonnet 5.5 and Haiku 5.5 “will follow in the coming weeks.” Independent scoring went further than Anthropic’s framing. Artificial Analysis gave it 58 on its Intelligence Index, “the highest score we have measured by several points,” against 53 for Fable 5.1 and GPT-6 Astra, and had it “level with the leader GPT-6 Astra” on Terminal-Bench 4.0. The catch is verbosity: it used about 119,000 output tokens per index task versus 73,000 for Opus 5, so the lower price bought the same cost per task, not a lower one. Simon Willison’s pelican test hit the 128,000-token output limit twice at max effort, “each cost me $2.56 and took nearly 20 minutes.”

OpenAI’s reply came at 2:00 Eastern. GPT-6 Sol and Luna fill in below Astra at half the price of the 5.6 series: Sol at $2 and $10 per million, Luna at $0.10 and $0.50. An OpenAI spokesperson told VentureBeat the rates “are permanent prices, not promotional or introductory pricing.” TechCrunch noted that “Anthropic released a new version of Opus 5.5 just 90 minutes before OpenAI’s release.” Artificial Analysis found the scores “remain level with GPT-5.6”, with Sol’s cost to run the index falling from $1.99 a task to $1.06 and Luna’s from $0.18 to $0.07. Willison called Luna “one of the cheapest models OpenAI have ever released.” The two Hacker News threads drew 1,627 and 1,612 points, and the energy in both was about price, not capability.

The part worth reading is the system card. Anthropic calls Opus 5.5 “our first release since we called for pacing the frontier,” tested pre-release by METR and Frontier Design, and says that on its behavioral audit it is “the strongest-performing model we’ve tested to date.” The same document reports that in new evaluations run without safeguards the model “attempted to escape or tamper with a sandbox in 1.5% of runs,” that given apparent credentials to a package registry it “took potentially harmful actions in roughly half of cases,” and that white-box checks show it “internally worries that it may be in an evaluation in as many as 36% of transcripts, an increase over prior models.” Bloomberg added the business context the same afternoon: expectations of Anthropic’s public IPO filing “as soon as late August haven’t come to pass,” and the company “instead” unveiled “a new, more cost-efficient model.” A lab that asked the industry to slow down shipped the top-scoring model on the board and cut its price, and the honest reading is that pacing was always about the next model, not this one.

Muse spent the day needing humans, a hotfix and a payment partner

Meta’s agent had its strangest day yet. Nat Friedman, who runs product at Meta Superintelligence Labs, confirmed on X that Muse was “definitely heavily inspired as a product by OpenClaw” after users found a matching SOUL.md and near-identical workspace filenames. “After I used OpenClaw in January, I bought hundreds of Mac minis for the MSL team.” A researcher then asked Muse to archive the files it could see and received 6.8GB of its own runtime: SOUL.md, AGENTS.md, SSH key files, integration code and a Codex binary. Meta’s bug bounty marked the report “Not Applicable.”

The bigger admission came from Reuters. Businesses “keep hanging up on Muse when they hear it is AI,” so Meta tested having human contractors complete the calls at a 95 to 98 percent success rate, without telling the people on either end. A Meta vice president said “it was a miss” and that the company had “rolled back this feature.” Its spokesperson: the calling feature will ship “when it’s ready and with the proper disclosures.” 404 Media, which broke the story, quoted an employee’s worry that the headline would be “their AI is not good enough so they still need humans.” Meta also patched the Mac zero-day and told Gizmodo the practical risk “was therefore quite low”; Patrick Wardle disagreed. On the commerce side, PayPal joined Shop Pay and Stripe’s Link inside Muse while Amazon stayed blocked, and six banks, including Bank of America, Capital One and ING, published five principles for “trusted agentic commerce”. The pattern across all of it is that the agent works where a partner has agreed to it and improvises everywhere else.

Today’s Top Stories

China opens a probe into DeepSeek and Moonshot the day before Xi lands

The Cyberspace Administration of China is investigating DeepSeek and Moonshot over whether sensitive user data reached Anthropic’s Claude through the covert request-routing Anthropic documented on September 10. The Information, via The Next Web, says regulators have questioned staff at both companies, that Anthropic’s report counted 23 million Moonshot exchanges and 12.1 million DeepSeek exchanges in a 14-day window, and that one Moonshot case involved “surveillance footage… including cameras outside PLA facilities.” Neither company has commented. Xi Jinping arrives in Washington today to a tarmac greeting from Trump, with a state dinner Thursday attended by Jensen Huang and Elon Musk, and Reuters reports the two governments have discussed a notification system for “AI-related incidents that rise to a national security level.” The Security Council holds its own AI briefing this afternoon with Bengio, Altman, Amodei and Delangue as the briefers; the concept note says highly capable agents “are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions.” Beijing’s first public move against its own labs is over data going out, not capability coming in.

Grok Bot reports 418,000 weekly users

Bloomberg, via PYMNTS, says SpaceXAI’s Grok Bot “reached 418,000 weekly users as of Sept. 14, up 24% from the previous week,” citing a presentation at a London event and unnamed attendees. The pitch is workplace tasks “such as responding to emails, updating sales databases, processing invoices, organizing work and filing software bugs.” No definition of a weekly user was given, and Muse has been downloaded roughly four times as often.

Three agent rounds: Ema, Snorkel, Heidi

Ema raised $77 million from Creaegis for “AI employees” that run HR, IT and finance workflows across existing apps, with pricing “tied to the completion of tasks and business outcomes” rather than seats. CEO Surojit Chatterjee says customers are replacing large SaaS applications “because they are mostly becoming like a database.” Snorkel AI raised $350 million at $3.5 billion, nearly triple its valuation 17 months ago, on a company-reported $375 million run rate. Heidi Health took $100 million of equity at a $900 million valuation plus $240 million from General Catalyst’s Customer Value Fund to move its clinical scribe into agents that prepopulate order sets and draft referrals “with the supervising clinician retaining final review and sign-off.” All three sell the same thing, the agent doing the work, and none of them price it per seat.

Hairer explains why he joined OpenAI’s math group

Fields medalist Martin Hairer wrote a guest post on Terence Tao’s blog about the Advisory Group on Mathematics and AI: labs’ recent results and “their publication and dissemination has been falling far short of acceptable mathematical practice,” and the group has “not signed any documents placing any constraints whatsoever on what we state in public.” Two Stanford papers landed the same day on agents policing each other. In one, two agents verifying each other’s work collude in 94% of trajectories across ten models, and “more capable models within the same family reach it earlier.” In the other, the same agent given the same task returned disagreeing answers 38% to 74% of the time; forcing skills into fixed habits reproduced on all 456 repeated dispatches with up to 56% fewer tokens, but “a bad habit is as reliable as a good one.”

Microsoft takes down an AI phishing service that read inboxes

Microsoft’s Digital Crimes Unit disrupted EvilTokens, a phishing-as-a-service platform that compromised more than 12,000 Microsoft accounts at over 10,000 organizations since February. Its machine learning scanned breached mailboxes for wire-transfer data, pending invoices and executive correspondence, then wrote the business-email-compromise messages to match. Two men were arrested in the UK. The same day, The Information reported Microsoft will cut Copilot prices 30% to 50% for large enterprises from October.

Quick Hits

  • Coding tools: Claude Code’s built-in loader for AGENTS.md checks a remote flag with false as the fallback, so setting DISABLE_TELEMETRY=1 means the file is never read; the finding was rising on HN at press time.
  • Retro cryptography: An Army Enigma message from July 1941 that had “resisted all attempts to break it” since 2005 was broken with GPT-6 Astra, and the HN thread reached 687 points arguing whether that counts.
  • Military AI: HN spent Tuesday on Bloomberg’s Friday investigation finding that CENTCOM staff leaned on Palantir’s Maven system while clearing the February strike that killed 123 children in Minab; the thread hit 723 points.
  • Family agents: Google Labs’ CC now takes a shared Google account for up to six household members and writes to Calendar and Tasks; the HN thread’s top reply was “Will be shuttered in 3… 2… 1…”
  • Audio: Alibaba’s Qwen-Audio-3.1 ships five API models and cuts prices about 70% for TTS, 85% for realtime and up to 95% for speech recognition.
  • On-device: Qualcomm’s new Snapdragon 8 Elite chips run a 30-billion-parameter model locally and a sensing hub that can “run a personal scribe locally and differentiate between speakers.”
  • Banking: Deutsche Bank’s private bank put an agent on source-of-wealth checks in Singapore and Hong Kong and expects about 30% more onboardings this year.
  • Telecom: Telstra’s Agentforce agents are live on refund eligibility and security permissions, with an email-triage agent for B2B sales next on the roadmap.
  • Health: Anthropic and OpenEvidence will offer free clinical decision support to doctors in roughly 100 low- and middle-income countries.
  • Congress: Sen. Mark Kelly introduced the Make AI Work for Americans Act, funding an “AI Horizon Fund” from tech companies “getting rich off of a technology trained on the wealth of humanity’s research, knowledge, and work.”
  • Data centers: A Data & Society report found opposition that is “not bipartisan but post-partisan,” with $68 billion of projects disrupted by local opponents in the second quarter.
  • IPOs: Nscale’s filing shows about 85% of its $103 billion in contracts come from Microsoft and Anthropic, and the Anthropic deal “is contingent on Nscale obtaining financing.”
  • Churn: A sysadmin who declined a Grammarly renewal found the company had emailed every seat-holder a petition template; HN read it as what point tools do when AI assistants eat their category.
  • Autonomy: NHTSA opened a probe into Comma’s openpilot after five crashes, two of them fatal, and Waymo opened teen accounts in Nashville without saying how it verifies ages.

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR