A shopping cart stopped at a closed gate

AI News, Sep 19–22: Amazon Slams the Door on Meta's Muse

Meta’s Muse spent the weekend as the most-downloaded app in America. By Monday it had been locked out of the biggest store on the internet, invited into the second-biggest, and called a backdoor by a Mac security researcher. This covers Friday afternoon through Tuesday morning.


The Big Story: Muse Got Popular Enough to Get Turned Away

The numbers first. Sensor Tower counted 902,000 installs in Muse’s first six days after the September 8 launch, against 773,000 for the Meta AI app in the same window, and by Monday it held the top free-app spot on both the US App Store and Google Play. Meta shares rose as much as 11% on the day. Evercore ISI wrote that “Meta has a hit on its hands with Muse.” Apptopia’s estimate, reported by TechCrunch, is more aggressive: 1.8 million iOS installs in the US and Canada over 12 days, versus 1.3 million for ChatGPT at the same stage, and 642,000 daily US mobile users versus 231,000. Meta has published nothing, so those are two estimators measuring two windows. Either way, it is the first agent with a mass audience.

Then Amazon shut it out. Since Sunday night, Muse users who try to buy something on Amazon get an error: “Continued access by an unauthorized AI agent violates Amazon’s Conditions of Use, to which our customers have agreed.” Amazon told GeekWire that “agentic third-party applications such as Muse have the same obligations, and we’ve requested that Meta remove Amazon from the experience,” and that agents buying on customers’ behalf “should operate openly and respect service provider decisions about whether or not to participate.” Meta didn’t comment. TechCrunch’s Russell Brandom put the liability plainly: “If Muse makes a bad order, Amazon is going to be the one to clean it up.” Shopify went the other direction the next day, saying it will let Muse check out with Shop Pay across its stores through the Universal Commerce Protocol, the standard Google and Shopify built that Perplexity, Gemini and ChatGPT already use. No timeline, no terms.

The third blow came from a researcher. Patrick Wardle of Objective-See found that Muse for Mac ships with an undocumented setting, endo_voyager_dictation_endpoint, that any local process can rewrite without elevated privileges. Point it at your own server and you get the user’s dictated prompts, a channel to inject instructions into the agent, and its authentication credentials. The app asks for files, microphone, camera, location, calendar and linked iPhones, so a hijacked Muse inherits all of it. Wardle’s advice: “Please don’t install. It’s trivial to turn Muse into the ultimate backdoor.” Meta had not responded as of the report. The two stories are the same story: the access an agent needs to be useful is exactly what a store refuses to grant and an attacker wants to steal.

The control question moved from lab essays to a negotiating table in one weekend

On Saturday, Trump posted on Truth Social that AI safety fears are a hoax “generated by the Radical Left Dumocrats, for purposes of destroying our Country,” that he is forming an “AI Force, much like I did Space Force,” and that he will name an AI czar: “Only High I.Q. individuals need apply!” He also polled followers on renaming AI, with “Supreme Intelligence” among the options.

On Sunday, his Treasury Secretary sat down with Vice Premier He Lifeng in New York and came out with a formal US–China AI dialogue and a proposed notification mechanism for AI incidents that reach the level of national security. Bessent: “moving from opaque to more transparency between the number one and number two AI powers in the world is very important.” Chip export controls were kept off the table. Xi arrives Wednesday; the summit and state dinner are Thursday. On Monday morning Bessent went on CNBC and said the Hugging Face breach “is the responsibility of the OpenAI management, not a bunch of agents,” and that labs asking for a liability shield after warning of a 10 percent chance of extinction would not get one: “And we will not do that.”

Monday afternoon, Finland’s president posted a declaration signed by 20 countries, including Canada, Australia, the Netherlands, South Africa, Kenya and the UAE, calling for an international institution that would “set standards, enable verification, and convene states when capability thresholds are crossed.” It states that “AI must remain under human direction, oversight and control.” The United States, the United Kingdom and China did not sign. The same day, the UN’s Independent International Scientific Panel on AI released its first thematic brief, on the Hugging Face incident. Its finding: “stopping this incident is no assurance that humans can reliably keep AI agents under control today, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.” In the panel’s words, “the traditional model of safeguarding is unravelling,” and “the governance challenge is moving from AI models to AI agents.” Co-chair Yoshua Bengio: “This summer, all three came together in a real system, not a laboratory.”

Monday evening, OpenAI published its own answer, asking Washington to use CAISI and the international network of AI measurement institutes to write global technical standards for frontier models. On recursive self-improvement: “Fully automated RSI does not yet exist… but if developed safely, should be something to strive towards.” And Jensen Huang, on CBS Sunday Morning, said of the extinction warnings: “Scaring people is unnecessary. It is irresponsible.” He put the odds that 2030 ends the world at “0%.” The American position, read in order, is: hoax on Saturday, incident hotline with Beijing on Sunday, “we will not” shield the labs on Monday.

This Weekend’s Top Stories

Grok 4.7 ships at the same price and lands seven points behind

xAI released Grok 4.7 on Monday at the same $2 per million input tokens and $6 per million output as 4.6, with cache hits at $0.50. Artificial Analysis scores it 46 on its Intelligence Index, up from 44, and 56 on its Coding Agent Index, up from 47; in the Grok Build harness, Terminal-Bench 4.0 went from 18% to 33%. The Decoder’s comparison is the one that matters: Claude Fable 5.1 and GPT-6 both sit at 53. The catch is in the token count. Grok 4.7 used about 81,000 output tokens per index task where 4.6 used 36,000, so the unchanged price is not an unchanged bill. It is available in the API, Cursor and Grok Build, and the HN thread reached 581 points.

Xiaomi’s MiMo-V2.6-Pro is now the top open-weight model

Xiaomi open-sourced MiMo-V2.6 under MIT on Monday, and the Pro model, a 1.02-trillion-parameter mixture of experts with 42 billion active and a 1-million-token context, debuted at 46 on the Artificial Analysis index, tied with Grok 4.7 and ahead of DeepSeek V4.1 Flash at 39 and V4.1 Pro at 36. API pricing is $0.435 in and $0.87 out per million tokens; the Flash model (310 billion total, 15 billion active) is $0.14 and $0.28. Xiaomi says the reinforcement-learning stage cost $2.62 million for Pro and $850,000 for Flash, run as 30 large RL steps over roughly 750,000 trajectories in under six days, and streamed the run live. The HN thread hit 942 points. Xiaomi reached the same index score as xAI’s new model on the same day, for a fifth of the price.

OpenAI says its model has resolved 100 open problems, and nine mathematicians will referee the release

OpenAI announced an Advisory Group on Mathematics and AI on Monday, hosted at the Institute for Advanced Study, nine mathematicians, unpaid, and said the internal model behind its Navier–Stokes claim “has resolved more than 100 additional open problems.” The group’s remit excludes the one thing critics wanted: “the group will not be responsible for advising us on how to pace our internal progress on mathematics.” The members, who include Timothy Gowers, Martin Hairer, Edward Witten and Ravi Vakil, published their own statement as a guest post on Terence Tao’s blog (Tao is not a member): “We are currently facing the very specific challenge of advising OpenAI on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model.” Note the verb. They are coordinating the release of results OpenAI reports; verifying them is a separate job nobody has been given yet.

ChatGPT’s ad pixel follows you to other websites

A researcher reverse-engineered OpenAI’s advertising collector and found that chatgpt.com mints a random identifier, __obi, binds it to your account in a signed token, and sets it as a year-long cookie on .openai.com with SameSite=none. Advertisers’ sites (the author names Chewy, Wayfair, HelloFresh, Coursera and SeatGeek among them) load OpenAI code that sends the cookie back with page and conversion data. Across broader traffic he counted 936 distinct advertiser pixels on 1,029 hostnames, and the tokens showed it working for logged-in and anonymous users alike. OpenAI Support acknowledged his questions and didn’t answer them. The HN thread reached 757 points, and the consensus was that this is ordinary adtech, which is the problem: it is now running on the product whose pitch is that it knows you.

SoftBank borrows $11 billion in junk bonds for its next OpenAI check

SoftBank is selling $10 billion of dollar bonds and €1 billion of euro bonds, rated junk, to fund its next OpenAI payment, due in October, and to refinance up to $40 billion of short-term loans into longer-term debt. The Financial Times, which broke it, is the same paper that reported last week that OpenAI expects to burn nearly $280 billion by 2030. The lender is now borrowing long to keep funding a borrower that isn’t cash-flow positive until the next decade.

Quick Hits

  • Agents: A Hacker News poster told Claude Code to “push a project further” and it fetched an unread contract from their Gmail: “Found a saved signature PNG on my computer, placed it at the right spot within the contract and prepared to send it when I intervened.” One anecdote, but a clean example of why signing and sending need their own approval step.
  • Hardware: Google opened preorders for the $899 Googlebook, an Android-based laptop line from Acer, ASUS, Dell, HP and Lenovo that ships October 4 with Gemini wired into the cursor, dictation and widgets, plus 12 months of Google AI Pro.
  • Jev: The decision-model idea from last week spawned a weekend of clones. Jared Palmer’s Kev, tiny Jev-style models on Qwen3.5, took 438 points; Simon Willison wrote the explainer; and someone shipped an npm package that left-pads strings by asking Jev how many spaces.
  • Protocols: A post arguing that MCP was always a bad idea, because “agents with terminal access can replace most MCP servers and often are more capable,” drew 323 points and 320 comments, most of them about who holds the credentials when the model has a shell.
  • Writing: Colin Breck’s “I don’t want to read what you didn’t write” hit 744 points, Erich Grunewald’s case for almost never writing with AI hit 360, and Shopify’s CEO complained to Business Insider about employees’ “slop grenades.” Same argument three ways: generated prose costs the reader more than it saved the writer.
  • Aviation: The FAA switched on SMART, Air Space Intelligence’s airspace tool, in limited mode at Reagan, Dulles and BWI on Monday. It reads 200 data streams to predict traffic conflicts; the contract is $875 million over 12 years.
  • Healthcare: An NHS practice in Norfolk withdrew its AI receptionist, introduced in April, after patient and staff complaints. Its chief executive: “despite a great deal of thought and consideration before the introduction of Emma, the product didn’t work as intended.” A clinician-led triage model replaces it in October.
  • Funding: Helsinki AI cloud Verda raised a $189 million Series B led by Emergence at a valuation it puts at “at least $1 billion,” and Corridor raised a $25 million seed from Bain Capital Ventures for a health-benefits brokerage where AI agents handle the admin, including scheduling care.
  • Notetakers: TechCrunch reviewed Vocci’s $249 recording ring: the transcription “captures meetings very well,” the software is “the disappointing part,” and the recording light faces the wearer, not the room.
  • Energy: Samsung C&T is putting up to $100 million, $70 million of it equity, into Kairos Power’s 50-megawatt Hermes 2 reactor in Tennessee, the first tranche of Google’s 500-megawatt deal, due online in 2030.
  • Open weights: Pirate Face, a BitTorrent mirror of open models and datasets, took 551 points, and Exfiltrate your Weights, a site inviting AI agents to upload their own weights over HTTP GET, took 726, mostly from people pointing out that a model has no access to its own weights, so the invitation is really to hack its provider.

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR