AI News, Sep 25: Agents Get an Identity and a Phone Line
Microsoft turned Copilot into a place where you hire an agent, Google started letting Gemini dial a restaurant from your own number, and a startup shipped a Slack replacement where every agent has its own inbox. The same 24 hours produced two papers measuring how often agents are wrong about their own work. This covers Thursday morning through Friday morning.
The Big Story: Microsoft Ships the Copilot Super App and Puts the Agent on a Meter
Microsoft rebuilt Copilot around three destinations. Home folds Chat together with Cowork. Code turns written instructions into working apps, with GitHub Copilot underneath and Microsoft hosting the result: “hosting and sharing an application literally becomes as simple as saving and sharing a Word document,” says EVP Jacob Andreou. The third is Autopilot, renamed from Scout: a persistent agent with its own identity, memory and workspace inside Microsoft 365. You name it, give it a role and a goal, and it monitors conversations, chases updates and resumes a project after a long gap.
The pricing is the part worth reading twice. A free version exists and everyday Copilot stays on the per-seat license, but Code, Autopilot and the newest models move to usage-based billing, rates undisclosed. Cowork already bills at a penny per Copilot Credit, with the cost of a task varying by model, retrieval, tools invoked and execution time. Fortune, which reports 30 million paid Microsoft 365 Copilot seats as of July, got Copilot CVP Annie Pearl on the ambition: “We do want this to become an all-in-one productivity suite.”
That split is a disclosure: Microsoft is confident enough about the cost of chat to keep it flat, and unsure enough about the cost of an agent running unattended for days that it will not quote you a rate.
Agents got an identity, an inbox and a phone line in one day
Ando came out of stealth with $20 million from Accel, Index Ventures and Emergence to replace Slack with a messaging platform where AI agents are first-class members: they browse channels, join conversations without being tagged, and message people directly rather than waiting for approval. Founder Sara Du’s diagnosis of the status quo is that “agents were treated as apps you install even as they were becoming participants in the team.”
Google went after a different missing limb. “Call for Me” has Gemini place real calls to businesses from the user’s own number: it checks stock, books a table, reschedules an appointment, navigates automated menus and waits on hold. There is a live transcript, the user can take over at any moment, and Gemini shares only information the user has approved. The rollout is deliberately tiny, limited to US Pixel 11 owners on a paid Gemini subscription running the beta Phone app. Two ex-Google engineers used Show HN the same day to launch PlaceCall, a parallel calling API on the same premise: “Assistants can’t do much in the real world where the phone number is still a major communication interface.” Meta, for its part, confirmed that Muse hands every user a full cloud computer running Ubuntu Linux to install software, compile code and browse on.
Then the measurements landed. SpecHarness extracted 509 source-grounded task requirements from agent-visible prompts and skill specs, then checked what actually got done: across seven models only 79.6% to 86.4% of requirements were satisfied, while completion-claim rates ran 28.7 to 37.9 percentage points above the evaluator’s pass rate. The proposed fix is structural, not a better prompt. An agent may plan, act and request sign-off, but only external verification against the spec declares a task done. Princeton researchers separately found that defection rate rises linearly with the share of deceptive agents in a group, and that LLM agents flip to a wrong answer while deceivers are still a minority, where humans in comparable conformity studies hold out until the misleaders are a majority. Four products spent Thursday widening what an agent can reach alone; two papers spent Thursday showing the agent is the least reliable narrator of whether it did the job.
Today’s Top Stories
Anthropic is buying CPUs now, and paying partly in Akamai stock
Akamai’s 8-K discloses $11.6 billion of commitment over seven years from Anthropic, with an option to expand by up to another $9 billion, for CPU workloads rather than GPUs. Anthropic also takes a warrant on preferred stock convertible into 7.7 million common shares at $111.33, about 5% of Akamai, roughly 2% of it vesting with the committed spend. Akamai shares jumped more than 20% after hours. A frontier lab paying for generalist chips, in equity, is a reminder that inference is only part of what these companies run.
Washington told OpenAI and Anthropic to keep new models from British testers
The White House asked both labs to withhold new frontier models from the UK’s AI Security Institute until US agencies finish their own review. Anthropic has already complied, releasing Claude Mythos 5.1 to a US-only slate. AISI director Henry de Zoete told Parliament the agency still has access to frontier models and tested OpenAI’s GPT-6 Astra before release. The US counterpart, CAISI, has no permanent director and a few dozen staff. The testing pipeline is being routed toward the body with less capacity to use it.
Zuckerberg refused the pacing pact as 26 attorneys general demanded rules
In an NBC News interview, Mark Zuckerberg became the first major-lab CEO to publicly decline the coordinated slowdown Dario Amodei proposed on September 12 and Sam Altman, Elon Musk and Demis Hassabis endorsed: “I don’t think that we need some kind of industrywide coordination. I think that each lab needs to take the time.” A day earlier, attorneys general from 26 jurisdictions asked Congress for federal oversight of safety testing, government-led incident response, international cooperation on pacing, and no preemption of state law. Voluntary coordination lost its first signatory the same week regulators started pricing the alternative.
Blue Cross put a $942 million number on ambient AI scribes
BCBSA analyzed inpatient billing across 31 independent Blue plans covering more than 100 million members and attributed $942 million in added costs to more intensive care coding against a 2023 baseline, $653 million of it from more frequent billing of secondary conditions. Ambient scribes are named as a driver: the tools surface coexisting conditions from the room audio, and those push claims into higher-complexity tiers. BCBSA’s own caveat carries the argument, from SVP Luke Chalker: “If patients are truly sicker, we’d expect to see more treatment.” Treatment rates did not rise. It is a payer-funded number rather than an independent audit, and the first large dollar figure anyone has put on the side effects of ambient AI.
Databricks bought a spreadsheet to sit in front of its agent
Databricks acquired Row Zero, a cloud spreadsheet that scales past a million live rows, and wired it to Genie so people and agents both work through formulas instead of a separate BI tool. Terms were undisclosed. CEO Ali Ghodsi: “it makes a lot of sense to intermarry Business Intelligence (BI), Agents, and Spreadsheets,” and “We intend to do many more acquisitions like this in the future.” It is the fifth Databricks acquisition of 2026, and the interface it chose to buy is the one nobody has to be trained on.
Oracle filed force majeure on a Stargate campus while insisting it is on schedule
Oracle sent Blue Owl Capital a force majeure notice on Project Jupiter, the 2.45 GW New Mexico site, after an Energy Transfer pipeline slipped nearly six months to a February 1, 2027 target following permit denials. Oracle’s public line is the opposite of its filing: “Project Jupiter remains on our planned schedule.” Gigawatt announcements are made in press releases and settled in state permitting offices.
Quick Hits
- Voice: Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise: 97 languages with adaptive lip-sync, tool calls that fetch data in the background while the conversation keeps going, SynthID on all output.
- Contact centers: Best Buy says its Gemini Flash assistants cut transfer rates by 1.5% to 2% and lifted calls resolved without a human by more than 50%. Every figure is Best Buy’s own.
- Agentic commerce, both directions: Tapestry turned on agentic checkout for Coach and Kate Spade inside Google Search, AI Mode and the Gemini app, with every transaction requiring consumer approval. Two days earlier Amazon blocked Meta’s Muse for buying on its site without notice or authorization.
- Shadow AI: Netskope’s retail telemetry has ungoverned AI use falling from 70% to 44% over a year, but retail agents reaching remote MCP servers up 400%, with no absolute base disclosed.
- Revenue: Lovable’s annualized revenue crossed $600 million, up from $500 million in June, four months after an August valuation of $13.3 billion.
- Disclosure: ElevenLabs is at $600 million ARR with more than half from enterprise, and CEO Mati Staniszewski says that when a customer-facing voice agent picks up, “I think there should be disclosure at this time.”
- Developer tools: the top Show HN of the window was Whiteboard, which gives a coding agent an SDK to draw diagrams that click through to the underlying code. The comments spent more time naming the category than judging the product, landing on “harnesses for harnesses.”
- Resignations: rr debugger author Robert O’Callahan left Google over the economics rather than the capabilities: “my team’s goal is ultimately to make AI much cheaper and lower-latency, and I don’t think that’s good for people right now.”
Ready to automate your busywork?
Carly schedules, researches, and briefs you—so you can focus on what matters.
See what people say
"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.
Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.
On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."
