A dark control panel with one switch flipped and a signal line running out of frame

AI News, Aug 26: ChatGPT Tasks Now Trigger on Your Email

Two labs shipped the same missing piece on the same day, from opposite ends. Then a $18 billion settlement landed and a stealth model took its mask off.


The Big Story: ChatGPT Tasks Stopped Waiting for the Clock

OpenAI added event-based triggers to ChatGPT’s scheduled tasks on August 25. Paid tiers (Plus, Pro, Business and Enterprise) can now fire a prompt off a webhook from Gmail, Slack or GitHub: new mail arrives, a channel gets a message, a pull request moves, and the task runs. Free accounts got the scheduling hub for the first time, capped at three active tasks running once or no more than daily, with no webhook triggers. Every tier can now share a task so a recipient gets an independent copy of it.

Until now, a ChatGPT scheduled task was a cron job. You picked a time and it ran at that time, whether or not anything had happened. The webhook version is a different product: the model sits idle until the world does something. That is the difference between a reminder and an automation, and it is the feature that has separated assistants from workflow tools for the last two years.

The three sources OpenAI picked are worth reading closely. Gmail, Slack and GitHub are the developer-and-inbox triangle, not the calendar. There is no meeting trigger, no booking event, no “when this gets rescheduled.” The scheduling coverage from Dataconomy confirms the same three. OpenAI built the half of the trigger surface that its own Codex users were already asking for.

Every agent release today was about what happens when you are not typing

Anthropic shipped the other half. Claude now carries one memory across chat and Claude Cowork, so a Cowork task starts from what was said in a chat window, using Anthropic’s own example of conference headcount and speakers mentioned in conversation flowing into a drafted agenda. Memory is now written continuously during a conversation rather than summarized at the end, and everything stored shows up as editable files under Topics in settings. Health, race, religion, politics and gender identity are excluded by default and opt-in; government IDs, Social Security numbers, criminal history and immigration status are never stored at all. It is on by default for Free, Pro and Max across web, desktop and mobile.

Set those two side by side. OpenAI gave its agent a reason to wake up. Anthropic gave its agent something to remember when it does. Both are answers to the same complaint, which is that an assistant that only exists inside the text box you are currently looking at is not an assistant.

The research on arXiv the same day was doing the identical thing from underneath. Recuris splits agent memory into working memory for progress tracking and experiential memory for skills, and adds up to 32.2 points on the longest-horizon tasks, improving 35 of 37 model-benchmark pairs it was tested on. Prime Agent preserves histories, skills and subagent specs across trajectories so a model can program over its own past, taking ARC-AGI-3 Best@1 from 30% to 95.5% and edging the 95.4% human-expert baseline. And Headlong, an open-source harness from Laude Institute and MIT written in under 10,000 lines of Bash, runs an agent as an infinite self-guided loop where your incoming message arrives as an observation inside a monologue that was already running.

Four different teams built persistence and a heartbeat in the same twenty-four hours. Nobody built a calendar trigger.

Today’s Top Stories

Meta Settled With 48 States for Up to $18 Billion

Meta reached a settlement mid-trial in Oakland federal court with 48 state attorneys general, ending the landmark case alleging its platforms were engineered to be addictive to minors. The terms are the interesting part: a default two-hour daily cap across Facebook and Instagram for under-18s, an overnight block from midnight to 6 a.m., push-notification blackouts during school hours, and a teen-selectable feed with no algorithm in it. Only 70% of the money is paid unless TikTok and YouTube adopt comparable one-hour defaults and pay states roughly $5.3 billion of their own. Reported totals vary between $16.68 billion and $18 billion depending on the outlet.

This is a recommender-system case rather than a generative-AI one, but it is the first time a US settlement has written specific algorithmic defaults into a product as a remedy.

Ox Alpha Was GLM All Along

Bloomberg got Z.ai to confirm that Ox Alpha, the stealth model that had been quietly beating benchmarks on OpenRouter, is a new GLM-series iteration, with weights slated for August 28. That closes a guessing game that ran for most of last week, and it closes it against the consensus: the fingerprinting argument that Ox Alpha could not be GLM, because GLM has no vision encoder and Ox Alpha accepts images, was the reasoning that pushed everyone toward Kimi K3.5.

Testers report strong long-horizon coding, and there is still a visible gap between the unofficial benchmark numbers circulating and the more modest official ones.

23% of the ICML Papers Anyone Checked Had a Falsified Claim

Science covered the Hugging Face and alphaXiv reproduction challenge run against ICML 2026, and the numbers are worse than the framing suggests. In 19 days, 1,221 participants published 6,816 reproduction logbooks covering 2,226 of the conference’s 6,352 accepted papers and evaluated 35,908 individual claims. Half the papers had at least one claim independently verified. But 496 of them, 23%, had at least one claim falsified or contested, and 49 had every claim falsified.

Organizers confirmed at least a dozen contain real errors human reviewers missed, including a paging-algorithm paper claiming H_k + O(1) robustness that actually achieves H_k + Θ(log k), and a self-distillation paper whose code used forward KL while the theory analyzed reverse KL. One contestant reproduced 363 papers alone. Peer review was never checking this, and now something is.

Amazon Is Shutting Down Mechanical Turk

AWS posted notice that Mechanical Turk, the crowdsourced task marketplace Jeff Bezos once called “artificial artificial intelligence,” closes on September 30 after 21 years. Amazon had already stopped accepting new MTurk customers on July 30.

The joke writes itself and is also the whole story: the service named for humans pretending to be a machine is being closed because the machines absorbed the labeling and transcription work, with what remains going to enterprise data-labeling specialists.

Stability AI Raised $76 Million From All Three Major Labels

The Stable Diffusion maker closed a $76 million Series B with an investor list that reads as a peace treaty: Universal Music Group, Sony Music Group, Warner Music Group and Electronic Arts, alongside AMD Ventures and Pacific Alliance Ventures, with Coatue, Greycroft, Sean Parker and Eric Schmidt returning. Total raised is now $232 million. It is the first AI company backed by all three major labels, and the plan is co-development rather than licensing.

Stability largely prevailed in Getty Images’ UK copyright suit and the US case is still pending, which makes the label money look less like vindication and more like the industry deciding to own a seat instead of buying one.

Quick Hits

  • Agent infrastructure: Arga Labs raised a $10 million seed led by General Catalyst to build digital-twin sandboxes of Salesforce, Workday and email clients, permission systems and webhooks intact, so agents can be trained and tested on cross-system workflows without resetting production.
  • Open weights: Alibaba released Qwen3.8-Flash-Next, 125B total and 6B active with a 51B N-gram embedding layer, claiming it trained at one-ninth the cost of Qwen3.7-Plus while beating it across the board. The N-gram technique is a first public release and reads as a Qwen 4 preview. Neither llama.cpp nor vLLM supports the architecture yet.
  • Enterprise: OpenAI shipped an Admin plugin for ChatGPT Work and Codex letting workspace admins review adoption and credit usage, manage members and groups, gate model access by role, and approve or deny spending requests conversationally.
  • Departures: OpenAI lost Chris Malone, head of data centers since March 2025, after an infrastructure reorg moved him out of Greg Brockman’s reporting line. Business Insider counts 13 executive exits in 2026, against a $500 billion Stargate buildout.
  • Robotics money: Generalist reached a $3 billion valuation on roughly $200 million led by 8VC, two months after closing at $2 billion, and self-driving truck startup Gatik raised $200 million led by the Qatar Investment Authority on the back of $600 million in contracted freight revenue.
  • Robotics reality: A useful counterweight to both: general-purpose robots still land around 80% task success, which is not enough to pay for. Unitree floated at $66 billion and promptly lost half of it, and Foxglove’s CEO predicts an Apple II moment rather than a ChatGPT one.
  • Agent startups: Bengaluru’s Runable took $21 million at a $65 million valuation on 1.7 million registered users and $2 million ARR three weeks after enabling payments, with gross margins currently negative because it subsidizes inference. India’s Ringg added $10 million from Peak XV on 20 million monthly call attempts, and QueryStory left stealth with $6 million and an analytics product whose pitch is showing you the SQL it ran.
  • Consolidation: Gamma acquired Lica, the Accel-backed design startup, folding both founders into a new design research division; Gamma is at a $2.1 billion valuation with $100 million ARR. Accenture is separately acquiring Dutch SAP specialist McCoy and its 380-plus consultants into its mid-market unit.
  • Adoption: CompTIA’s first Corporate AI Adoption report, from a June survey of 1,027 professionals, found 59% now prioritize integrating AI into their existing stack rather than handing out more tools, while 78% concede their data practices need work. The bottleneck moved from access to operations.
  • Sentiment: Pew surveyed 3,488 US adults and found 39% say chatbots hurt more than help people using them for loneliness, against 19% who say they help. Adults under 30 are the most negative group, at 49%.
  • Labor: Stanford payroll research finds employment declines concentrated in early-career workers in AI-exposed occupations. The live dispute is whether that is substitution or the post-ZIRP hiring contraction landing on the same cohort.
  • Regulation: Pennsylvania’s attorney general sued Snap over Snapstreaks, infinite scroll and disappearing messages, alleging Snap understated nudity and drug content to win an age-appropriate app-store rating. Separately, an EPA rule change would let data center air permits issue without public notice, removing the main lever communities have used against the buildout.
  • Information operations: The Guardian documented a fake US think tank funded by Israel with no legal entity, no address, no named staff and no bylines, whose reports appear optimized to be cited by LLMs and AI overviews. It is the second state-funded fake institute surfaced this month.
  • Harness tuning: Microsoft’s AutoSaddler treats agent-harness improvement as offline learning, diagnosing failures from execution traces and generating code patches, for gains of 9.0 on GAIA2, 9.6 on SWE-Bench Pro and 10.0 on Terminal-Bench 2.0. Manual prompt engineering is the thing being automated.
  • Silent bugs: A two-forward-pass audit requiring no training and no gradients found causality violations in shipped models Zamba2 and Nemotron-H, an inter-chunk axis error in chunked-scan code. Across 192 injected faults, standard attention-mask inspection caught zero and the audit localized all 192 to the exact layer.
  • Datasets: LAION released LAION-BVD, 80 million videos totaling 10 million hours pulled from 1.3 billion CommonCrawl URLs, with synthetic video and audio captions and extracted scene-change frames, under CC BY 4.0. It is the largest open multimodal video corpus so far.
  • Self-awareness: An analysis sampling Hacker News top stories in February and June put roughly half of daily front-page items in the AI-topic or AI-generated bucket, up from about 40%. Most of the 331-comment thread argues about the detector’s false-positive rate.
  • Working from home: Anthropic moved San Francisco staff remote ahead of a possible strike by Allied Universal security workers, who are employed by the building contractor rather than Anthropic. SEIU, which is organizing those workers across California, said the specific strike threat was news to them.
  • Essays: Bill Gates published a piece on the turbulent AI era arguing for designated human-reserved jobs and warning that AI risk exceeds what Big Tech admits. OpenAI CFO Sarah Friar separately published a full-stack compute-economics argument, citing GPT-5.6 Sol hitting a quality high on 54% fewer output tokens and invoking Jevons paradox; that one is on OpenAI’s own site with no outside coverage.

Ready to automate your busywork?

Carly schedules, researches, and briefs you—so you can focus on what matters.

See what people say

"Before Carly, I relied on a Calendly link, but the whole process felt impersonal and not very professional. Carly changed that by handling all the back-and-forth, so I'm no longer stuck in endless email threads trying to line up schedules.

Now Carly reaches out to candidates, shares my real-time availability, lets them pick a slot, then sends a Zoom link and drops it straight into my calendar. She sends reminders to both of us before each call, which has significantly reduced no-shows and last-minute confusion.

On top of scheduling, Carly acts like a full executive assistant, sending me my schedule the night before so I can prepare for each call. It reminds me of the old x.ai assistant, but Carly is noticeably smarter, faster, and better suited to my healthcare recruitment business."

Gus Ibrahim, Founder & Director, IHR