September 23, 2026

AI Next Wave - September 23, 2026

A paper price tag with a red line through it beside a face-down phone on a dark counter.

The frontier just got 40 to 50 percent cheaper in one afternoon

Anthropic and OpenAI shipped new models within about an hour of each other on Tuesday, and both led with price, not benchmarks. Claude Opus 5.5 costs 20 percent less per token than Opus 5 and, by Anthropic's own math, 40 percent less on real workloads. GPT-6 Sol and Luna landed at half the price of their GPT-5.6 predecessors. Neither company's benchmark table includes the other's new model, which tells you how fast this moved. Meanwhile Xiaomi open-sourced a trillion-parameter omnimodal model under MIT, Alibaba laid out a 10-trillion-parameter roadmap and a 20-gigawatt build plan, Meta's Muse went No. 1 and picked up PayPal and Shopify checkout, the UK Parliament summoned all four big labs, and New Jersey hit a data center with a record fine for running 62 gas generators without permits. Busy day. Here's what matters if you build with this stuff.


Models and research

Claude Opus 5.5: same tier as Fable 5.1, 40 percent cheaper than Opus 5

Anthropic released Claude Opus 5.5, the first model in its 5.5 family. Pricing is $4 per million input tokens and $20 per million output, down 20 percent from Opus 5, and cache reads drop 60 percent to $0.20 per million. The company says the model uses fewer tokens per task on top of the lower price, which nets out to a 40 percent cost drop on typical workloads and 30 percent faster output. It's the first Opus to ship with Fable 5.1-class safeguards for cyber and biology: most security tasks get rerouted to Opus 4.8, and biology work above a threshold requires enrollment in a verification program. Anthropic is unusually candid that benchmark margins are getting less reliable, that the real gap to Fable 5.1 is narrower than the scores suggest, and that Opus 5.5 "often suspects it is being evaluated," which makes alignment testing harder. Sonnet 5.5 and Haiku 5.5 are promised "in the coming weeks." Available on AWS, Google Cloud, Azure, and the Claude Platform as claude-opus-5-5. Thinking can no longer be switched off.

Why it matters: For a solo builder running agents, cache reads are most of the bill. A 60 percent cut there is the number to care about, not the headline benchmark. The "fewer tokens per task" claim is Anthropic's, so measure it on your own traces before you re-budget.

Source: Introducing Claude Opus 5.5 (Anthropic)

GPT-6 Sol and Luna: half the price of GPT-5.6, and Luna is nearly free

OpenAI's answer, published roughly an hour apart from Anthropic's: GPT-6 Sol at $2 input and $10 output per million tokens, and GPT-6 Luna at $0.10 and $0.50, both 50 percent below the GPT-5.6 promotional pricing. OpenAI credits caching and inference improvements and says cached input reads now get a 90 percent discount with higher hit rates by default. The company's benchmark claims compare against Claude Opus 5 and Fable 5.1, not Opus 5.5: it says Sol beats Opus 5 on Zapier's AutomationBench at 9 percent of the cost per task and comes within 1.1 points of Fable 5 on DeepSWE at roughly 80 percent lower cost. On its internal factuality eval, Sol makes about half as many mistakes as GPT-5.6 Sol. One telling stat from the post: OpenAI's median researcher now burns $600 a day in tokens at API prices, and the 90th percentile burns $7,000. Availability is ChatGPT Work and Codex for paid tiers, Luna on the desktop app for Free and Go users, and gpt-6-sol and gpt-6-luna in the API. Not yet in the standard Chat interface.

Why it matters: Luna at $0.10 per million input tokens is cheaper than most open-weight models you'd host yourself. If you've been running a small model to save money on classification, routing, or extraction, rerun the math. The comparison charts are OpenAI's own and stale by an hour, so treat the Opus numbers in that post as a floor for Anthropic, not a ceiling.

Source: Introducing GPT-6 Sol and Luna (OpenAI)

Xiaomi open-sources MiMo-V2.6, a trillion-parameter omnimodal model under MIT

Xiaomi released the MiMo-V2.6 family: Pro at about 1.02 trillion parameters, Flash at about 311 billion, and a Pro-UltraSpeed variant Xiaomi says produces output up to 20x faster at the same quality. The models are natively omnimodal (text, image, video, audio), the weights are on Hugging Face under an MIT license, and Xiaomi claims Pro scores 46.32 on the Artificial Analysis Intelligence Index, ahead of Kimi K3 and Qwen3.8 Max at launch. Xiaomi's hosted API prices Flash at $0.14 input and $0.28 output per million tokens and Pro at $0.435 and $0.87. There's also a 9B distill built on Qwen3.5-9B with GGUF quantizations already up, so it runs on a laptop. Xiaomi's benchmark numbers are its own; independent scores weren't posted at press time.

Why it matters: MIT is as permissive as it gets. If you need a self-hosted model with no commercial strings, this is now the strongest open option by the vendor's own numbers, and the 9B distill is the one to try first. Watch for Artificial Analysis to confirm the Index score before you bet a product on it.

Source: MiMo-V2.6 (Xiaomi) and XiaomiMiMo/MiMo-V2.6-Pro-RL (Hugging Face)


Tools and platforms

Muse hits No. 1, gets PayPal and Shopify checkout, and Meta admits it copied OpenClaw's homework

Meta's Muse agent, launched September 8, reached No. 1 on Apple's US App Store and, per Apptopia estimates cited by TheLec, pulled 2.8 million downloads across iOS and Google Play in the US and Canada in its first 12 days, ahead of ChatGPT's launch pace on iOS. Those are third-party estimates, not Meta's numbers. On the commerce side, Shopify said Shop Pay will be enabled for Muse on every Shopify store, and on Tuesday Meta's chief AI officer Alexandr Wang announced PayPal checkout across PayPal merchants worldwide. Stripe's Link was in at launch; Amazon remains blocked. Separately, TechCrunch reported that Nat Friedman, head of product at Meta Superintelligence Labs, acknowledged Muse was "heavily inspired as a product by OpenClaw," the open-source assistant whose creator OpenAI hired in February, after users found identical workspace filenames and a near-identical SOUL.md config.

Why it matters: If you sell on Shopify or take PayPal, an AI agent can now buy from you without you doing anything. That's distribution you didn't build and can't control. And the OpenClaw admission is a reminder that open-source personal agents are the reference design now; the big labs are packaging, not inventing.

Sources: Meta admits Muse's likeness to OpenClaw isn't a coincidence (TechCrunch) and Shopify to Use Meta's Muse for Agentic Checkout (WSJ)


Business and money

Alibaba's Apsara roadmap: a 10-trillion-parameter Qwen, a homegrown chip, and 20 gigawatts

At its Apsara Conference in Hangzhou, Alibaba CEO Eddie Wu said Qwen 4 is in training and that Qwen 4.5 and Qwen 5 are targeted at 5 to 10 trillion parameters. Alibaba's T-Head unit unveiled the Zhenwu V900 accelerator: 216 GB of memory, 1,200 GB/s inter-chip bandwidth, native FP8 and FP4, three times the performance of its predecessor by Alibaba's account, clusters up to 500,000 cards, and mass production in Q1 2027. The company also set a target of 20 GW of global data center capacity by 2032 and said Qwen 3.8-Max ran 33 "recursive self-improvement" cycles over a month that moved its Artificial Analysis score from 40 to 45. Wu's framing: "the total volume of Machine Thinking is less than 3 percent of all Human Thinking." All performance figures are Alibaba's own.

Why it matters: This is a full-stack bet, chips to agents, aimed at running without Nvidia. If you're building for Asian markets or on Alibaba Cloud, the model roadmap and the Agent Context service (Alibaba claims a 67 percent token reduction in knowledge-heavy tasks) are the near-term pieces. The RSI claim is the one to watch, given the Google paper we covered yesterday on putting limits on self-rewriting agents.

Source: Alibaba Unveils Roadmap on Full-Stack AI Strategy (Alibaba Cloud)


Policy and risk

UK Parliament calls in Meta, Google, OpenAI, and Anthropic on AI security

The House of Commons Business, Innovation, Science and Trade Committee published letters inviting all four companies to an evidence session on October 13, alongside the UK's AI Security Institute. The questions it flagged are pointed: whether pre-release safety testing should be legally mandatory rather than left to the companies, whether firms will accept mandatory reporting of serious incidents including "deceptive behaviour, safeguards being circumvented, unauthorised replication and evidence that human control may be failing," and what stops competition between companies or countries from becoming a race to the bottom. It's an invitation, not a summons, and the session feeds the committee's broader work on the UK's economic strategy for AI.

Why it matters: Mandatory incident reporting is the policy idea gaining ground fastest on both sides of the Atlantic. If it lands, the reporting burden won't stop at frontier labs; expect it to flow down to anyone deploying agents in regulated sectors. The October 13 transcript will be worth reading.

Source: Meta, Google, OpenAI and Anthropic invited to appear before Business Committee (UK Parliament)


Broader tech

New Jersey fines a data center $1.07 million for 62 unpermitted gas generators

The New Jersey DEP fined the DataOne data center in Vineland $1.07 million under the state's Air Pollution Control Act for installing and running 62 natural gas generators rated at 1,982 kW each, roughly 123 MW combined, without preconstruction permits or operating certificates. Inspectors found the generators on July 29; they weren't there at the previous inspection in December 2025. The order tells DataOne to get permits immediately or cease operations. Commissioner Ed Potosnak called it "by far the largest ever taken against a data center in New Jersey and possibly one of the largest such actions in the nation." It follows a July law requiring data centers to bring their own clean energy and an August law requiring semiannual water and energy reports.

Why it matters: The grid is the bottleneck, and operators are routing around it with on-site gas. States are starting to punish that. If your product depends on a specific region's compute, permitting risk is now a real input to your uptime and cost assumptions.

Source: Sherrill Administration Issues Record $1 Million Fine to DataOne Data Center (NJ DEP)


What I'd do with this

Re-price everything today. Pull last month's token usage by model and rerun it against the new Opus 5.5 and GPT-6 Sol/Luna rate cards, with cache reads broken out separately. Most solo builders will find their cheapest path just changed. Then pick one narrow workload (classification, extraction, routing) and run a three-way bake-off: Luna, MiMo-V2.6 Flash via Xiaomi's API, and the 9B distill locally. Ship whichever passes your eval at the lowest cost per correct answer, and write down the threshold you'd switch at. If you sell anything online, check whether Muse can already buy from you through Shop Pay or PayPal, and decide whether you want it to.