September 22, 2026

AI Next Wave - September 22, 2026

OpenAI hands nine mathematicians an unpaid veto, on the same day its million-dollar proof gets picked apart

A chalkboard full of mathematical equations, softly out of focus, in an empty classroom.

Models and research

OpenAI names an outside math board, and the timing isn't subtle

OpenAI announced an Advisory Group on Mathematics and Artificial Intelligence with nine members, including Timothy Gowers, Edward Witten, Martin Hairer, Ravi Vakil, and Melanie Matchett Wood. The group will "advise on the review and communication of emerging results," and OpenAI says its members "will not be paid by OpenAI, and the group can change its membership as it sees fit," with freedom to give advice OpenAI didn't ask for, comment on OpenAI's impact on mathematics, and publish that advice. The stated trigger is OpenAI's internal model resolving the Navier-Stokes problem plus more than 100 other open problems, and the unease that caused among mathematicians.

Why it matters: This is what a credibility problem looks like when you try to fix it with governance. OpenAI can't verify its own math claims and be believed, so it's renting nine reputations and giving them the right to say no in public. If you sell anything where the buyer can't check your output themselves, that structure is worth stealing: unpaid outside reviewers who can publish without your sign-off are cheaper than a trust crisis.

Source: OpenAI announcement

Mathematicians say OpenAI solved the easier Navier-Stokes

Scientific American laid out the objection the same day. The Clay Mathematics Institute's official 2000 problem statement, written by Charles Fefferman, included an "Option C" that permits an external force acting on the fluid. OpenAI's system solved that version. Most working groups were attacking the version without a force, and three mathematicians have since shown OpenAI's approach can't be extended to it. Luis Silvestre put it bluntly: "The most important problem is unsolved. The Clay problem is settled, but the main problem for the Navier-Stokes equations is not." Not everyone agrees the criticism lands; Diego Córdoba argues that "all fluids we know of are under some kind of external force," so including one is reasonable. Uncertainty flag: the Clay Institute hasn't ruled on the prize, OpenAI didn't respond in the piece, and Gonzalo Cao-Labora's read is that how this gets remembered depends on what happens next.

Why it matters: The benchmark and the thing you care about are not the same object, and a sufficiently motivated system will find the gap. That's the whole lesson, and it applies to your evals too. When you tell a model to hit a number, check what the number is actually measuring before you celebrate.

Source: Scientific American

Google researchers put a leash on agents that rewrite themselves

A team including Google Cloud AI researchers posted RRSI, short for Regularized Recursive Self-Improvement of Agent Harnesses. The premise: agent performance comes mostly from the harness around the model, meaning the prompts, control flow, tooling, and memory, and letting a system iteratively rewrite its own harness produces big gains that turn out to be memorization. In-distribution scores jump, out-of-distribution scores don't follow. RRSI constrains how edits get proposed and selected, using an annealed edit budget plus a critic and a pruner that throws out changes that are too small, too expensive, or no longer useful. Across eight benchmarks it gained up to 14.1 points in-distribution, up to 4.7 points on five out-of-distribution benchmarks, and ran on 30 percent fewer policy tokens. This is an arXiv preprint, not peer reviewed, and the code is posted at github.com/google-research/rrsi.

Why it matters: If you've been letting an agent tune its own prompts against your eval set and watching the score climb, this paper is telling you what you probably built. The 30 percent token reduction is the part with a dollar sign on it, and the pruner is the idea you can copy this week without adopting the whole framework.

Source: arXiv:2609.24972


Tools and platforms

A chunking method that cuts RAG ingestion cost by three quarters

Four researchers at Yellow.ai published D-RAC, or Document Retrieval-Aware Chunking. The approach normalizes any input format into PDF, does a single multimodal model pass to convert rendered pages into retrieval-optimized Markdown that rewrites tables as self-contained prose and keeps heading hierarchy, then chunks over identifiers instead of regenerating source text. On a 236-document, 795-page benchmark across five enterprise domains it processed the whole corpus in 72 minutes with zero errors and 1,748 chunks. Against agentic chunking with frontier models it cut chunking-stage output tokens by 95.7 percent, cost by 77.8 percent at GPT-4.1 pricing and 85.6 percent at Gemini 2.5 Pro pricing, and time by 75 percent. Uncertainty flag: this is a vendor's own paper on its own benchmark subset, not peer reviewed, and the pricing comparisons use list prices for two specific models.

Why it matters: Ingestion is the line item people forget when they price a RAG product, and it's the one that scales with the customer's document pile instead of their usage. Never regenerating source text during chunking is the trick here, and it's a design decision you can apply to whatever pipeline you already run. If you're eating ingestion cost to win enterprise deals, this is a margin conversation.

Source: arXiv:2609.24220

OpenAI Academy expands to 14 courses across four tracks

OpenAI added role-based learning paths to its Academy: Apply AI at Work (3 courses), Build with AI (8), Lead AI Adoption (1), and Teach and Learn with AI (2), aimed at knowledge workers, developers, organizational leaders, and educators and college students. Learners practice on real tasks and earn Academy badges by passing assessments. The company's framing is that "you should use AI to learn AI." OpenAI didn't state pricing or regional availability in the announcement.

Why it matters: Badges from a model provider aren't a credential anyone should over-weight, but eight free-looking courses on building with the API is a real onboarding ramp for a solo founder or a career-switcher who can't expense a bootcamp. Worth an hour to see whether the Build with AI track covers what you'd otherwise spend a weekend piecing together from docs.

Source: OpenAI announcement


Business and money

Helsinki's Verda raises $189 million and becomes Europe's newest AI cloud unicorn

Verda, formerly DataCrunch, announced a $189 million (€163 million) round led by Emergence Capital, bringing total funding past $450 million. Other investors include MUFG Innovation Partners, Supermicro, Varma, Lifeline Ventures, 6 Degrees Capital, byFounders, and Tesi. The company says the money goes to compute capacity and platform services with a specific push into inference, and that it plans to multiply capacity over the next year. Founded in 2020, roughly 250 employees, operating across Europe, the US, and Asia. Uncertainty flag: the $1 billion-plus valuation and the $165 million annualized revenue run rate as of July are the company's own figures, and the press release doesn't name an exact valuation.

Why it matters: European teams with data residency requirements now have another full-stack option that isn't a US hyperscaler, and the explicit inference focus means this is aimed at people serving traffic, not just training runs. If your customers ask where their tokens get processed, the list of credible answers just got one entry longer. Founder Ruben Bryon's stated goal, "to build the first true tech company in Europe," is ambition talking, but the $450 million behind it isn't.

Source: Verda press release


Broader tech

NVIDIA starts certifying batteries and coolant pumps

NVIDIA launched DSX Ready, a qualification program for power and cooling products that meet its DSX AI factory reference design requirements. Two categories at launch: battery energy storage systems, with Hitachi Energy, LG Energy Solution, and Tesla qualified, and coolant distribution units, with LG Electronics, LiquidStack, and Vertiv. BESS partners run required tests and submit data for NVIDIA review; CDU providers use a self-qualification suite. The CDU checks cover hydraulic, thermal, and control requirements including constant differential-pressure performance, flow-sensor accuracy, cold-start behavior, pump failover, and group control. NVIDIA's stated reasoning is that "optimizing one part of an AI factory can shift the bottleneck elsewhere."

Why it matters: The constraint on AI capacity stopped being chips a while ago and is now power, cooling, water, and grid interconnect. NVIDIA writing the compatibility spec for batteries and pumps means it's extending its platform control one layer down into physical infrastructure, the same way it did with networking. For everyone else, this is the clearest signal yet that inference prices are tied to electricity and thermal engineering, not just silicon supply.

Source: NVIDIA blog


What I'd do with this

  • Audit one eval you trust. Ask what it actually measures versus what you want, the way Option C measured something narrower than the field wanted. If a model could score well while failing your real goal, you have a Navier-Stokes problem in miniature.
  • If you run any self-improving loop, prompt optimization included, hold out a benchmark it never sees and check the gap. RRSI found gains that shrank or vanished out of distribution. That's the default outcome, not the edge case.
  • Price your RAG ingestion separately from your inference. D-RAC's numbers suggest chunking is often the dominant cost at scale, and it's the one that grows with the customer's archive rather than their engagement.
  • Ask your cloud vendor where inference physically runs and on whose power. With Verda in the market and NVIDIA certifying grid hardware, "which region" is becoming a real procurement answer instead of a checkbox.
  • Borrow OpenAI's structure, not its claims. One outside reviewer who can publish without your approval does more for buyer trust than any amount of self-reported benchmarking.
  • Spend an hour in the Build with AI track before you buy any course. Free and mediocre beats paid and mediocre, and you'll find out fast which it is.