September 18, 2026
AI Is Building Its Successor and Moving Into the Lab
Anthropic says Claude now leads 26% of its model research and development while the company builds a physical biology operation. A disputed model on a federal website, internal copyright evidence, camera glasses and Toyota's 400,000-robot estimate show what happens when software gains a longer reach.

The Short Version
Anthropic has put two forms of acceleration next to each other.
The company says Claude led 26% of its model research and development work in August, up from about 1% in March. In Anthropic's definition, leading means completing most of a task from a high-level prompt while a person supervises. Claude collaborated on over 90% of the company's model research and development. It operated neither independently nor without controls.
Anthropic has also built a wet lab in the San Francisco Bay Area. Its life-sciences group is doing physical biology work in-house and with outside partners. The company wants Claude to direct robotic laboratory equipment with limited intervention, according to a person familiar with the work, while keeping people involved for safety. Anthropic says it plans preclinical programs in areas that drug companies tend to overlook and has drawn a boundary before clinical trials.
The two developments create a consequential loop. AI helps build a stronger AI. A stronger AI can propose and run more physical experiments. Results from the lab can shape the next hypothesis, tool and model. Each handoff can shorten a research cycle. It can also carry an error from software into biological material or from an experiment back into the system that chose it.
The public evidence remains early and company-reported. Anthropic has disclosed no drug candidate, disease program, validated discovery, lab throughput, cost reduction or patient outcome from the new operation. Most drug programs fail before approval. A wet lab proves physical capability and intent. It provides no clinical return yet.
Four other developments published Thursday and Friday show the same expansion into operating systems people already depend on.
The Federal Register briefly offered an AI search tool using Alibaba's Qwen model for browsing public comments. The tool disappeared Wednesday as posts about it spread. Reuters found no immediate security issue, and experts noted that the underlying material was public. The National Archives did not explain who approved the model, where queries were processed or what evaluation preceded deployment.
News organizations have made previously redacted evidence public in their copyright case against OpenAI and Microsoft. Their filing cites internal statements describing generative AI as substitutive for publisher sites and acknowledges the economic threat to creators. These are plaintiffs' arguments and selected evidence, not a court finding. They sharpen the commercial question around products trained on, and designed to answer from, professional work.
Paris prosecutors have opened a criminal investigation following complaints about people using smart glasses to film women without consent and post the videos online. France's privacy regulator has received fewer than ten workplace complaints and says businesses increasingly ask whether the devices can be banned. The brand involved in the criminal complaints was not disclosed.
Toyota estimates that modernizing factories across the company, group businesses and major suppliers could cost 1 trillion yen, about $6.4 billion, a year from 2028. It told investors that roughly 400,000 robots could be needed, including replacements, new equipment, automated logistics and future human-robot collaboration. Toyota gave no commitment that the full spending will occur or how many years it might continue.
Software is acquiring a body, a worksite and a route into public institutions. The useful control now has to follow the action all the way to the lab result, government answer, reader, bystander and factory worker.
The Research Loop Reaches the Lab Bench
Anthropic is measuring how much AI builds AI while giving it physical experiments to run
Anthropic's 26% figure sounds like a measure of autonomous research. The company's definition is narrower and more useful.
A Claude-led task begins with a high-level prompt and ends with the model completing most of the work while a person supervises. Collaboration covers a wider range in which Claude performs large portions under close human direction. Anthropic says no measured part of its research runs fully autonomously.
The share still moved quickly. Claude led almost none of the company's model research and development in February, about 1% in March and 26% in August. Over 90% of the work involved some form of human-AI collaboration by August. The company says it will publish the measure regularly and has urged other frontier labs to use public methods that make progress easier to compare.
These are self-reported workflow classifications. They show where work sits on Anthropic's scale, not whether Claude produced 26% of the intellectual value, saved 26% of the cost or improved the next model by 26%. A task can be large or small. A human may spend little time assigning it and significant time verifying it. Failed paths and correction work can disappear inside a completion label.
Anthropic disclosed operating details that help define the denominator. About 30,000 agents were active on its main internal research platform at any one time in August. The company says every proposed action was screened and roughly one in 47,000 decisions was blocked across over one billion decisions that month. In one sampled week in July, it assigned about 6% of research computing to safety work and 12% of the computing used for AI-led research.
Those figures describe Anthropic's controls and accounting. They lack independent reproduction. A low block rate can indicate that agents usually behave within policy, that the controls rarely trigger or that the policy misses important behavior. A safety-compute share shows resource allocation, while safety performance requires evidence from incidents, interventions and releases changed.
The wet lab makes the classification more consequential. Anthropic life-sciences chief Eric Kauderer-Abrams confirmed that the company is performing physical biology work in its own facility and through partners. One person familiar with the program told Reuters that Anthropic wants Claude to direct robotic units through experiments with limited human intervention. The company says people remain essential to safety.
Anthropic has said it wants to pursue preclinical work on conditions that traditional drug developers find financially unattractive. It launched Claude Science, bought Coefficient Bio and released a Model Hardware Standard designed to help AI operate laboratory equipment. Its lab also gives the model maker direct experience with the messy realities that disappear in a software simulation: calibration drift, contaminated samples, failed assays, ambiguous results and equipment that responds differently from its specification.
The boundary is currently preclinical. Anthropic says it isn't running clinical trials and doesn't plan to compete with companies that bring drugs to market. It has not identified the diseases under study or disclosed progress toward a molecule. A biological result still needs replication, toxicology, manufacturing, regulatory review and human trials before becoming a treatment.
That long route changes the sensible return calculation. Faster hypothesis generation creates options. More experiments per week create capacity. A higher share of reproducible results improves quality. Lower cost per validated target or candidate is operating efficiency. A licensed molecule or approved treatment can create revenue and patient benefit. None of those outcomes should be collapsed into a claim that AI made drug discovery ten times faster.
A laboratory considering similar automation should choose one bounded protocol with known controls. Preserve the hypothesis, model and prompt version, source evidence, approved protocol, robot commands, materials, deviations, observations and human interventions. Run blinded replicates against the current process. Keep the system away from organisms, reagents and procedures outside its authorization.
The scorecard should include successful runs, reproducibility, scientist review time, contamination, protocol deviations, equipment downtime, cost per accepted result and time to the next decision. A faster run that creates a doubtful result adds review debt. A slower system that makes the experiment easier to reproduce may create greater scientific value.
Anthropic's disclosures show a company shortening both sides of the research loop. Claude is taking a larger role in building Claude. Physical experiments give its scientific work contact with reality. The operating challenge is to keep evidence, authority and containment intact as the loop tightens.
A Public AI Tool Appeared Without a Public Decision Record
The Federal Register episode turns model procurement into a basic transparency test
The Federal Register is where the U.S. government publishes proposed rules, final rules and public notices. Its website briefly offered people an AI search option powered by Alibaba's Qwen model for browsing public comments.
Reuters found the option on Wednesday and reviewed screenshots and archived source code. The Qwen search disappeared around the time social-media posts drew attention to it. The National Archives, which operates the Federal Register, did not answer Reuters' questions about when the tool arrived or why it was removed.
The model choice carried political baggage. The FBI had accused Alibaba the previous week of copying Anthropic technology on an industrial scale. Alibaba and the Chinese embassy rejected the accusation. Qwen is an open-weight model, which means an organization can download key model components and run a version on its own infrastructure.
That deployment detail decides much of the actual risk. A locally hosted model searching public comments may send no query or document to Alibaba. An externally hosted service could expose user questions, usage patterns and system instructions. Reuters' experts saw no immediate cybersecurity risk from the known facts because the Federal Register content was already public. They could not determine whether data left the government's boundary.
The episode therefore supports no claim that Alibaba accessed federal information or that Qwen compromised the site. It does expose a missing public record. People could see a model name. They couldn't see the host, data route, evaluation, retention, accessibility test, error process or official responsible for the result.
That matters because regulatory search can shape participation. A small business, nonprofit or citizen may use the tool to find a proposal that affects licensing, benefits, workplace rules or costs. A confident answer that omits a relevant notice can narrow who submits a comment. A summary that blends a proposal with a final rule can change a person's understanding of current law.
Public agencies need a short model card for the deployed service, written for users rather than vendors. Name the model and version, hosting environment, information sent outside the agency, retention period, covered records, known limitations, citation method, evaluation date, responsible office and route for correcting an answer. Preserve a conventional search path beside it.
The pilot should use questions gathered from actual public-service work. Compare the AI results with the current search system and reviewed answers. Measure relevant records found, material omissions, citation accuracy, user completion, accessibility, review time, query exposure, corrections and cost per successful search.
A cheaper open model can be a sound public-sector choice. Local operation can improve custody and reduce recurring fees. The value disappears when nobody can show how the choice was made or how an affected person can verify the answer.
Internal Evidence Meets the Copyright Case
Publishers are using the AI companies' own words to argue that substitution was the plan
The copyright fight between news organizations, OpenAI and Microsoft has moved from public principles into internal documents and sworn testimony.
A combined summary-judgment brief filed Thursday by the news plaintiffs quotes OpenAI and Microsoft employees describing generative AI products as substitutes for publisher sites. The filing cites OpenAI's head of ChatGPT saying publishers faced an existential threat and that the products were largely substitutive. It cites Microsoft chief executive Satya Nadella agreeing in testimony that chatbot conversations can replace a visit to the underlying website.
The filing also alleges that an OpenAI employee shared a method for getting around The New York Times paywall while scraping, and that OpenAI president Greg Brockman responded positively. Microsoft told Reuters that Nadella's comments concerned broad changes in information consumption and remain consistent with the company's legal position. It said a separate employee's criticism of AI training reflected an individual view rather than company policy. OpenAI did not immediately comment to Reuters.
Every legal qualifier matters. This is evidence selected and characterized by the plaintiffs in support of summary judgment. OpenAI and Microsoft dispute the copyright claims and argue that training is transformative fair use. The judge has issued no ruling on the new filing, the quoted statements or liability.
The evidence still changes the commercial conversation. Fair use analysis considers, among other factors, the purpose of the use and its effect on the market for the original work. A product designed to replace visits, trained on the work it replaces and supported by internal recognition of that substitution gives publishers a more concrete argument than a general objection to machine learning.
The filing says Microsoft's data showed sharply lower click-through rates from Bing Chat than from conventional Bing search for the plaintiff publishers. Those figures appear in an adversarial court brief and have yet to be tested through a ruling. They point toward a measurable product choice: whether an answer sends a person to the source, pays for the source or captures the value while leaving the source with neither traffic nor compensation.
Creators, publishers and research businesses should track which uses produce revenue or saved labor for the AI product and which uses return value to the source. Citation presence alone is weak evidence. Useful measures include referral visits, licensed queries, content used under contract, subscriber conversion, revenue shared, removal requests, repeat infringement and the cost of producing the underlying work.
AI product teams can reduce uncertainty by separating licensed corpora, public-domain material, customer-owned data and material collected under another legal theory. Keep acquisition source, rights status and permitted uses attached to the data. Test whether answers reproduce protected expression, satisfy a query without a visit or falsely attribute work.
The outcome of this case may reshape training, retrieval and licensing across the industry. The immediate lesson requires no prediction about the judge. Internal strategy statements become evidence when a product's economic design reaches court.
Camera Glasses Turn Every Workplace Into a Recording Policy
France's investigation shows how a tiny indicator light can carry too much responsibility
Paris prosecutors have opened at least one criminal investigation after complaints linked to a social-media trend in which people used smart glasses to film women on the street without consent and post the recordings online.
The prosecutors did not identify the brand involved. Their cybercrime unit tested smart glasses two weeks ago to better understand potential offenses. An investigation establishes scrutiny, not guilt or a completed prosecution.
Meta's AI-enabled glasses, produced with EssilorLuxottica, lead the global category with an estimated 76% share. Reuters reports that seven million units sold last year. The products combine an ordinary eyewear shape with a camera, microphones, speakers and AI services. A small light signals recording.
The form factor creates the social problem. People recognize a raised phone as a camera. Glasses stay pointed wherever the wearer looks. A bystander may miss the light, misunderstand it or see the recording only after the interaction has begun. A visible indicator communicates activity. It does not create consent to capture or distribute a person's image.
France's data protection authority, CNIL, told Reuters that it had received fewer than ten workplace complaints involving smart glasses. It has also received more questions from businesses asking whether they can prohibit the devices at work. Australia is considering a bar on camera-equipped glasses in government workplaces because of privacy and security concerns.
Employers have several interests to reconcile. A technician can use glasses for remote assistance, hands-free documentation or accessibility. The same device can record a customer, whiteboard, patient, production process or colleague without a clear shared moment of consent. A universal ban can remove a useful tool. An informal rule leaves workers and bystanders guessing.
A workable policy starts with location and purpose. Define approved devices, tasks, recording zones, notice, consent, storage, retention, access and deletion. Separate live assistance from recording. Mark spaces where glasses must be removed or physically covered. Provide a non-recording alternative for employees who need vision correction or accessibility support.
Test the policy during an ordinary shift. Ask whether a visitor can identify recording, whether the wearer can confirm storage, whether a supervisor can retrieve and delete a file, and whether an employee can report suspected misuse without confronting the wearer. Track authorized sessions, complaints, unapproved capture, deletion time, security incidents and useful work completed.
The camera is becoming wearable enough to disappear into a normal interaction. Product safeguards and workplace rules have to make the recording legible again.
Toyota Put a Price on Factory Automation
The 400,000-robot estimate is a planning range, with the workforce design still open
Toyota told investors that modernizing factories across the automaker, its group companies and major suppliers could require roughly 400,000 robots and annual spending of 1 trillion yen, about $6.4 billion, from 2028.
The estimate includes replacements for existing machines and new installations. It covers industrial robots, automated logistics, humanoid and non-humanoid systems, and future collaboration between people and machines.
Toyota did not commit to the full investment, specify how many years the annual spending might continue or divide the 400,000 units between replacement and added capacity. The number describes a possible buildout, not a purchase order or deployed fleet.
The scale still reveals what factory automation requires after the demonstration. A useful robot needs tooling, safety systems, floor space, power, maintenance, spare parts, software integration and a process stable enough to automate. Suppliers need capital and technical staff. Workers need new routines around setup, exception handling, quality and recovery.
Labor shortages and aging equipment create a credible reason to invest. A robot can take repetitive lifting, inspection or movement from a hard-to-fill job. It can also create downtime when a rare fault stops a connected line. Humanoid form can help in spaces designed for people, while a simpler fixed machine may perform one task more reliably and cheaply.
The pilot belongs on a complete production cell. Establish current throughput, defects, changeover time, injuries, overtime, maintenance and unplanned stops. Add the robot, including the people who program, supervise, repair and feed it. Record every exception that sends the work back to a person.
Return should be separated by type. Lower paid hours are labor savings. More output from the same shift is capacity. Fewer defects improve quality and avoid rework. Reduced lifting and exposure lower safety risk. Faster changeovers improve flexibility. Training expenses, integration delays and supplier financing belong in the cost.
For workers, the transition becomes credible when the company names the tasks that leave, the tasks that remain and the qualifications attached to the new work. A maintenance technician, robot integrator, safety specialist or process analyst may gain leverage. A supplier without the capital or staff to automate may lose business even when its current quality is strong.
Toyota has supplied an unusually large planning number. The valuable evidence will come from cells that produce safely, suppliers that can finance the change and workers who can move into the jobs the automated system still needs.
Opportunity Radar
Audit trails for physical AI
Laboratories and factories are connecting models to robots faster than their existing records can explain a machine's decision. A lab-information provider, industrial integrator or assurance firm could build an action-complete record for one protocol or production cell.
The service would connect the authorizing person, model and software version, source data, approved procedure, machine command, material, sensor reading, deviation, intervention and final result. Biotech companies, contract research organizations, manufacturers and insurers could pay for faster investigation, reproducibility and evidence for a customer or regulator.
The provider must validate timestamp integrity, equipment compatibility, sensitive-data controls and recovery when a vendor system goes offline. Time to reconstruct a run, unexplained deviations, repeatability, containment time, review effort and accepted results form the scorecard. A detailed log that scientists and operators cannot interpret would add storage without accountability.
Public AI deployment records
Small agencies, schools, courts and local governments are adopting models without teams that can document hosting, data movement, evaluation and public recourse.
A civic-technology firm, privacy engineer or public-interest university center could prepare a deployment record and test one citizen-facing tool. The package would map data flows, evaluate representative questions, verify citations and accessibility, define retention and publish a plain-language notice with a correction route.
Agencies could pay for procurement evidence they can show residents, auditors and oversight bodies. The service has to remain independent of the model seller and prove that its test reflects the people and decisions in scope. Material errors, records found, query exposure, corrections, completion time, user trust and operating cost should decide whether the tool stays.
What You Can Do With This
If AI can operate equipment
Choose one bounded protocol or cell and preserve the complete path from instruction to physical result. Run known controls, blind the evaluation where possible and record every intervention. Measure reproducibility, exceptions, review work and cost per accepted result before increasing autonomy.
If you buy AI for public or regulated work
Require the model version, host, data route, retention, evaluation, limitations and responsible owner in the approval record. Give users a conventional route and a way to challenge the answer. Recheck the deployment whenever the model or provider changes.
If your work supplies an AI answer
Track the value that returns to the source. Separate citations from visits, licenses, subscriptions and revenue. Keep rights and acquisition records with training or retrieval data, and test whether the product satisfies demand by substituting for the work that made the answer possible.
If cameras or robots are entering the workplace
Define approved tasks, zones, signals, storage, safety boundaries and reporting before rollout. Include employees and affected bystanders in the test. Count useful work beside complaints, exceptions, incidents, downtime and the new labor required around the system.
The Bigger Picture
The most consequential AI development this week may be the shortening distance between a prompt and a physical or institutional effect.
Anthropic's research agents can take a high-level assignment through much of the work needed to build their successors. Its life-sciences operation can carry an AI-directed idea into a physical experiment. The Federal Register episode placed a model inside the public route to proposed rules. Smart glasses place cameras and AI in a person's line of sight. Toyota's estimate puts automation across a supplier network and production floor.
The copyright filing shows the economic version of the same movement. An answer can travel directly to the user while the organization that produced the underlying reporting loses the visit. The technology completes the task by crossing a boundary that used to send a person elsewhere.
Each crossing needs its own evidence. A lab needs a reproducible protocol and chain of custody. A public agency needs a deployment record. A creator needs rights and compensation attached to use. A bystander needs an intelligible signal and recourse. A worker needs a task and transition plan grounded in the actual cell.
Speed remains useful when it shortens discovery, finds a regulation, helps a technician or removes dangerous lifting. It becomes expensive when the action outruns the record needed to verify, reverse or compensate it.
That is where the next wave of practical work sits. The model proposes. The surrounding system decides what it may touch, preserves what happened and assigns responsibility for the result.
AI Next Wave