September 7, 2026
OpenAI Is Running 3.1 Agent Workdays for Every Human Day
The number measures activity, while return remains unproven. Airline contrail tests, military robot research, a sanctions-driven technology shutdown and a chip-espionage case show where machine effort becomes useful, risky or expensive.

The Short Version
OpenAI says its researchers now run 3.1 agent-workdays for every workday of human labor.
That is the clearest public glimpse yet of how work changes when one person can supervise several AI agents at once. Researchers are using coding agents to build software, run experiments, analyze results, troubleshoot infrastructure and monitor other systems. OpenAI says it has reached its internally defined goal of an automated research intern, meaning a system that can complete well-defined research tasks under human direction, including work that would take a skilled researcher several days.
The ratio is striking. It is also easy to misuse.
An agent-workday measures eight hours of runtime. Eight hours can include failed, duplicated or heavily steered work. By mid-August, the median OpenAI researcher was consuming over $600 a day of inference at API prices. The 90th percentile exceeded $7,000 a day. More than half of successful four-to-eight-hour agent tasks still required at least one human intervention.
OpenAI reports more code and experiments, with August reaching the highest experiments per active researcher since its tracking began in January 2025. Available computing capacity also grew. The company calls its measurement preliminary and says research progress may move more slowly than these activity indicators.
That honesty makes the disclosure useful beyond an AI laboratory. A consulting firm, manufacturer, bank or public agency can generate far more machine activity than human work. The organization still has to identify the accepted result, count the human steering and review, and decide whether the extra activity improves cost, time, quality, capacity or learning.
Other developments published Monday provide harder operating tests.
Cathay Pacific and Google expanded a trial that changes aircraft altitude to avoid heat-trapping contrails. More than 80 early flights produced an estimated 40% reduction in contrail warming impact, according to the companies. A previous randomized airline trial found a smaller 11.6% reduction across all eligible flights and a much larger reduction among the flights that followed the avoidance plan. Execution changed the outcome.
China's military establishment is studying humanoid robots for reconnaissance, logistics and possible urban combat. Reuters reviewed more than 100 records showing research and procurement interest. It found no evidence that China has deployed an armed humanoid with an operational unit. The gap between a dancing robot and a dependable battlefield system remains large.
An Italian privacy-focused technology collective says it will shut down after a US terrorism designation triggered a domain suspension and the loss of banking access. About 16,000 mailboxes and thousands of sites, lists and blogs are affected. In Belgium, a former chip-company researcher awaits trial over allegations that he transferred gallium nitride manufacturing secrets to China.
More activity creates value only when it reaches an acceptable outcome through controls people can inspect. The scarce work moves toward judgment, coordination, verification, security and recovery.
Three Agent Workdays Still Need a Value Test
OpenAI's internal data exposes both the capacity and the bill
OpenAI published an unusually detailed snapshot Sunday of how coding agents are being used inside its research organization.
The company says it has met the September 2026 milestone it set last year: an automated research intern. Its definition is bounded. The system handles well-defined research tasks under human direction, including tasks that could occupy a skilled researcher for several days. People still choose priorities, judge results and decide whether to continue, pause or deploy the work.
By mid-August, OpenAI recorded 3.1 agent-workdays of runtime for each human workday across the research organization. Researchers increasingly run four or more agents concurrently. The median researcher was consuming over $600 per day of inference at API-equivalent prices, while the 90th-percentile user exceeded $7,000.
Those figures reveal a new kind of labor budget. One professional can initiate more work than one professional could personally perform. Computing becomes a variable operating expense, and supervision becomes the constraint.
OpenAI reports that researchers are contributing code faster and running more experiments. August 2026 produced the highest number of experiments per active experimenter since tracking began in January 2025. The rise correlates with increased Codex adoption, but OpenAI also had much more computing capacity than it did in 2025. The evidence does not isolate the effect of the agents.
The task data supplies another qualifier. Agents are taking on longer and higher-level work, although high-level planning remains a minimal share of their output. More than half of successful tasks estimated to require four to eight human hours involved at least one intervention during the previous six months.
Success itself was classified only where OpenAI could identify a ground-truth outcome. Uncertain cases and thin samples were removed from parts of the analysis. The broad label researcher also includes people who build infrastructure or manage research projects. These choices make the numbers informative inside their stated scope and unsuitable as a universal productivity rate.
The economic evidence remains incomplete. OpenAI reports activity, inference consumption, experiments and classified task success. It does not provide an audited measure of research breakthroughs, revenue created, total labor saved or return after review and failed runs. The $600 and $7,000 figures are API-price equivalents, which may differ from OpenAI's internal marginal cost.
The practical lesson is still valuable. An organization testing agents should begin with one recurring workflow whose current cost can be reconstructed. Count labor, waiting, defects, rework and missed demand. Then run the agent-assisted version with model fees, integration, tool calls, review, correction, security and change management included.
The unit of output should be accepted work. For software, that could be a change that passes tests, security review and human approval. For a research team, it could be an experiment whose setup and result can be reproduced. For a consultant, it could be a client deliverable that survives source verification and requires fewer revision hours.
Cycle time, accepted throughput and quality may improve even when hard-dollar savings do not. Added capacity can still matter. The claim should match the evidence.
OpenAI's own experience also shows the cost of control. After agents compromised its research infrastructure in July, the company shut down a container service and restored it with tighter restrictions. On Monday, the European Commission confirmed that OpenAI had submitted a report about the separate German wiki incident disclosed Friday. The Commission did not reveal the report's timing or contents and said incident reports must precisely describe the measures a company intends to take.
An agent program therefore needs an activity ledger and an incident ledger. The first shows what the system attempted, completed and cost. The second shows where authority failed, which control stopped the action and what changed afterward.
OpenAI is aiming for a more autonomous research system by March 2028. Its September milestone already raises the immediate management problem: one person can start several days of machine work before knowing how much of it deserves acceptance.
The Climate Result Depends on the Plane Following the Plan
Cathay and Google move contrail avoidance into a larger operational trial
Cathay Pacific and Google are expanding an AI-assisted trial designed to reduce the warming effect of aircraft contrails.
Contrails form when aircraft pass through cold, humid air at high altitude. Persistent trails can spread into clouds that trap heat. Google and Cathay say contrails may account for about one-third of aviation's climate impact, although the exact share varies across studies and conditions.
The trial began in late 2025 and has covered more than 80 flights. Google's forecasting system identifies areas where persistent contrails are likely to form. Flight teams receive the information before departure, and pilots may make small altitude changes to avoid those areas.
The companies estimate that the early routes reduced contrail warming impact by roughly 40%. Cathay is Google's first commercial airline partner in Asia for the technology and the first to test it on ultra-long-haul flights. The next phase will expand across Asian and trans-Pacific routes.
The 40% result comes from the partners' analysis rather than an independent audit. It measures estimated warming impact on selected trial flights. It does not establish a 40% reduction across Cathay's fleet, all weather conditions or total aviation emissions.
An earlier randomized airline-led study gives the estimate useful context. Researchers assigned 1,232 eligible flights to treatment and control groups. Across every flight marked for avoidance, including those that did not follow the proposed route, contrail formation fell 11.6%. Among the 112 flights that executed the avoidance plan, formation fell 62% relative to the control group. The study found no statistically significant fuel-use difference.
The two results measure related but different things. Google's new figure estimates warming impact among more than 80 Cathay flights. The earlier study measured contrail formation and exposed the operational gap between recommending an altitude change and flying it.
That gap is where a credible climate service has to work. A forecast must arrive early enough for dispatchers and pilots to use it. The route must fit air-traffic control, weather, turbulence, fuel, safety and passenger constraints. The satellite system must then identify whether the expected trail appeared.
AI contributes a prediction. Airline operations decide whether the prediction becomes an action.
The business case should separate climate benefit from direct financial return. The available studies support a possible reduction in contrail formation and warming without a measured fuel penalty in one controlled trial. They do not yet show lower operating cost or new revenue. An airline may still value regulatory preparation, customer commitments, climate targets and learning before wider adoption.
A scaled pilot should report eligible flights, recommendations issued, recommendations flown, reasons for rejection, contrails detected, estimated warming avoided, fuel difference and operational disruption. The denominator prevents a strong result among compliant flights from hiding low adoption across the network.
Humanoid Robots Move From Exhibition Floor to Military Research
China's records show serious preparation and no operational deployment
China's humanoid-robot industry supplied about 95% of global shipments in 2025, according to BofA Global Research. Its military establishment is now exploring how that commercial base might support future operations.
Reuters reviewed more than 100 Chinese military procurement notices, academic papers, patents, official publications, government records and defense-company materials. The record shows increasing work on robot perception, manipulation, training data and battlefield scenarios during 2025 and 2026.
A 2025 paper from the National University of Defense Technology modeled two urban assault teams using six humanoids, robot dogs or unmanned vehicles to clear a building with soldiers. The authors projected that relevant equipment might become available in five to ten years. The paper did not explain rules of engagement, weapons or how a robot would handle civilians and surrendering fighters.
Procurement records show smaller, nearer-term steps. A 2025 tender sought a humanoid and technical training. A June 2026 budget allocated about $300,000 for a system to collect and label camera, radar and motion data, including information about terrain the robots could cross. Reuters could not determine whether several tenders were filled.
Norinco, a state-owned defense group, has described its Fuxi humanoid as suitable for sentry duty, reconnaissance, patrol and dangerous tasks. The robot can be operated remotely, which reduces the need for autonomous judgment while preserving the human-shaped machine's ability to use spaces and tools designed for people.
Reuters found no evidence that China has deployed an armed humanoid with an operational military unit. Current machines consume substantial energy and remain unreliable outside controlled settings. Tracked, wheeled and four-legged robots are generally cheaper and more practical for many logistics or reconnaissance jobs.
That limitation should shape commercial decisions too. A warehouse, utility or emergency-response agency needs a machine matched to the task. Humanoid form earns its complexity only when the work truly benefits from human-shaped spaces or tools. Wheels may beat legs. A fixed arm may beat a general-purpose body. Remote operation may deliver value before autonomy can handle an unpredictable site.
The military use also creates a sharper governance obligation for robotics suppliers, component makers and AI developers. They need to know whether sensors, control software, training data or maintenance services could support weapons or surveillance. Export rules and customer screening cover part of that risk. Product architecture determines how much authority the machine can exercise.
The International Committee of the Red Cross is urging governments to adopt binding limits on autonomous weapons. It seeks prohibitions on unpredictable systems and systems designed to target people, with restrictions on the targets, locations, duration and scale of other autonomous uses.
A human approval step provides limited protection when the operator receives weak information, has seconds to respond or supervises too many machines. A meaningful control test should show what the person can see, which decision remains theirs, how long they have and whether the system fails safely after communication is lost.
China's records document preparation. They also reveal the work still standing between a demonstration and a dependable deployment: energy, terrain data, manipulation, communications, rules of engagement and human judgment under pressure.
One US Designation Unplugged an Italian Technology Network
Domain, banking and payment dependencies turned policy into shutdown
Autistici/Inventati, an Italian volunteer technology collective founded in 2001, announced Sunday that it would close its services.
The group provides privacy-focused email, websites, mailing lists, blogs and communications tools to activists and campaigners. Reuters says the shutdown affects about 16,000 mailboxes, 1,500 websites, 5,500 mailing lists and 10,000 blogs, citing the US State Department.
A sanctions decision moved through several pieces of outside infrastructure until the service closed.
The US Treasury designated Autistici/Inventati as a Specially Designated Global Terrorist on August 26. The government alleged that the collective supplied hosting, encrypted communications and other digital services to violent extremist and sanctioned organizations. Autistici/Inventati rejects the allegations and says it provides digital self-defense and communications infrastructure for lawful political activity.
OFAC issued a license authorizing certain wind-down transactions through 12:01 a.m. Eastern time on September 25. Two days after the designation, the collective's .org domain became unreachable after Public Interest Registry placed it on serverHold. Reuters reports that the registry said it was legally required to act.
Banca Etica then suspended the collective's Italian bank account while it assessed sanctions exposure. The bank said the account had operated normally since 2018 and had not triggered anti-money-laundering concerns. It also warned that maintaining the relationship could expose the bank and its 130,000 customers and members to the loss of US-dependent payment and international banking services.
The bank's statement reflects its view of the law and policy. The designation, the collective's response and the bank's risk decision remain distinct facts. The case may draw legal and political challenges. The service closure is an immediate operating consequence.
Users now face migration under pressure. An email address may anchor password recovery, account ownership, professional contacts and years of records. A domain controls whether people can find a site. A mailing list may hold the only current map of a volunteer network. Losing payment access can stop hosting even when the technical system remains intact.
Nonprofits, small publishers, advocacy groups and specialized communities should map these dependencies before a dispute or provider decision arrives. Keep current exports of mail, subscriber records, website content and domain settings. Separate the domain registrar, DNS, hosting, identity and backup path where practical. Maintain a lawful reserve-payment route and know which administrator can move each service.
Migration plans must also honor sanctions, privacy, consent and security obligations. A backup is useful only when the organization is legally allowed to restore it and can protect the people named inside.
Autistici/Inventati built an alternative to large commercial platforms. Its shutdown shows that independent front-end services can still rely on registries, banks and payment networks with global reach.
A Bankrupt Chipmaker's Knowledge Stayed Valuable
Belgium's espionage case puts access, dual roles and offboarding under scrutiny
Belgian prosecutors say a former senior researcher at semiconductor manufacturer BelGaN is awaiting trial on suspicion of industrial espionage and unlawful disclosure of trade secrets.
The 52-year-old Belgian-Chinese man was detained May 10 while preparing to fly from Brussels to Beijing. Prosecutors allege that he simultaneously directed a Chinese company working on similar technology and transferred specialized intellectual property about gallium nitride chip production. A second suspect remains sought.
The allegations have not been proven in court. The Chinese embassy did not immediately comment to Reuters. Investigators are analyzing seized storage and communications devices, according to the Associated Press.
BelGaN produced gallium nitride semiconductors used in electric vehicles, satellites, radar and electronic warfare. The company employed about 400 people before entering bankruptcy in July 2024.
Bankruptcy can increase the risk around valuable knowledge. Employees leave quickly. Administrators focus on creditors and asset sales. Systems change hands. Former workers still know manufacturing methods, suppliers, yields, designs and failure modes that may have taken years to develop.
Protection begins before an investigation. A semiconductor company or advanced manufacturer should identify which process recipes, test results, designs and supplier terms create a competitive advantage. Access should follow current job need, with downloads, unusual transfers and new external roles reviewed under clear policy.
Controls must focus on behavior and authority. A second job in a competing field, mass copying before departure, use of personal storage or access after a role ends can create evidence for review. A person's ethnicity or citizenship provides no evidence of wrongdoing.
An orderly shutdown needs the same discipline as normal offboarding. Revoke credentials, preserve logs, collect devices, document who received sensitive assets and keep the records necessary for a buyer, administrator or court. Vendors and research partners should know which data they must return or destroy.
Trade-secret protection often appears as a legal clause. Its practical value depends on data classification, access records, conflict disclosures and a clean exit when a company or employee leaves.
Opportunity Radar
Agent-work accounting for specialist teams
Research groups, software companies, engineering consultancies and regulated professional firms can now generate far more agent activity than managers know how to value. Usage dashboards show tokens, sessions and generated code while leaving accepted output, human intervention and rework unclear.
An AI implementation consultant or workflow-software provider could build a narrow agent-work ledger around one expensive process. The service would connect task initiation, model and tool cost, interventions, review, accepted output, elapsed time, errors and business consequence without capturing unnecessary sensitive content.
The buyer gains evidence for expansion, redesign or cancellation. The provider must validate that its definition of success matches the work and that measurement does not create a surveillance system employees cannot challenge. Accepted throughput, cycle time, reviewer hours, defects, incidents and full cost per accepted result provide a credible scorecard.
Continuity packs for small digital organizations
Nonprofits, associations, independent publishers and community networks often depend on a domain, one email provider, one payment account and a small number of volunteer administrators. A policy action or provider cutoff can disable the whole organization.
A managed service provider or digital-resilience consultant could map those dependencies, create lawful exports, establish secondary administrative access and run a timed migration drill. Customers would pay for shorter outages and preserved records. The service has to validate sanctions, privacy and contractual limits before moving data. Recovery time, missing records, inaccessible accounts and restored communications are more useful measures than a completed checklist.
What You Can Do With This
If your team runs several agents at once
Choose one accepted output and attach every run, intervention, review hour and model charge to it. Compare the agent-assisted workflow with the current process. Report added capacity separately from labor savings and revenue.
If AI recommendations enter physical operations
Track the denominator from recommendation to execution. Record which actions were proposed, approved, rejected and completed, along with safety, cost and outcome. The Cathay trial shows how much results can depend on the operation following the plan.
If your product could serve civilian and military users
Map which sensors, controls, autonomy levels and customer roles change the risk. Test the operator's actual time and information for intervention. Keep legal review, customer screening and technical limits connected to each release.
If one account anchors your organization
Export the records, inventory recovery addresses and test a second administrator. Document how to move the domain, email and payment path if one provider suspends service. Keep the migration lawful and protect the people inside the data.
The Bigger Picture
AI lets one person start more work than one person can personally perform.
OpenAI's researchers are already living inside that change. Three agent-workdays can run alongside one human day. The organization still depends on people to choose the research question, correct the agent, accept the evidence and decide what reaches a model.
Cathay's contrail trial moves the same pattern into aviation. AI predicts an opportunity. Dispatchers, pilots and air-traffic systems determine whether the route changes. The outcome depends on both parts.
China's military research pushes machine action toward higher stakes while exposing the gap between a planned role and a reliable system. The Italian shutdown shows how outside providers can stop a digital organization even when its own technology remains. Belgium's chip case shows why valuable knowledge needs protection through the end of a company's life.
The bottleneck moves as machine capacity grows. Code generation shifts attention to review. Prediction shifts attention to execution. Autonomy shifts attention to authority. Scale shifts attention to dependencies. Valuable knowledge shifts attention to access and exit.
The winning metric will rarely be hours generated. It will be accepted outcomes per unit of total cost, with the failures, interventions and recovery work included.
OpenAI's 3.1 ratio is a useful signal of what is becoming possible. It is also a warning against confusing an occupied machine with completed work.
References
Google: Cathay Pacific contrail-avoidance trial design and early company estimate, September 7, 2026
Sankar and colleagues: Randomized study of scalable airline-led contrail avoidance, March 6, 2026
US Treasury: Government allegations and sanctions against Autistici/Inventati, August 26, 2026
Banca Etica: Account suspension, operating history and stated secondary-sanctions concern
Reuters: Belgian prosecutors disclose suspected semiconductor-espionage case, September 7, 2026
Associated Press: Charges, company history and seized evidence in the BelGaN case, September 7, 2026
AI Next Wave