September 28, 2026

Hospital AI Finds More Diagnoses. Your Premium Picks Up the Tab.

Blue Cross plans say AI billing software added $942 million to hospital claims in two years while the care stayed the same. Hospitals say the patients are sicker, and which side is right decides part of what you pay next year.

A printed hospital bill and a stethoscope on an office desk under a single lamp, with an empty window-lit space beside them.

The Short Version

Last Thursday the Blue Cross Blue Shield Association put a number on something health plans have been grumbling about for a year. Between early 2023 and the end of 2025, hospitals billed a steadily larger share of inpatient stays as medically complex. The association estimates that shift cost its member plans about $942 million. The surgeries stayed the same. What changed was the list of secondary diagnoses on each claim, conditions like low sodium or anemia after an operation, which can move a stay into a higher-paying billing category. The association's researchers went looking for the treatment that should follow those diagnoses and found less of it than the diagnoses implied. Their explanation is AI billing software that scans lab results and doctors' notes for anything codable. The American Hospital Association calls that unsubstantiated and says patients really are older and sicker. The whitepaper works from claims, and the hospitals answer with severity statistics, so a look inside the charts is what's missing. What's settled is that health plans told PwC in June they expect employer health costs to rise about 9 percent next year, and they rank these tools among the top reasons. That part lands on your premium.

What the Insurers Found

The analysis draws on de-identified claims from Blue plans that cover about one in three Americans, from early 2023 through the end of 2025. Over that stretch, the share of inpatient stays billed with a complication code rose from roughly 37 percent to about 40 percent. Three points sounds small until you price it. That's how the association reaches $942 million.

Most of the money traces to one mechanism. Against the 2023 baseline, hospitals classified about 55,000 additional stays as complex because of secondary diagnoses, and those stays paid roughly $11,800 more apiece. That's about 70 percent of the total, and it's the part the paper documents in detail.

To see what was driving it, the researchers zoomed in on major bowel procedures, which mostly treat colon cancer and diverticular disease. The fastest-growing add-on diagnoses were ones you can pull from a single lab value: unspecified acidosis, low sodium, and acute blood-loss anemia. The growth was lopsided. A quarter of hospitals accounted for most of it. Within that group the anemia code rose about five and a half points, while at everyone else it stayed flat.

Then comes the test the whole paper rests on. If patients have anemia serious enough to code, some of them should get a transfusion. At the top-growth hospitals, the anemia diagnosis rate on bowel-surgery patients ran 38 percent higher than at their peers. Yet among patients with that diagnosis, the transfusion rate was lower, about 17 percent versus 19 percent. ICU use, reoperations and length of stay looked the same or lower too. "If patients are truly sicker, we'd expect to see more treatment," said Luke Chalker, the association's senior vice president of product and data science. His conclusion is that the software is surfacing billable conditions while the patients stay about as sick as they were.

Where the Argument Gets Hard

The American Hospital Association answered this line of research in July with a fact sheet, and its case deserves a fair hearing. Hospitals treat an older population with more chronic disease, and the simplest cases have moved to outpatient settings, which leaves the inpatient mix sicker by design. An AHA and Vizient analysis puts the rise in case-mix index, the standard measure of patient severity, at about 5 percent from 2019 to 2024. The coding rules are unchanged, the association argues, and AI tools work inside them with human validation. Then it turns the accusation around. Insurers run automated programs that downcode claims without reading the chart, and insurers have themselves been caught adding unsupported diagnoses to lift Medicare Advantage risk scores. "These plans want it both ways," the fact sheet says.

Some of that lands. Read the whitepaper closely, though, and the sicker-patients argument runs into the concentration finding. Aging and chronic disease should push up coding at every hospital in a peer group. Yet the growth clustered in a quarter of them, including among teaching hospitals compared with other teaching hospitals. That pattern points to coding practice.

The whitepaper has its own soft spots, and I'd weigh them before repeating the headline number. It's built entirely from claims. It leaves out which hospitals actually installed AI coding software, so the link between tool and trend is inferred from timing and from the kinds of diagnoses that grew. About 30 percent of the $942 million sits outside the secondary-diagnosis story the paper explains. And the authors work for the payer that would benefit from a tougher line on hospital claims. What survives is a well-built argument from one side of a payment dispute, and the other side has yet to produce the chart-level evidence that would end it.

The Software Between the Doctor and the Bill

The tools in question sit between the clinician and the claim. Ambient listening apps record the visit and draft the note. Documentation software reads that note plus the lab panel and suggests conditions a coder might have missed: a sodium reading a little low, a hemoglobin a little under range after surgery. Each suggestion becomes a secondary diagnosis, and enough of them push the stay into a higher-severity billing group. Hospitals bought these tools for reasons any clinic manager would recognize. Coders are scarce, notes are long, and insurers deny claims that lack specificity.

PwC's actuaries heard the plan side of that story in June. Nearly 70 percent of the health plans they surveyed ranked AI documentation and coding tools among their top three cost inflators for 2027, and about one in five called them the biggest. Those plans expect employer health costs to rise about 9 percent next year, the highest rate in more than a decade. Glenn Hunzinger, who leads PwC's U.S. health practice, said the technology lets providers "code things that they were never able to."

The response is an arms race, and executives on both sides describe it that way. Elevance's chief executive told investors last year that AI coding tools let hospitals increase documented acuity, and Centene's said it needed AI in its own payment-integrity work to keep pace. Payers said they would expand clinical-validation audits and keep downcoding disputed claims. Hospital executives see the same fight from the other trench. HCA's finance chief said hospitals are behind the payers on automated claims review, and Shiv Rao, who founded the ambient-documentation company Abridge, described the endgame as "bots fighting bots, agents fighting agents."

Chalker rejects the idea that it's a fair fight at all. Speaking to TechCrunch, he called it "a completely one-sided blood bath," with the plans losing. Either way, a lab-derived diagnosis that lifted a claim into a higher-paying group in 2024 may draw an audit letter in 2027.

A Community Hospital Decides

Picture a 180-bed community hospital whose revenue-cycle director has a vendor proposal on her desk. The software promises to read every lab panel and draft note and surface missed diagnoses, and the sales deck projects a mid-single-digit lift in case-mix index. Her compliance officer has read the Blue Cross paper. Between them, they've got a decision the coding rules leave open.

What has to be in place first is a definition of success the compliance side can defend. A lift in case-mix index is revenue, and revenue alone is the wrong target, because a payer auditor will measure the same lift and call it upcoding. The better target is accuracy: diagnoses a physician confirms and the chart supports. So before the pilot starts, the hospital pulls an eight-quarter baseline: the share of stays billed as complex, the rate of each diagnosis the whitepaper flags, and the share of those patients who got the matching treatment. Anemia against transfusions. Malnutrition against dietitian consults. Those ratios are the hospital's own version of the whitepaper's test, and a Blue plan will run them anyway.

The pilot stays small: one service line for one quarter, with every suggestion routed through a physician query before it reaches the coder. Every accepted suggestion is counted, and so is every one a physician declines. The added review work belongs on the cost side of the ledger, including a documentation specialist's hours, physician time answering queries, and a compliance sample of 30 charts a month checked against the treatment record. If accepted diagnoses rise while the matching treatments stay flat, the tool is finding paperwork, and the compliance officer can pause it before the next quarter's claims go out. If diagnoses and treatments rise together, the hospital was under-documenting sicker patients all along, and it has the evidence to say so when the audit letter arrives.

Where this fails is predictable. The vendor's settings may flag every out-of-range value, which is the exact behavior the whitepaper describes. Physicians may rubber-stamp queries to clear an inbox. Leadership may move the goal back to revenue once the pilot ends. The fix for all three is the same: publish the treatment-match ratios internally every quarter, and give the compliance officer the final say on the tool's settings.

Opportunity Radar

If you've spent years in hospital coding or nursing audit, there's a service in this fight. Community hospitals are buying AI documentation tools without anyone checking whether the added diagnoses match the treatment in the chart, and insurers have said they'll audit exactly that. The offer is a quarterly pre-bill review of the lab-derived diagnoses that carry the most payment risk, delivered as a treatment-match report the compliance officer can hand to a payer. The buyer is the compliance or revenue-integrity lead, who already pays outside auditors and now has a whitepaper to justify the line item. Test it with one service line at one hospital for a fixed fee, and count how many flagged diagnoses they withdraw or defend. Walk away if the hospital only wants help capturing more codes, because then you are the upcoding, and the audit letter will have your name on it.

What You Can Do With This

If you buy health coverage for a business

Your 2027 renewal is being priced right now, and plans are building coding growth into it. Ask your broker what share of your group's inpatient claims were billed as complex in each of the last three years, and whether the plan validates lab-derived diagnoses before paying. A plan that can answer has a lever on your trend; a plan that shrugs is passing the increase through.

If you code, chart or document for a living

Treat every AI suggestion as a query, and answer it with the treatment, or the reason none was needed, written into the note. An anemia code with a hemoglobin value and no transfusion is exactly what the whitepaper was built to find. A diagnosis with clinical support and a documented plan survives an audit, and it's also just good medicine.

If you're a patient with a hospital bill

Ask for the itemized bill and the diagnoses on the claim. If you see conditions nobody mentioned during your stay, ask the billing office which lab value or note supports them. A hospital that adds a diagnosis should be able to show you the care that went with it.

The Bigger Picture

The consequence this story points to is a change in who gets to define "accurate." For decades, accurate coding meant what a trained human could find in the chart in the time available. AI removed the time limit, and the definition moved with it. Every out-of-range lab value is now a candidate diagnosis, and the coding rules, written for humans, permit it. Insurers will answer with their own machines, and you'll pay for both sides of that arms race through premiums and review hours. My read is that the number that settles this fight will come from a chart audit, and whichever side publishes one first, with the treatment record beside each diagnosis, will set the terms of the next contract talks. A prediction, labeled as one: by the time 2027 premiums are final, at least one large payer will make treatment-matched validation a condition for paying lab-derived secondary diagnoses. Payers told investors last year they were heading there. Whether your hospital's patients are sicker is a question only the chart can answer.

References

BCBSA analysis examines AI use in hospital billing practice (Blue Cross Blue Shield Association press release, September 24, 2026) Hospital Coding Intensity Analysis: Major Bowel Procedures, whitepaper (Blue Cross Blue Shield Association, September 2026) Hospitals' use of AI coding tools cost BCBSA plans $942M more for similar care (Fierce Healthcare, September 24, 2026) Hospitals say they're losing the AI billing war. A new BCBS study suggests otherwise (Becker's Hospital Review, September 24, 2026) Insurers claim AI is already increasing healthcare costs (TechCrunch, September 26, 2026) Fact Sheet: Artificial Intelligence and Coding Intensity (American Hospital Association, July 31, 2026) Medical cost trend 2027: Behind the Numbers (PwC Health Research Institute, June 11, 2026) Health plans say AI is pushing healthcare costs higher (Healthcare Dive, June 11, 2026) AI's healthcare side hustle: inflating your bill (Tech Brew, June 12, 2026) Reeling payers plan to increase scrutiny of providers' coding practices (HFMA, July 29, 2025)