Introduction
A 2019 study of a health care algorithm is one of the most cited examples of algorithmic bias, and its abstract packs the whole argument into six sentences. News summaries usually keep the finding and drop the cause.
This reading goes through the abstract sentence by sentence to recover the chain of reasoning from a design choice to an unequal outcome.
Locating the Passage
The study is "Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations," published in the journal Science in October 2019. Its lead author is Ziad Obermeyer of the University of California, Berkeley, writing with Brian Powers, Christine Vogeli, and Sendhil Mullainathan. The journal's own page is behind a paywall. The US Federal Trade Commission hosts a free copy of the published article, and the link is in the References.
The passage is the abstract, the single paragraph at the head of the article. This reading refers to its sentences by number, first to sixth. One step also draws on the body of the article. All quotations are from the study (Obermeyer et al. 2019).
The authors capitalize "Black" and "White," and quotations here follow them.
Walking Through the Passage
Step 1: Read what the algorithm was for
The first sentence reads: "Health systems rely on commercial prediction algorithms to identify and help patients with complex health needs."
Three things are in this sentence. The algorithms are commercial, meaning companies sell them to hospitals and insurers. Their job is to identify patients. And the purpose is to help: patients picked out by the algorithm get extra attention, such as a dedicated care team.
So a high score is a good thing to receive, and the algorithm is deciding who gets a benefit. The question to carry forward is whether the patients it selects are the ones with the greatest need.
The second sentence says the algorithm studied is "typical of this industry-wide approach" and affects "millions of patients." The authors are presenting it as representative of a class.
Step 2: Read the finding that at the same score Black patients were sicker
The second sentence continues: the algorithm "exhibits significant racial bias: At a given risk score, Black patients are considerably sicker than White patients, as evidenced by signs of uncontrolled illnesses."
The phrase to read slowly is "at a given risk score." The comparison is between patients the algorithm rated the same. If the score measured health need, two patients with the same score should be about equally sick. The authors found they weren't. Among patients with equal scores, the Black patients had more illness.
Put the other way round, a Black patient had to be sicker than a white patient to receive the same score. The same level of sickness earned a lower score and a smaller chance of extra help.
The authors measured sickness directly from medical records, using signs of illnesses that were not under control. That gave them a yardstick independent of the score.
Step 3: Read the cause the authors identify
The fourth sentence gives the cause: "The bias arises because the algorithm predicts health care costs rather than illness."
The builders wanted to find patients with high health needs. Need isn't recorded in a database as a number. Cost is. Every patient has a billing history, and people who are sicker generally cost more to treat. So the algorithm was trained to predict next year's cost, and cost stood in for need.
A measurable substitute of this kind is called a proxy. The fifth sentence concedes that the proxy looked reasonable: health care cost appears "to be an effective proxy for health by some measures of predictive accuracy." The algorithm predicted costs well. The trouble lay in what cost left out.
Step 4: Follow the reasoning from unequal access to unequal spending to unequal scores
The fourth sentence continues: "but unequal access to care means that we spend less money caring for Black patients than for White patients."
The reasoning runs in a line.
- Black patients, on average, have had less access to care.
- Less access means fewer appointments, tests, and procedures, so less money is spent on a Black patient than on a white patient who is equally sick.
- The algorithm learns from spending. It sees a lower bill and predicts a lower future bill.
- A lower predicted cost is a lower score, and a lower score means the patient is less likely to be selected for help.
The algorithm reproduced an inequality that already existed in the health system. Its cost predictions were accurate, and that accuracy is how the bias got in.
Step 5: Read what changing the predicted quantity did
The third sentence states the size of the effect: "Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%."
This is a calculation of what would happen if patients were selected by how sick they were. The share of Black patients among those selected would rise from under a fifth to nearly half.
The body of the article reports a practical test. The authors worked with the algorithm's manufacturer on a version that predicted a combination of health and cost. They report that this produced "an 84% reduction in bias" by their measure. The data and the patients stayed the same, and the target changed.
The sixth sentence draws the general lesson: "the choice of convenient, seemingly effective proxies for ground truth can be an important source of algorithmic bias in many contexts." Ground truth means the real quantity you care about, as opposed to the stand-in you can measure.
Key Considerations
The algorithm didn't use race as an input. The body of the article says so: "the algorithm specifically excludes race." The bias came through cost, which was correlated with race because of unequal access to care. Leaving a characteristic out of the inputs doesn't stop a system from producing different results for groups defined by it.
The common mistake is to assume that bias requires a biased variable or biased intent. This case had neither. Every input was a routine billing or medical record, and the builders were trying to direct help to patients who needed it. The bias came from one design decision, the choice of what to predict, that seemed sensible and was never examined by group.
A related mistake is to treat accuracy as proof of fairness. The algorithm was accurate at its stated task of predicting cost. Whether it was accurate and whether it was predicting the right thing were separate matters.
Two limits on the study should be kept in view. It examined one algorithm at one health system, though the authors argue it is typical. And the manufacturer cooperated in testing a fix, which the abstract doesn't mention and which doesn't always happen.
Summary
The abstract traces an unequal outcome to a single design choice. In plain words, the causal chain has four steps:
- The builders wanted to find the patients with the greatest health needs, and chose to predict future health care cost as a stand-in for need.
- Because of unequal access to care, less money was spent on Black patients than on equally sick white patients.
- The algorithm, trained on spending, gave Black patients lower scores than equally sick white patients.
- Lower scores meant that fewer Black patients were selected for extra help: 17.7 percent of those selected, where selection by illness would have made it 46.5 percent.
References
- Obermeyer, Ziad, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019. "Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations." Science 366 (6464): 447–453.