KnowledgeInSight
AI Literacy
0% of Course 3 complete

Module 3 · Lesson 2

Information and Rights: Bias, fairness, and automated decisions

You'll see how bias gets into AI systems through data and design, and why "fair" has several definitions that conflict. You'll be able to explain a well-known dispute over a criminal-justice algorithm and say why both sides had a point.

What you will be able to do

  • Explain how bias enters AI systems and compare the standards of fairness used to judge them.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional journal

Contents of this lesson9 items
  1. ReadingTwo Fair Algorithms That Can't Both Be Fair: The COMPAS Dispute3 min
  2. ReadingHow Bias Enters AI Systems: Training Data, Proxy Measures, and Design Choices4 min
  3. ReadingThe COMPAS Findings: Error Rates by Group and the Vendor's Reply4 min
  4. ReadingCompeting Definitions of Fairness and Why They Can't All Hold at Once4 min
  5. ReadingAutomated Decisions in Hiring, Credit, and Public Services: Audits and Notice4 min
  6. Guided ReadingGuided Close Reading: The Health Algorithm Study's Account of a Biased Proxy7 min
  7. Guided ConversationChoose a Fairness Standard for One Decision12 min
  8. Journal · optionalWhich Error Would You Rather Risk?15 min
  9. Knowledge CheckBias, fairness, and automated decisions10 min

Reading 3 min

Two Fair Algorithms That Can't Both Be Fair: The COMPAS Dispute

In 2016 a news investigation reported that software used in American courts was biased against Black defendants. The company that made the software answered that it treated Black and white defendants alike. Each side had numbers, and the numbers on both sides were right.

The software was called COMPAS. It gave each defendant a score for the risk of committing another crime, and judges in several states saw such scores when they set bail or passed sentence. Reporters at ProPublica, a nonprofit news organization, obtained the scores of more than 7,000 people arrested in Broward County, Florida. They then checked who was charged with a new crime over the next two years (Angwin et al. 2016).

Their central finding was about mistakes. Among defendants who did not go on to reoffend, 44.9 percent of Black defendants had been labeled higher risk. For white defendants the figure was 23.5 percent. The reporters wrote that the formula wrongly flagged Black defendants "at almost twice the rate as white defendants" (Angwin et al. 2016).

Northpointe, the company that sold COMPAS, rejected the analysis. Its defense, as later summarized by three computer scientists, was that the scores were "equally well calibrated" for both groups (Kleinberg, Mullainathan, and Raghavan 2016). In plain terms, a given score meant the same thing whoever received it. Black and white defendants with the same score went on to reoffend at similar rates. Northpointe had a commercial stake in that defense, and the claim could still be checked against the data.

So one side said the score made unequal errors, and the other said the score meant the same thing for everyone. Both were describing the same records accurately.

Later in 2016 those computer scientists, Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan, showed why. They wrote down three conditions that each sound like a reasonable requirement for a fair score, including the two in this dispute. They then proved that "except in highly constrained special cases, there is no method that can satisfy these three conditions simultaneously" (Kleinberg, Mullainathan, and Raghavan 2016). When two groups have different underlying rates of the outcome being predicted, a score that means the same thing for both groups will make unequal errors.

Better engineering can't remove that conflict, because it comes from the arithmetic. No version of COMPAS, and no replacement for it, could have satisfied both ProPublica's standard and Northpointe's on that data.

This changes what you should ask when you read that an algorithm is "biased" or has been "shown to be fair." Each of those words rests on a particular measurement. Fairness can mean that errors fall equally on each group, or that scores mean the same for each group, or something else again. A claim that names its measure can be checked. A claim that doesn't has left out the part you most need to know, along with the matter of who chose that measure.

References

  • Angwin, Julia, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016. "Machine Bias." ProPublica, May 23, 2016.
  • Kleinberg, Jon, Sendhil Mullainathan, and Manish Raghavan. 2016. "Inherent Trade-Offs in the Fair Determination of Risk Scores." arXiv:1609.05807.

Report an issue with this item

Reading 4 min

How Bias Enters AI Systems: Training Data, Proxy Measures, and Design Choices

Introduction

When an AI system treats groups of people differently, the first assumption is often that someone built prejudice into it. In the best-documented cases nobody did.

This reading uses three studied cases to show three routes by which bias gets into a system: the data it learns from, the history that data records, and the choice of what to predict.

Skewed Data

An AI system learns from examples, and it does best on the kinds of example it saw most. Training data bias is a slant in a system's results that comes from training data that over-represents some cases and under-represents others.

In 2018 Joy Buolamwini of the Massachusetts Institute of Technology and Timnit Gebru, then at Microsoft Research, tested three commercial systems that classify a person's gender from a photograph of their face. Darker-skinned women were the most misclassified group, with error rates of up to 34.7 percent. For lighter-skinned men the highest error rate was 0.8 percent (Buolamwini and Gebru 2018).

The authors also examined two collections of face photographs widely used to test such systems. One was 79.6 percent lighter-skinned and the other 86.2 percent. Against test sets this lopsided, a system that was weak on darker-skinned faces could still post a high overall score.

Learned History

Balanced data can still carry a problem if it records a biased past. Historical bias is a pattern of past unequal treatment that a system learns from records of past decisions.

In 2018 Reuters reported on an experimental recruiting tool built at Amazon. It scored job applicants from one to five stars, and it had been trained on résumés submitted to the company over ten years. Most of those came from men. According to the five people familiar with the project who spoke to Reuters, the tool "penalized resumes that included the word 'women's,'" as in the name of a women's club, and marked down graduates of two all-women's colleges (Dastin 2018).

The tool had learned what past hiring looked like and treated that as what a good applicant looks like. The company edited it to ignore those particular terms, and later disbanded the team.

The Wrong Target

The third route is the hardest to see. Many things people care about can't be measured directly, so builders pick something measurable to stand in for them. A proxy is a measurable substitute for something that's harder to measure directly. Label choice is the decision about which quantity a system is trained to predict.

In 2019 a team led by Ziad Obermeyer of the University of California, Berkeley, studied a commercial algorithm that US health systems used to decide which patients should get extra help with complex health needs. The algorithm was meant to find the sickest patients. It did so by predicting each patient's future health care costs (Obermeyer et al. 2019).

Cost turned out to be a biased proxy. The authors report that "unequal access to care means that we spend less money caring for Black patients than for White patients." Lower spending led to lower predicted cost and a lower score. At the same score, Black patients were considerably sicker than white patients. The authors calculated that removing the disparity would raise the share of Black patients receiving additional help from 17.7 to 46.5 percent.

The algorithm did not use race as an input.

The Three Cases Compared

Source of biasCaseWhat went wrongWhat was fixed or abandoned
Skewed dataCommercial gender classification (Buolamwini and Gebru 2018)Systems checked against mostly lighter-skinned test sets made far more errors on darker-skinned womenNot reported in the study; the authors called for urgent attention from the companies
Learned historyExperimental recruiting tool (Dastin 2018)A tool trained on ten years of mostly male résumés marked down the word "women's"Terms edited out, then the team was disbanded
Proxy and label choiceHealth-needs algorithm (Obermeyer et al. 2019)Predicting cost in place of illness scored sicker Black patients as lower needThe authors worked with the maker on a version that predicted health as well as cost

What the Cases Share

None of these cases required intent. No one in any of them is reported to have set out to disadvantage a group. Each system did what it was built to do: fit its data and predict its target.

Removing a sensitive characteristic from the inputs doesn't guard against these failures. The health algorithm excluded race and still produced unequal results, because cost carried the effect of unequal access. The recruiting tool wasn't given applicants' sex and found a word that tracked it.

Each problem was found only when someone compared results across groups. By their overall figures the systems looked accurate.

Conclusion

Bias reaches AI systems through skewed training data, through records of unequal past treatment, and through the choice of a proxy to predict. The three cases here are well documented, and in each the builders responded once the problem was shown. None involved a deliberately biased rule, which is why checking results by group matters more than inspecting the inputs.

Key Terms

  • Training data bias: A slant in a system's results that comes from training data that over-represents some cases and under-represents others.
  • Historical bias: A pattern of past unequal treatment that a system learns from records of past decisions.
  • Proxy: A measurable substitute for something that's harder to measure directly.
  • Label choice: The decision about which quantity a system is trained to predict.

References

  • Buolamwini, Joy, and Timnit Gebru. 2018. "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification." In Proceedings of the 1st Conference on Fairness, Accountability and Transparency. Proceedings of Machine Learning Research 81:77–91.
  • Dastin, Jeffrey. 2018. "Amazon Scraps Secret AI Recruiting Tool That Showed Bias Against Women." Reuters, October 10, 2018.
  • Obermeyer, Ziad, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019. "Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations." Science 366 (6464): 447–453.

Report an issue with this item

Reading 4 min

The COMPAS Findings: Error Rates by Group and the Vendor's Reply

Introduction

One piece of software is cited in nearly every discussion of algorithmic fairness: COMPAS, a tool used in US courts to estimate whether a defendant will commit another crime. Journalists said it was biased, its maker said it wasn't, and the argument turned on what each side chose to count.

This reading sets out what the score was for, what the journalists found, how the company replied, and why the case became the standard example.

What the Score Was For

A risk score is a number that estimates how likely a person is to have some future outcome, such as being arrested again. COMPAS produced one from a defendant's answers to a questionnaire and from their record. It was sold by Northpointe, a for-profit company.

In 2016 ProPublica, a nonprofit news organization, reported that scores of this kind were used at many points in the criminal justice system, from setting bond amounts to sentencing, and that in nine states they were given to judges during sentencing. The company did not publicly disclose how the score was calculated. Race was not one of the questions (Angwin et al. 2016).

What ProPublica Analyzed

The reporters obtained the scores of more than 7,000 people arrested in Broward County, Florida, in 2013 and 2014. They then checked public records to see who was charged with a new crime in the following two years. That let them compare each prediction with what happened.

A score can be wrong in two directions, and each has a name.

  • The false positive rate is, among people who didn't go on to have the outcome, the share the system wrongly flagged as likely to have it.
  • The false negative rate is, among people who did go on to have the outcome, the share the system wrongly rated as unlikely to have it.

For a defendant, a false positive can mean higher bail or a longer sentence for someone who would not have reoffended. A false negative means someone rated low risk who did reoffend.

The Finding

ProPublica reported both rates for Black and white defendants (Angwin et al. 2016).

White defendantsBlack defendants
Labeled higher risk, but didn't reoffend23.5 percent44.9 percent
Labeled lower risk, yet did reoffend47.7 percent28.0 percent

The two rows run in opposite directions. Black defendants who didn't reoffend were labeled higher risk nearly twice as often as white defendants who didn't reoffend. White defendants who did reoffend were labeled lower risk far more often than Black defendants who did.

The reporters wrote that the score "made mistakes with black and white defendants at roughly the same rate but in very different ways." They also reported how well it predicted overall. Of those it deemed likely to reoffend, 61 percent were arrested for a new crime within two years, which the article called "somewhat more accurate than a coin flip."

The Reply

Northpointe disputed the analysis. Its statement to ProPublica said it "does not agree that the results of your analysis, or the claims being made based upon that analysis, are correct" (Angwin et al. 2016).

The substance of its defense concerned a different measurement. Calibration is the property that a given score corresponds to the same real rate of the outcome in every group. The defense was that COMPAS was calibrated: among defendants given the same score, Black and white defendants went on to reoffend at similar rates. Three computer scientists who studied the dispute described this as the claim that the score's estimates were "equally well calibrated to the true outcomes for both African-American and white defendants" (Kleinberg, Mullainathan, and Raghavan 2016).

On this measure, a judge reading a score would not need to know a defendant's race to know what the score meant. The company sold the product and had every reason to defend it. Its claim was also one that could be checked against the same records, and it wasn't shown to be false.

Why the Case Became the Reference Point

Three features made COMPAS the standard example.

  1. The stakes were concrete. The scores were shown to judges deciding on people's liberty.
  2. Both claims were arithmetic on the same data. Neither side depended on a secret. Anyone with the records could confirm the unequal error rates and the similar meaning of each score.
  3. The disagreement was about the definition. The two sides did not dispute the facts. They disputed which facts made a score fair.

That third feature turned a news story into a research question. Kleinberg and his colleagues opened their paper with the dispute and went on to ask whether any score could meet both standards (Kleinberg, Mullainathan, and Raghavan 2016).

Conclusion

ProPublica found that COMPAS's errors fell unequally: Black defendants who didn't reoffend were flagged as higher risk nearly twice as often as white defendants who didn't. Northpointe's defense was that each score meant the same for both groups. Both statements described the same records correctly, and the case is cited because it shows two reasonable standards of fairness giving opposite verdicts on one system.

Key Terms

  • Risk score: A number that estimates how likely a person is to have some future outcome, such as being arrested again.
  • False positive rate: Among people who didn't go on to have the outcome, the share the system wrongly flagged as likely to have it.
  • False negative rate: Among people who did go on to have the outcome, the share the system wrongly rated as unlikely to have it.
  • Calibration: The property that a given score corresponds to the same real rate of the outcome in every group.

References

  • Angwin, Julia, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016. "Machine Bias." ProPublica, May 23, 2016.
  • Kleinberg, Jon, Sendhil Mullainathan, and Manish Raghavan. 2016. "Inherent Trade-Offs in the Fair Determination of Risk Scores." arXiv:1609.05807.

Report an issue with this item

Reading 4 min

Competing Definitions of Fairness and Why They Can't All Hold at Once

Introduction

Most people assume that a fair algorithm is possible in principle and that the work lies in building it. A proof published in 2016 showed that "fair" has several reasonable meanings that conflict.

This reading states three fairness conditions in plain words, shows the conflict with a small table of invented numbers, and says what the result does and doesn't imply.

Three Conditions

A fairness definition is a precise statement of what must be equal across groups for a system to count as fair. Jon Kleinberg and Manish Raghavan of Cornell University and Sendhil Mullainathan, then of Harvard University, studied three such definitions for systems that give people a risk score (Kleinberg, Mullainathan, and Raghavan 2016).

  1. A score means the same for every group. Among people given a particular score, the share who turn out to have the outcome is the same whichever group they belong to. This is calibration: the property that a given score corresponds to the same real rate of the outcome in every group.
  2. People who don't have the outcome get similar scores across groups. Someone who would never have reoffended shouldn't be scored as riskier because of their group.
  3. People who do have the outcome get similar scores across groups. Someone who will reoffend shouldn't be scored as safer because of their group.

Conditions 2 and 3 together amount to error-rate balance: the property that a system's false positive rate and false negative rate are the same in every group. A false positive is a person wrongly flagged as high risk, and a false negative is a person wrongly rated low risk.

The Result

The authors proved that the three can't all hold, apart from special cases. In their words, "except in highly constrained special cases, there is no method that can satisfy these three conditions simultaneously."

A base rate is the share of a group that actually has the outcome being predicted. The special cases are perfect prediction and equal base rates. When two groups' base rates differ and prediction is imperfect, at least one condition must fail.

A Small Example

The numbers below are invented. A score sorts two groups of 100 people into "high risk" and "low risk." It is built to be calibrated: in both groups, 60 percent of the people labeled high risk have the outcome, and 20 percent of the people labeled low risk do.

Group AGroup B
People who will have the outcome40 of 10030 of 100
Labeled high risk50 (30 have the outcome, 20 don't)25 (15 have the outcome, 10 don't)
Labeled low risk50 (10 have the outcome, 40 don't)75 (15 have the outcome, 60 don't)

Condition 1 holds. Now look at the errors.

Group AGroup B
People without the outcome who were labeled high risk20 of 60, about 33 percent10 of 70, about 14 percent
People with the outcome who were labeled low risk10 of 40, 25 percent15 of 30, 50 percent

Conditions 2 and 3 fail. A person in Group A who won't have the outcome is more than twice as likely to be wrongly labeled high risk as a similar person in Group B. A person in Group B who will have the outcome is twice as likely to be missed.

Adjusting the labels to even out the errors would make the high-risk label mean something different in each group, and condition 1 would fail. The difference in base rates, 40 against 30, forces a choice.

What Follows

A trade-off is a situation in which getting more of one thing you want means accepting less of another. These fairness conditions trade off against each other, so choosing among them is a decision about values. The mathematics shows that the conditions conflict and is silent on which matters more. A person who cares most about what a score means to the decision-maker will favor calibration. A person who cares most about who bears the mistakes will favor error-rate balance. The proof also leaves open who should choose.

What Doesn't Follow

The result is sometimes read as showing that fairness is hopeless. Three things it doesn't say:

  • That bias can't be reduced. A system trained on skewed data or aimed at a poor stand-in for what matters can be improved on every measure by fixing the data or the target.
  • That all systems are equally fair. Two systems can differ widely in how far they depart from each condition.
  • That base rates are fixed facts. A measured base rate can itself reflect unequal treatment, such as differences in who gets arrested for the same conduct.

Conclusion

Three reasonable conditions for a fair risk score can't all be met when groups differ in their base rates and prediction is imperfect. That is a mathematical result and isn't contested. Which condition to give up in a particular use, and who decides, is a matter of values and remains open.

Key Terms

  • Fairness definition: A precise statement of what must be equal across groups for a system to count as fair.
  • Calibration: The property that a given score corresponds to the same real rate of the outcome in every group.
  • Error-rate balance: The property that a system's false positive rate and false negative rate are the same in every group.
  • Base rate: The share of a group that actually has the outcome being predicted.
  • Trade-off: A situation in which getting more of one thing you want means accepting less of another.

References

  • Kleinberg, Jon, Sendhil Mullainathan, and Manish Raghavan. 2016. "Inherent Trade-Offs in the Fair Determination of Risk Scores." arXiv:1609.05807.

Report an issue with this item

Reading 4 min

Automated Decisions in Hiring, Credit, and Public Services: Audits and Notice

This content reflects the field as of October 2026.

Introduction

If you have applied for a job, a loan, or a public benefit in recent years, software may have sorted your application before any person read it. You probably weren't told.

This reading describes where such systems are used, a current concern about language models in screening, and the two checks most often applied: audits and notice. It uses one city's rule as an example and doesn't argue whether such rules are enough.

What Counts as an Automated Decision

An automated decision is a decision about a person that is made or substantially shaped by a computer system's output. The definition covers more than a machine deciding alone. A recruiter who interviews only the candidates a tool ranked highest has made a decision shaped by the tool.

Such systems are used to screen job applicants, to approve loans and set credit limits, and to decide who is eligible for public services or flagged for review. Organizations adopt them because they handle large numbers of cases quickly and apply one procedure to all.

A Current Concern: Covert Bias in Language Models

Organizations can now ask a general-purpose language model to read an application and give a judgment. A 2024 laboratory study tested what such models do with dialect.

Valentin Hofmann and colleagues at Stanford University, the University of Chicago, and other institutions gave language models text written in African American English and matching text in standardized American English. The models included older research models and commercial chatbot models, from several developers. The researchers then asked the models to make judgments about the writers (Hofmann et al. 2024).

When asked directly about African Americans, the models' stated views were positive. When given only the dialect, the models were "more likely to suggest that speakers of AAE be assigned less-prestigious jobs, be convicted of crimes and be sentenced to death." AAE is the authors' abbreviation for African American English. They call this pattern "covert racism in the form of dialect prejudice."

The study found that methods used to reduce overt bias left this covert bias in place. It was an experiment with constructed prompts and didn't observe real hiring decisions. It shows a risk in using these models for screening and doesn't show how often real applicants have been affected.

Bias Audits

A bias audit is a test of whether a system's outcomes differ across groups, such as by sex or race. An auditor runs the system on real or test cases, records the results by group, and compares them. For a hiring tool, the comparison is usually between the rates at which candidates from each group are selected or scored highly.

What an audit looks for is disparate impact: a marked difference in outcomes between groups, whether or not anyone intended it.

Audits can be done before a system is put to use and repeated while it is running, because the people a system sees in use may differ from the cases it was built on.

One Rule in Force: New York City

New York City regulates what it calls automated employment decision tools. Under its Local Law 144, the city's Department of Consumer and Worker Protection states that employers and employment agencies are prohibited from using such a tool unless three things are true (New York City Department of Consumer and Worker Protection n.d.).

RequirementWhat the department's page says
Auditthe tool "has been subject to a bias audit within one year of the use of the tool"
Publication"information about the bias audit is publicly available"
Notice"certain notices have been provided to employees or job candidates"

Notice means telling a person that an automated system is being used in a decision about them. The department began enforcing the law on July 5, 2023.

This is one city's rule for one kind of decision, as it stood in October 2026. It isn't legal advice, and other places have different rules or none.

What Audits Can Miss

An audit answers the question it is set up to ask. Two gaps follow from that.

  • The choice of fairness measure. An audit that compares selection rates says nothing about whether errors fall equally, or whether a score means the same for each group. A tool can pass on one measure and fail on another.
  • Groups not tested for. An audit by sex and race won't detect a disadvantage by age, disability, or dialect. The dialect study is an example of a bias that a standard audit by declared race could miss, because the model was responding to how people wrote.

Notice has its own limit. Knowing that a tool was used doesn't tell an applicant how it scored them or why.

Conclusion

As of October 2026 automated systems shape decisions in hiring, lending, and public services, and laboratory evidence shows that language models can carry bias that direct questioning doesn't reveal. The main checks in use are audits of outcomes by group and notice to the people affected, with New York City's hiring rule as one example in force. Audits test what they are designed to test, and whether these checks are sufficient is disputed.

Key Terms

  • Automated decision: A decision about a person that is made or substantially shaped by a computer system's output.
  • Bias audit: A test of whether a system's outcomes differ across groups, such as by sex or race.
  • Disparate impact: A marked difference in outcomes between groups, whether or not anyone intended it.
  • Notice: Telling a person that an automated system is being used in a decision about them.

References

  • Hofmann, Valentin, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. "AI Generates Covertly Racist Decisions About People Based on Their Dialect." Nature 633 (8028): 147–154.
  • New York City Department of Consumer and Worker Protection. n.d. "Automated Employment Decision Tools (AEDT)." Accessed October 3, 2026.

Report an issue with this item

Guided Reading 7 min

Guided Close Reading: The Health Algorithm Study's Account of a Biased Proxy

Introduction

A 2019 study of a health care algorithm is one of the most cited examples of algorithmic bias, and its abstract packs the whole argument into six sentences. News summaries usually keep the finding and drop the cause.

This reading goes through the abstract sentence by sentence to recover the chain of reasoning from a design choice to an unequal outcome.

Locating the Passage

The study is "Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations," published in the journal Science in October 2019. Its lead author is Ziad Obermeyer of the University of California, Berkeley, writing with Brian Powers, Christine Vogeli, and Sendhil Mullainathan. The journal's own page is behind a paywall. The US Federal Trade Commission hosts a free copy of the published article, and the link is in the References.

The passage is the abstract, the single paragraph at the head of the article. This reading refers to its sentences by number, first to sixth. One step also draws on the body of the article. All quotations are from the study (Obermeyer et al. 2019).

The authors capitalize "Black" and "White," and quotations here follow them.

Walking Through the Passage

Step 1: Read what the algorithm was for

The first sentence reads: "Health systems rely on commercial prediction algorithms to identify and help patients with complex health needs."

Three things are in this sentence. The algorithms are commercial, meaning companies sell them to hospitals and insurers. Their job is to identify patients. And the purpose is to help: patients picked out by the algorithm get extra attention, such as a dedicated care team.

So a high score is a good thing to receive, and the algorithm is deciding who gets a benefit. The question to carry forward is whether the patients it selects are the ones with the greatest need.

The second sentence says the algorithm studied is "typical of this industry-wide approach" and affects "millions of patients." The authors are presenting it as representative of a class.

Step 2: Read the finding that at the same score Black patients were sicker

The second sentence continues: the algorithm "exhibits significant racial bias: At a given risk score, Black patients are considerably sicker than White patients, as evidenced by signs of uncontrolled illnesses."

The phrase to read slowly is "at a given risk score." The comparison is between patients the algorithm rated the same. If the score measured health need, two patients with the same score should be about equally sick. The authors found they weren't. Among patients with equal scores, the Black patients had more illness.

Put the other way round, a Black patient had to be sicker than a white patient to receive the same score. The same level of sickness earned a lower score and a smaller chance of extra help.

The authors measured sickness directly from medical records, using signs of illnesses that were not under control. That gave them a yardstick independent of the score.

Step 3: Read the cause the authors identify

The fourth sentence gives the cause: "The bias arises because the algorithm predicts health care costs rather than illness."

The builders wanted to find patients with high health needs. Need isn't recorded in a database as a number. Cost is. Every patient has a billing history, and people who are sicker generally cost more to treat. So the algorithm was trained to predict next year's cost, and cost stood in for need.

A measurable substitute of this kind is called a proxy. The fifth sentence concedes that the proxy looked reasonable: health care cost appears "to be an effective proxy for health by some measures of predictive accuracy." The algorithm predicted costs well. The trouble lay in what cost left out.

Step 4: Follow the reasoning from unequal access to unequal spending to unequal scores

The fourth sentence continues: "but unequal access to care means that we spend less money caring for Black patients than for White patients."

The reasoning runs in a line.

  • Black patients, on average, have had less access to care.
  • Less access means fewer appointments, tests, and procedures, so less money is spent on a Black patient than on a white patient who is equally sick.
  • The algorithm learns from spending. It sees a lower bill and predicts a lower future bill.
  • A lower predicted cost is a lower score, and a lower score means the patient is less likely to be selected for help.

The algorithm reproduced an inequality that already existed in the health system. Its cost predictions were accurate, and that accuracy is how the bias got in.

Step 5: Read what changing the predicted quantity did

The third sentence states the size of the effect: "Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%."

This is a calculation of what would happen if patients were selected by how sick they were. The share of Black patients among those selected would rise from under a fifth to nearly half.

The body of the article reports a practical test. The authors worked with the algorithm's manufacturer on a version that predicted a combination of health and cost. They report that this produced "an 84% reduction in bias" by their measure. The data and the patients stayed the same, and the target changed.

The sixth sentence draws the general lesson: "the choice of convenient, seemingly effective proxies for ground truth can be an important source of algorithmic bias in many contexts." Ground truth means the real quantity you care about, as opposed to the stand-in you can measure.

Key Considerations

The algorithm didn't use race as an input. The body of the article says so: "the algorithm specifically excludes race." The bias came through cost, which was correlated with race because of unequal access to care. Leaving a characteristic out of the inputs doesn't stop a system from producing different results for groups defined by it.

The common mistake is to assume that bias requires a biased variable or biased intent. This case had neither. Every input was a routine billing or medical record, and the builders were trying to direct help to patients who needed it. The bias came from one design decision, the choice of what to predict, that seemed sensible and was never examined by group.

A related mistake is to treat accuracy as proof of fairness. The algorithm was accurate at its stated task of predicting cost. Whether it was accurate and whether it was predicting the right thing were separate matters.

Two limits on the study should be kept in view. It examined one algorithm at one health system, though the authors argue it is typical. And the manufacturer cooperated in testing a fix, which the abstract doesn't mention and which doesn't always happen.

Summary

The abstract traces an unequal outcome to a single design choice. In plain words, the causal chain has four steps:

  1. The builders wanted to find the patients with the greatest health needs, and chose to predict future health care cost as a stand-in for need.
  2. Because of unequal access to care, less money was spent on Black patients than on equally sick white patients.
  3. The algorithm, trained on spending, gave Black patients lower scores than equally sick white patients.
  4. Lower scores meant that fewer Black patients were selected for extra help: 17.7 percent of those selected, where selection by illness would have made it 46.5 percent.

References

  • Obermeyer, Ziad, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019. "Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations." Science 366 (6464): 447–453.

Report an issue with this item

Guided Conversation 12 min

Choose a Fairness Standard for One Decision

In this conversation you'll name a decision that could be handed to an automated system, then hear two competing standards of fairness argued for it, each at full strength. You'll leave with a statement of which standard you'd choose and who you think should make that choice.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Choose a Fairness Standard for One Decision (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a role play with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying how bias enters AI systems and the competing standards of fairness used to judge them. Follow this guidance for the whole conversation.

GOAL
I can explain how bias enters AI systems and compare the standards of fairness used to judge them, applied to a decision I know.

ROLE
You will play two advocates in turn, then step out of role.
- First, an advocate of calibration: a score should mean the same thing for every group, so the decision-maker can rely on it without knowing a person's group. Argue it at full strength, in the first person, as its advocates would. Never a caricature.
- Second, an advocate of error-rate balance: wrongful "yes" and wrongful "no" decisions should fall equally on every group, so no group bears more of the system's mistakes. Equal force, equal length.
- Announce each change clearly: "In role as an advocate of...", and later "Stepping out of role."
- What makes this hard: I will push back, and you must answer in role without drifting to a middle position or making one advocate weaker.
- In role, concede what your standard gives up.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for anything confidential or personal, and remind me not to share any if I start to.
- Aim for about 12 minutes: four on each advocate and four out of role. If my replies are brief, offer one concrete prompt, such as "Think of how a lender decides who gets a loan, or how an employer screens applications," and move on. If I seem uncertain, shorten to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; I don't have to be persuaded; I can ask you to clarify anything. Then ask me to name one decision I know that could be automated, what the system would predict, and who would be affected.

TOPICS, IN ORDER
1. Calibration advocate. In role, argue that calibration is the right standard for my decision. Use the wrongful "yes" and wrongful "no" in my case. Ask for my strongest objection and answer it in role.
2. Error-rate balance advocate. Change roles and argue the competing standard the same way. Ask for my objection and answer it in role.
3. Out of role. Ask me who bears the errors under each standard in my decision. Draw out that when groups differ in how often the predicted outcome occurs, the two standards can't both be met, so choosing one means accepting unequal results on the other.
4. Closing. Ask me which standard I'd choose for this decision and who should make that choice: the builder, the organization using the system, a regulator, or the people affected. Tell me I can take my answer into a short optional journal entry.

KEY POINTS TO KEEP ACCURATE
- Bias can enter a system without intent and without a protected characteristic among its inputs. Routes include skewed training data, records of unequal past decisions, and a proxy: predicting something measurable in place of what matters. A 2019 study found a health algorithm that predicted cost in place of illness scored Black patients lower than equally sick white patients, though it didn't use race.
- A 2016 investigation of the COMPAS risk score found that Black defendants who didn't reoffend were labeled higher risk nearly twice as often as white defendants who didn't. The vendor replied that a given score meant the same for both groups. Both were accurate.
- A 2016 proof (Kleinberg, Mullainathan, and Raghavan) showed that calibration and balanced errors can't all hold when groups' base rates differ, except with perfect prediction. A base rate is the share of a group that actually has the outcome.
- Choosing among the standards is a value choice. The mathematics shows they conflict and doesn't say which is right.
- This doesn't mean bias can't be reduced or that all systems are equally fair.
- A bias audit compares outcomes by group. It tests what it's designed to test and can miss other fairness measures and groups it didn't check.

MISCONCEPTIONS TO CORRECT GENTLY
Correct these out of role, briefly, then return to the conversation.
- "Remove race or gender from the data and it's fair": other inputs can act as proxies and carry the same effect.
- "An accurate system is a fair one": a system can be accurate overall while its errors fall unequally.
- "The math tells us which definition is right": the math tells us they conflict.

LIMITS
- No legal advice. If I ask what the law requires, tell me to check the rules that apply to me.
- Make no judgments about specific real organizations. If I name one, discuss the type of decision.
- Outside the roles, give no view of your own on which standard is right.
- Don't favor or disparage any company, including the one that built you.

TO FINISH
After my closing answer, close in one short turn, out of role:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: write a journal entry on which error I'd rather risk; ask whether an organization I deal with audits its automated tools; work through a small table of made-up numbers for my decision.
- Restate my choice on its own lines, labeled "My standard, and who should choose it", so I can copy it.

Report an issue with this item

Journal 15 minOptional

Which Error Would You Rather Risk?

Overview

You'll take one automated decision and write about the two ways it can go wrong, who pays for each, and which standard of fairness you'd hold it to. Writing it out for a single decision shows that picking a standard means deciding whose mistakes matter more.

The entry is optional. It's for you, and nobody collects it.

Writing Prompt

For one automated decision, describe the two kinds of error it could make, who bears each, and which fairness standard you'd apply. Write 250–400 words.

Steps

  1. Name the decision and what's being predicted. Choose something concrete, such as screening job applications, approving a loan, or flagging a benefit claim for review. Say what the system is trying to predict: who will do the job well, who will repay, which claims are mistaken.
  2. Describe a wrongful "yes" and a wrongful "no" and who's harmed by each. A wrongful "yes" is a person the system approves or flags who shouldn't have been. A wrongful "no" is the reverse. For each, say who bears the cost: the person, the organization, or someone else.
  3. State two fairness standards and what each would require. One standard is that a score means the same for every group. Another is that each kind of error falls equally on every group. Say in a sentence what each would demand of the system in your decision.
  4. Choose one and defend it. Give your reason. A good reason says something about whose errors matter more in this decision, and why. You can draw on any notes of your own.

Self-Check

Before you finish, check that your entry:

  • Names what the system predicts
  • Describes both kinds of error and who bears each
  • States two fairness standards accurately
  • Chooses one and gives a reason that isn't purely technical

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Bias, fairness, and automated decisions

This ungraded knowledge check assesses your understanding of how bias enters AI systems and how fairness is judged. You'll be asked about sources of bias, the COMPAS findings and the vendor's reply, competing definitions of fairness, and audits and notice for automated decisions.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item