FINANCE, POLICY & MARKETSPublished by Paul Ivinskas
fc.The Financial CurrentDAILY INTELLIGENCEWhat matters across finance
Deep-dive library

Credit-model fairness: competing metrics, base rates and the limits of a single score

6 min read · estimatedAI-generated analysis · Methodology
Current version · 1 version · Publication details

First published . This version published .

Initial full article. Primary sources checked October 4, 2026; historical research retains its dates, and numerical illustrations are hypothetical.

Related research, policy & entities ↓

At a glance

Excerpts from this version
What it covers
Fairness measures ask different questions about a credit decision. Calibration, error rates and approval rates can disagree even when calculated correctly, and no statistical metric by itself establishes legal compliance.
0% through article

Tap a dotted-underlined term for a definition; terms are highlighted once per section. Use Aa in the navigation for reading preferences.

In this article

A fairness claim needs a denominator

A model can be described as fair because equally scored applicants repay at similar rates, because repayment-capable applicants are rejected at similar rates, or because groups receive similar shares of approvals. Those statements describe different comparisons. Treating them as synonyms makes an argument appear more settled than its evidence supports. This article concerns statistical measurement. It does not state a current legal test under the Equal Credit Opportunity Act or determine whether any lending practice is lawful.

The distinction matters even before a sophisticated algorithm enters the picture. An ordinary cutoff applied to a spreadsheet produces approvals, rejections and errors. A highly complex model does the same after producing more detailed estimates. More computing power can improve predictions, but it cannot remove a contradiction between definitions. The key is to make the comparison explicit enough that a reader can reproduce it and understand whose experience it describes.

Scores, decisions and observed outcomes

Calibration within groups means that a stated probability has the corresponding outcome frequency within each group. If a score is expressed as a 10% default probability, approximately one in ten people in that score category would default over the defined horizon in a well-calibrated population. Kleinberg, Mullainathan and Raghavan's 2016 research analyzes calibration alongside balance in average scores among people with positive and negative outcomes. Their theorem shows that all three conditions generally cannot coexist unless prediction is perfect or group base rates are equal. The balance conditions concern scores, not simply rejection rates at one cutoff. [1]

A score is not yet a decision. A lender could use the same predicted probability for different products, maturities or loan amounts, producing different expected losses. For the numerical illustration below, all applicants are assumed to seek the same product on the same terms. Default is designated the positive outcome, and a high-risk classification leads to rejection. This convention is important: in another paper, positive might mean approval or repayment, reversing the everyday interpretation of a .

Hypothetical confusion matrices: equal error rates, different meanings

Assume two invented groups of 1,000 applicants whose repayment outcomes are known under identical loan terms. Group A contains 100 people who would default and 900 who would repay. Group B contains 200 who would default and 800 who would repay. These are assumed base rates, not estimates about any real demographic group. The example deliberately assumes complete outcome knowledge so the arithmetic is separate from the practical problem of missing outcomes for rejected applications.

Suppose the model flags 80% of eventual defaulters and mistakenly flags 10% of eventual repayers in both groups. In A, it flags 80 defaulters and 90 repayers, leaving 20 defaulters and 810 repayers unflagged. In B, it flags 160 defaulters and 80 repayers, leaving 40 defaulters and 720 repayers unflagged. Both groups have a 10% rate and a 20% false-negative rate. Those two error-rate comparisons are exactly equal.

Yet the flagged populations look different. In A, 80 of 170 flagged people default, about 47.1%. In B, 160 of 240 flagged people default, about 66.7%. The same high-risk label therefore has different predictive values. Approval rates also differ: 830 of 1,000 in A, or 83%, and 760 of 1,000 in B, or 76%. Equal error rates neither produce equal selection rates nor make the flagged category equally predictive.

The reason is multiplication rather than mysterious algorithmic behavior. A 10% false-positive rate operates on 900 repayers in A and 800 in B. An 80% detection rate operates on 100 defaulters in A and 200 in B. Combining those different underlying counts changes the composition of the flagged pool. A percentage divorced from its denominator conceals this mechanism. Chouldechova's 2016 research develops the relationship between predictive parity, unequal outcome prevalence and unequal error rates in recidivism prediction; applying the mathematical distinction here does not transfer that setting's legal or social conclusions to lending. [2]

Moving the cutoff changes more than one statistic

Continue the invented example and suppose an alternative threshold in A flags 80 defaulters but only 40 repayers. The default share among those flagged becomes 80 divided by 120, or 66.7%, matching B. A's rate becomes 40 divided by 900, or 4.4%, while B remains at 10%. Its approval rate rises to 88%, compared with B's 76%. Matching the predictive value has not matched the other comparisons. This is an assumed alternative classification, not a claim that every real model offers that exact threshold.

The illustration also separates a risk label from a calibrated probability. Calling the entire flagged category a 66.7% risk category says something about its aggregate composition. It does not establish that every finer score within the category is calibrated, nor that unflagged applicants share equal risk. A model may contain useful rankings inside each category that disappear when the published analysis compresses everything into approved and rejected.

Selection parity alone is similarly incomplete. Imagine approving 800 applicants from each group without specifying who receives those approvals. The approval percentages match, but the selected populations could contain very different numbers of eventual repayers. Conversely, two groups can have different approval rates because their observed application pools differ. Neither observation, by itself, explains how those pools arose or settles whether their treatment is acceptable.

Labels and history are part of the measurement

The numbers above assume a clean and fully observed target. Actual credit analysis faces a more difficult counterfactual: what would a rejected applicant have done if offered the same loan? Repayment under an expensive short-term contract is not automatically the same outcome as repayment under a cheaper long-term one. A default measure after six months also differs from one after three years. Changing price, horizon or servicing treatment changes the event the model is trying to predict.

Group information introduces another layer. Directly recorded group membership, self-identification and an inferred proxy are different kinds of evidence. In an invented audit of 100 people, assigning ten people to the wrong group can move both the numerator and denominator of a subgroup error rate. A proxy-generated percentage therefore carries classification uncertainty in addition to ordinary sampling uncertainty. Removing an explicit group field from prediction does not establish that all remaining variables are unrelated to group membership.

Historical outcomes describe a historical system. If earlier access, contract terms or collections practices helped shape those outcomes, using the resulting base rates as mathematical inputs does not explain or justify that history. The impossibility result is conditional on the quantities supplied to it. It is not a finding that present disparities are inevitable in size, that every model is equally good, or that better data and treatment cannot improve outcomes.

What a meaningful comparison can establish

A useful statistical account can report calibration, errors, selection and uncertainty together while keeping their meanings separate. In the hypothetical example, a reader can see exactly why one claim of equality coexists with another disparity. That visibility supports substantive discussion about access, repayment burden, mistaken exclusion and losses without hiding a choice inside the word fair.

The unresolved question is often normative or institutional rather than computational: which harm is being evaluated, under which comparison and with what evidence? Research can identify incompatibilities and quantify consequences. It cannot supply legal authorization for a group-specific rule, select society's preferred distribution of errors, or turn incomplete outcomes into observed facts. Statistical clarity narrows the disagreement; it does not make the underlying choices disappear.

Sources

  1. Kleinberg, Mullainathan and Raghavan, Inherent Trade-Offs in the Fair Determination of Risk Scores, revised November 17, 2016Technical reportBack to text: ↑
  2. Chouldechova, Fair prediction with disparate impact: A study of bias in recidivism prediction instruments, 2016 researchTechnical reportBack to text: ↑

Flag an error or suggest a correction →Public corrections log →