The missing outcome is structural
A lender observes repayment on loans it makes. For applicants it declines, it does not observe repayment on the loan that was never originated. That creates selection bias when the lender trains a model only on financed customers and then applies it to the entire applicant population. The missing outcome is not a blank field that can be filled with certainty; it is a counterfactual.
The peer-reviewed study Reject inference methods in credit scoring, published in the Journal of Applied Statistics in 2021, examines assumptions behind common methods and finds no uniformly dominant approach. Research proposing deep generative methods reports improvements in particular experiments, but that does not eliminate the underlying identification problem. Method performance depends on the data, selection process and evaluation design.
Why an approved-only test can mislead
Suppose an existing policy approves only applicants above a score threshold and with verified income. The training sample contains relatively little evidence about people below that threshold or without that verification. A new model may perform well within the approved sample while extrapolating poorly outside it. Its measured accuracy does not automatically justify a broad expansion of approvals.
Selection can also involve human judgment or information that was not retained in the modeling dataset. If loan officers used an unrecorded document or policy exception, the model developer may not be able to explain why some applicants were financed. Reweighting the observed sample cannot fully repair missing knowledge about the original selection mechanism.
A hypothetical expansion decision
Imagine 10,000 applicants, of whom the old policy finances 6,000. After a year, 300 of those financed borrowers meet the lender’s default definition, a 5% rate. A proposed model identifies 1,000 previously declined applicants for approval. Their default rate on the proposed loan is unknown. Assigning them the approved sample’s 5% rate would assume the very relationship the lender needs to investigate.
Even if some later obtained credit elsewhere, that outcome may involve a different amount, price, maturity or lender policy. It is useful evidence, but not a direct observation of performance on this lender’s intended offer. The hypothetical illustrates why an approval-growth estimate should present uncertainty about the newly included population rather than treat inferred labels as measured repayment.
What common methods assume
Reweighting gives observed borrowers different importance to approximate a target population, but it needs adequate overlap and assumptions about selection. Parceling or assigning likely outcomes to rejected applicants introduces modeled labels. Semi-supervised approaches can use information about the rejected population’s characteristics, while still depending on assumptions linking those characteristics to repayment.
These techniques can be useful tools for sensitivity analysis and model development. The danger is presenting their output as newly discovered ground truth. A model trained on labels produced by an earlier model can repeat the earlier assumptions and appear more certain because the synthetic training set is larger. Additional rows do not necessarily add independent information about outcomes.
Design an honest evaluation
Recommended analysis begins by documenting the population flow: applications, approvals, accepted offers, funded accounts and accounts with sufficiently mature outcomes. Attrition at each step can create another selection mechanism. Preserve the policy and offer terms associated with the observation. A borrower who declined an expensive offer is not equivalent to a borrower denied credit.
Evaluate the existing approved population separately from the proposed expansion group. Use sensitivity ranges for uncertain outcomes and compare alternative assumptions. Where appropriate and permissible under the lender’s policy, a carefully controlled pilot can produce direct evidence from a limited expansion. It should have defined exposure limits, monitoring and a stopping rule; the analytical problem is not solved by indiscriminately funding applicants previously considered unacceptable.
Governance and practical costs
Model documentation should label observed, externally obtained and inferred outcomes distinctly. Validators need access to the assumptions and the effect of changing them. Report whether the apparent benefit comes from a better model among familiar borrowers or from an untested expansion into a new population. Those are different business decisions with different evidence requirements.
The cost of obtaining better evidence may include a slower rollout, additional data, manual review and controlled credit exposure. That cost can be justified when an ambitious growth claim rests on weak extrapolation. Conversely, an excessively narrow demand for certainty can preserve an outdated policy. The goal is to make uncertainty explicit and proportionate to the amount of credit being committed.
What would change the conclusion
Confidence should rise when later observed performance supports the model in the actual expansion population, across relevant products and economic conditions. It should fall if the result depends on one arbitrary inferred-label rule, poor population overlap or evaluation against labels generated by the model itself. Independent replication and transparent test design are more persuasive than a large simulated approval increase.
Reject inference is therefore best understood as an attempt to manage a missing-evidence problem, with assumptions that remain visible. It can improve a decision process when used carefully, but it cannot reveal with certainty what every rejected borrower would have done. Readers evaluating an AI underwriting claim should ask how much of the claimed improvement was measured and how much was inferred.