A lending decision is part of a purchase journey
Affirm’s technical account describes an internal underwriting system using learned credit-report representations with a predictive model. Its measured incremental-approval experiment is one part of the evidence. The same account says the model is online making both approval and decline decisions for new users; it describes returning-user rollout as a future step. The public experiment and reported production scope should not be treated as identical. [2]
Analysis: the merchant is interested in completed, retained purchases and net contribution after financing and service costs. The borrower needs understandable terms, a usable repayment schedule and correct treatment of returns. The lender needs sustainable revenue after funding, fraud, servicing and losses. A conversion increase is relevant to all three, but does not establish that all three benefit equally.
This is a useful case for financial-technology readers beyond underwriting specialists. It connects a technical change to distribution and customer activity while showing why experimental boundaries matter. The company’s architecture and results remain company-authored evidence, rather than an independently reproduced performance estimate.
What the company has disclosed
Affirm’s September 17, 2026 announcement describes a transformer-based system in U.S. checkout underwriting and reports 3.4% more completed purchases in a controlled comparison. These are company-reported results. The system is an internal lending capability; the cited materials do not establish a generally available product, customer API or public licensing price.
The technical account describes learned representations of credit-report records feeding an XGBoost risk model alongside established features. This is predictive credit modeling, not evidence that a conversational chatbot makes the lending decision.
Read the metrics precisely
Affirm’s technical account reports a 1.2 percentage-point conversion increase, equivalent to 3.4% relative. It also describes 1.8 times the improvement in a specified risk-ranking metric versus its next conventional model candidate. That means a ratio of improvements against a common baseline—not 1.8 times the model accuracy or an 80% reduction in losses.
Analysis: conversion combines approval and customer completion. A ranking metric measures separation of outcomes under a defined evaluation, while profitability depends on policy, pricing and losses. Preserve the denominator, eligible population, outcome horizon and experimental restrictions for each claim.
The experiment and production deployment have different scopes
The published incremental-approval experiment allowed additional approvals without allowing the new model to reject applications the existing system approved. That is a narrower use than replacing both approval and decline decisions.
Analysis: a bank considering similar technology needs separate evidence for the policy it intends to deploy. A restricted rollout can control exposure while gathering information. It does not automatically validate a different cutoff, product duration, merchant mix or borrower population. Compare the incremental cohort with a credible control and allow repayment outcomes to mature.
The technical account separately reports first-look production use for new users, including both approvals and declines. Its incremental-approval experiment remains narrower than that deployment: the cited conversion result is not a complete test of all production decisions. [2]
Illustrative economics of the incremental approvals
Assume a hypothetical rollout adds 1,000 funded loans of $1,000 each. If revenue before funding, servicing, fraud and credit losses is $100 per loan, that creates $100,000 of revenue. If funding and servicing cost $30,000, $70,000 remains before fraud, losses and capital costs.
A 5% lifetime net credit-loss rate on the $1 million originated would consume $50,000; 8% would consume $80,000. The contribution before fraud and capital would move from positive $20,000 to negative $10,000. These are simplified assumptions, not Affirm results. Timing, amortization, recoveries and funding structure matter in a full cash-flow model.
Evidence needed beyond predictive performance
NIST describes its AI Risk Management Framework as voluntary guidance for managing AI risks across design, use and evaluation. Applying that lens here means deciding what constitutes an acceptable lending outcome before selecting a model. The following is an analytical evaluation plan.
Scroll horizontally to see all columns.
| Dimension | Question to test |
|---|---|
| Data integrity | Can the exact information available at decision time be reconstructed? |
| Credit outcomes | Do mature losses and calibration remain acceptable by cohort? |
| Fairness and access | How do decisions and errors vary across relevant populations? |
| Reliability | What happens when data, infrastructure or model outputs fail? |
| Change control | Can an approved version be identified, monitored and rolled back? |
Conversion should be followed through returns and costs
Hypothetical: a checkout change produces 100 additional purchases at $500 each, or $50,000 of sales. If 15% is returned or canceled, retained sales are $42,500. At an assumed 30% merchandise margin, gross contribution is $12,750 before financing, fulfillment, support and other costs. These are illustrative merchant economics, not Affirm results or fees.
The merchant should compare that contribution with the incremental cost and identify how much activity was genuinely additional rather than shifted from another payment option. A larger checkout conversion rate does not establish the retained sales benefit without that comparison.
For the borrower, completion is the start of the obligation, so later repayment and return handling remain part of the outcome. For the lender, the incremental-cohort example above shows how relatively small differences in losses can change contribution. The broader business question is whether a better prediction produces a durable service improvement across the purchase and repayment journey.
Explainability and what would change the assessment
Affirm says it developed a proprietary explanation method. That claim does not independently verify notice accuracy. Regulation B’s notification framework remains relevant to the actual decision and the specific reasons provided, regardless of model branding.
Analysis: confidence would increase with independently reviewed evidence, stable results on later cohorts, clear explanation testing and complete operational costs. It would decrease if gains vanish after cohort normalization, exception handling overwhelms savings, or the decision cannot be reproduced. Public evidence supports a promising company case study; it does not establish equivalent results for another lender.