FINANCE, POLICY & MARKETSPublished by Paul Ivinskas
fc.The Financial CurrentDAILY INTELLIGENCEWhat matters across finance
Deep-dive library

Cash-Flow Underwriting: Adoption, Performance & Risk

10 min read · estimatedAI-generated analysis · Methodology
Historical version · 9 versions · Publication details

First published . This version published .

Version history

About this historical version

Added a consent-and-coverage funnel, a prospective evaluation design separating model performance from access and take-up, and an illustrative uncertainty calculation. Retained the earlier empirical limits and worked affordability analysis.

Compare with an earlier version →
Related research, policy & entities ↓

At a glance

Excerpts from this version
What it covers
Cash-flow data can improve underwriting evidence, but coverage, timing and selection determine what a result means. This revision adds a prospective pilot design, uncertainty measures and controls for applicants who cannot connect accounts.
Controls, consent and operational costs
must connect the actual decision to understandable reasons. A vague reference to cash-flow risk does not by itself explain whether the issue was insufficient income, volatile inflows, existing obligations or missing information. Model feature importance can inform analysis, but operational notices need legal review for the applicable product and decision.Read in context
Limits of the evidence

The study excludes applicants rejected on every application during the sampled period, so it cannot observe how all rejected applicants would have performed. The report also discusses sample and generalizability limits. Its appendix compares logistic regression and XGBoost across different information sets, providing a useful way to separate the contribution of data from that of model architecture. Neither method makes an affordability determination automatic. [2][3]Read in context

0% through article

Tap a dotted-underlined term for a definition; terms are highlighted once per section. Use Aa in the navigation for reading preferences.

In this article

A bank statement can answer several different questions

uses transaction and balance information to assess repayment capacity or risk. It can reveal recurring income, essential expenses, volatility and buffers that a traditional credit file does not fully capture. But observed inflows are not automatically income, and successful collection is not the same thing as affordable repayment. Those distinctions are especially important for gig workers, seasonal earners and households moving money among several accounts.

Federal banking agencies and the CFPB acknowledged potential benefits and compliance considerations in their December 2019 alternative-data statement. [1] FinRegLab's July 2025 empirical research examines machine learning and cash-flow information in consumer underwriting. [2][3] These sources support evaluating the approach; they do not establish that every vendor model, product or borrower segment will experience the same improvement. This article makes no universal approval-lift or loss-reduction claim.

What the empirical evidence establishes—and what it does not

FinRegLab’s July 1, 2025 study used anonymized bureau and aggregator records linked to new credit accounts opened in 2018–2019. Its hybrid machine-learning model was the strongest overall predictor among the tested alternatives and produced comparatively favorable simulated access results at many risk thresholds. These are historical modeling results, not a live randomized rollout or evidence that today’s particular vendor will reproduce them. The project page identifies support from JPMorgan Chase and Capital One. That funding context belongs alongside the results. [2][4]

The study excludes applicants rejected on every application during the sampled period, so it cannot observe how all rejected applicants would have performed. The report also discusses sample and generalizability limits. Its appendix compares logistic regression and XGBoost across different information sets, providing a useful way to separate the contribution of data from that of model architecture. Neither method makes an affordability determination automatic. [2][3]

For a prospective deployment, define the outcome before examining performance. A score predicting serious answers a different question from a budget test asking whether repayment leaves enough money for essential expenses. Also track hardship, repeated overdrafts, payment reversals and complaints where lawfully available. A low default rate achieved by collecting ahead of rent is not, by itself, evidence of a better consumer outcome.

Reconstruct income before estimating capacity

A useful pipeline identifies the account owner, observation period, missing intervals and transaction source. It then distinguishes wages and benefits from transfers, loan proceeds, refunds and reimbursements. Counting an advance as recurring earnings can create a feedback loop in which borrowing appears to improve affordability. Counting a transfer twice can produce a similar error across linked accounts.

Classification uncertainty should remain visible. A payment platform deposit may combine sales, reimbursements and transfers. A business owner's gross receipts may precede substantial operating expenses and taxes. A lender should not silently treat every ambiguous credit as disposable household income. Conservative treatment, documentation requests or a transparent manual review can be preferable to false precision, depending on the product and stakes.

Observation coverage also affects interpretation. A single account may capture payroll but omit rent paid elsewhere, or show spending while income arrives in another account. A new account's short history can make a stable household appear volatile. Conversely, a long history can hide a recent job loss if the model gives too much weight to older months. Coverage and recency belong in both the decision and its explanation.

Build features that preserve the meaning of money

Illustrative monthly statement: $6,000 of credits consists of $3,600 net payroll, $1,200 transferred from the applicant’s own savings, $800 of new borrowing and a $400 merchant refund. Treating all credits as income overstates recurring earnings by $2,400, or two-thirds of the payroll figure. The transfer may support a buffer, but it should not simultaneously increase recurring income and available assets without reconciling both accounts.

Likewise, avoid counting a card payment and the purchases it repays as two separate consumption expenses when both the card and deposit data are linked. Debt-service cash needs still matter; the feature definition must state whether it measures consumption, contractual payments or actual cash outflow. Reversals, reimbursements and disputed transactions need explicit treatment. Store confidence flags alongside classifications so uncertain data can lead to a different process rather than a fabricated exact answer.

Scroll horizontally to see all columns.

FeatureUseful calculationTest before use
Recurring net incomeIdentified earnings after reversals, separated from transfers and financingManually reconcile a representative labeled sample
Income variabilityDispersion and low-income periods over a stated observation windowTest seasonal patterns and short histories separately
Liquidity bufferAvailable balances after known near-term obligationsCheck linked-account duplication, restricted funds and pending debits
Payment capacityCash available on relevant due dates under a stated stressDistinguish new payment from obligations already counted
Data completenessAccounts, days and fields actually observedKeep unavailable data separate from a measured zero

Worked example: average income conceals timing risk

Consider a hypothetical applicant with six monthly net-income observations of $2,000, $6,000, $2,000, $6,000, $2,000 and $6,000. The mean is $4,000. Assume essential monthly expenses of $2,500 and a proposed payment of $500. An average-based calculation shows a $1,000 monthly surplus. Yet every low-income month has a $1,000 deficit before any unexpected expense.

With a reliable $3,000 starting cash buffer and income arriving on schedule, the household may bridge those troughs. With only $200 available or delayed customer payments, the same average income can produce missed obligations. Neither assumption should be invented from the average. The analysis needs observed balances, payment timing, existing debt and the borrower's ability to access the buffer.

The example does not prescribe a regulatory affordability formula. Requirements differ across products, and a mortgage analysis has specific rules. It demonstrates an analytical distinction between a probability-of-default prediction and a budget stress. A model can rank repayment risk well while failing to show how a household manages the worst weeks of its cash cycle.

Design a test that separates data value from model value

To evaluate performance, compare a baseline using established information with a model adding cash-flow features, keeping the outcome window and population comparable. Separately test whether changing the modeling method improves results. Otherwise a reported benefit may conflate richer data with more flexible algorithms. FinRegLab's main report and technical appendix are useful methodological references, not a substitute for validation on the lender's own intended use. [2][3]

Out-of-time testing matters because income patterns, fraud behavior and economic conditions change. Evaluate thin-file applicants, irregular earners and incomplete-data cases separately where sample sizes permit. Prevent leakage from transactions recorded after the decision or from outcomes that would not have been known at origination. Document exclusions so the reported result is not driven by quietly removing difficult cases.

Selection bias is another limit. Applicants willing and able to connect an account may differ from those who cannot or decline. Outcomes observed only among approved borrowers do not automatically reveal performance for rejected applicants. An evaluation should explain these boundaries rather than translating a strong retrospective score into a claim of proven broad access gains.

Controls, consent and operational costs

Recommended controls include permission tracking, data minimization, retention limits and a fallback for connection failures. Validate account ownership and monitor changes in aggregator coverage. Retain the data and feature versions necessary to reconstruct a decision without collecting unrelated information indefinitely. A correction process should address misclassified income or missing accounts as well as conventional credit-report disputes where applicable.

must connect the actual decision to understandable reasons. A vague reference to cash-flow risk does not by itself explain whether the issue was insufficient income, volatile inflows, existing obligations or missing information. Model feature importance can inform analysis, but operational notices need legal review for the applicable product and decision.

Costs include data access, connection support, classification review, validation and handling applicants with incomplete coverage. More data can improve decisions while increasing privacy exposure and technical dependencies. A lender should evaluate net economics after those costs and after any additional manual reviews, not solely an offline discrimination statistic.

Choose the product and prove the net benefit

An installment decision commits the borrower to a fixed schedule, often beyond the available transaction history. Stress both the low-income month and the persistence of a reduced income level. Revolving credit adds uncertainty about future utilization, rates and minimum-payment rules. A low current balance does not establish that a large line is affordable if the borrower draws it after an income shock. Initial line setting and later line management therefore need different tests.

Hypothetical pilot economics: assume 10,000 applicants incur $2 each in data costs and 1,000 require an additional $8 manual review. Variable cost is $28,000. If the pilot produces 400 additional funded loans with an assumed $100 contribution per loan after expected credit loss, funding and servicing, the incremental contribution is $40,000 and the remaining benefit is $12,000 before fixed integration and validation expense. Break-even requires 280 such loans. Lower take-up or worse incremental losses can erase the benefit even if model discrimination improves.

For model governance, the current Federal Reserve reference is SR 26-2, issued April 17, 2026, which supersedes SR 11-7 and SR 21-8. Its applicability statement says the guidance is expected to be most relevant to Fed-regulated organizations above $30 billion in assets and emphasizes tailoring. Do not present a historical SR 11-7 checklist as the unchanged current standard. [5] Proportionate controls still include independent challenge, decision reconstruction, feature-version control and monitoring tied to the actual use.

Regulation B requires specific reasons for under its notification framework; saying that an applicant failed an internal standard is insufficient. [6] As an implementation test, trace a sample of notices back to the features that actually determined the outcome. Check that a missing bank connection has not been mislabeled insufficient income, and that manual overrides and alternative evidence are governed consistently. A model explanation is useful only when it faithfully describes the decision being communicated.

Evidence that would change the conclusion

Consistent out-of-time improvement, stable classifications, explainable decisions and measured outcomes across relevant applicant groups would support deployment. Fragile performance under missing data, unexplained disparities, frequent income misclassification or customer harm despite low defaults would weaken the case. The strongest conclusion is conditional: cash-flow data can add useful evidence, provided the lender demonstrates which question it answers and where the evidence stops.

Evaluate the whole applicant funnel

Recommended pilot reporting begins before account connection. Track eligible applicants, invitations, consent, successful connections, usable histories, decisions, take-up and observed outcomes. A benefit measured only among clean connected records cannot be assumed for everyone who applied.

When feasible and lawful, randomize an invitation or phased rollout among otherwise eligible applicants, retaining a safe established decision process. Analyze results according to the assigned group as well as actual usage, and explain noncompliance and missing outcomes. This is a proposed evaluation design, not evidence that a lender has run such a trial. Do not randomize away required protections or infer rejected applicants’ loan performance from nonexistent loans.

Scroll horizontally to see all columns.

StageMeasureInterpretation limit
Invitation and consentShare offered and share consentingWillingness to connect may be selective
ConnectionSuccessful and failed connectionsTechnical coverage is not financial capacity
DecisionApproval and terms by assigned groupDifferent populations can confound comparisons
Take-upFunded loans among offersApproval lift is not funded access
PerformanceMature comparable outcomesShort observation misses later losses
Customer experienceComplaints, hardship and correctionLow default alone does not prove affordability

Report uncertainty around the pilot result

Hypothetical: 10 of 500 funded loans reach a defined adverse outcome, an observed rate of 2%. A Wilson 95% interval is approximately 1.1% to 3.6%. That interval illustrates sampling uncertainty under a simple binomial model; it does not capture selection bias, correlated shocks, censoring or mismeasured outcomes.

If the expected advantage over the baseline is small, 500 loans may not distinguish improvement from noise. Set the comparison, observation window and minimum detectable effect before reading results. Do not repeatedly inspect outcomes and stop at the first favorable result without an appropriate statistical design. Report both statistical uncertainty and economically meaningful loss differences.

Missing data needs an operationally fair fallback

Recommended controls separate a declined consent request, a provider outage, an unsupported institution and a genuinely short account history. None automatically means zero income. Record the reason and offer appropriate alternative evidence under the product’s policy and applicable law.

Compare processing time, approval rates, terms and complaints for connected and fallback paths, acknowledging selection limits. Audit whether staff classify the same missing-data condition consistently. A technically stronger model can still deliver a worse product if difficult-to-connect applicants face unexplained delays or inaccurate reasons. Evaluate these frictions alongside credit performance and the earlier break-even economics.

Sources

  1. Federal Reserve and other agencies, joint alternative-data statement; December 3, 2019Official releaseBack to text: ↑
  2. FinRegLab, Advancing the Credit Ecosystem: Machine Learning & Cash Flow Data in Consumer Underwriting; July 1, 2025Source · PDFBack to text: ↑1↑2↑3↑4↑5
  3. FinRegLab, accompanying technical appendix; July 2025Source · PDFBack to text: ↑1↑2↑3↑4
  4. FinRegLab, research project and publication record; July 1, 2025SourceBack to text: ↑
  5. Federal Reserve SR 26-2, Revised Guidance on Model Risk Management; April 17, 2026; supersedes SR 11-7 and SR 21-8Official sourceBack to text: ↑
  6. CFPB, Regulation B §1002.9, notifications and specific adverse-action reasons; current text checked September 27, 2026Official textBack to text: ↑

Flag an error or suggest a correction →Public corrections log →