A forecast with an uncertainty layer
A cash-flow forecast of $100,000 says where a model expects an outcome to land. It does not explain how frequently errors exceed $5,000 or $50,000. Conformal prediction places an uncertainty layer around an existing forecasting model, using held-out observations to calibrate a set of plausible outcomes. It can wrap a simple regression or a much more complicated model. [1]
The central attraction is that a standard coverage result does not require a correct parametric model of the outcome distribution. But distribution-free does not mean assumption-free. Standard split conformal prediction relies on exchangeability: loosely, the calibration examples and the next example have a joint distribution that does not depend on their ordering. Financial data often challenge that premise.
How a calibration sample becomes an interval
In basic split conformal regression, one sample trains a model and a separate sample supplies absolute prediction errors. For a target coverage of 90%, the calibration threshold uses the appropriate finite-sample rank rather than simply any software package’s interpolated 90th percentile. With n calibration observations, the usual rank is the ceiling of (n + 1) times 0.90, with boundary handling if that rank exceeds n. [1]
For a hypothetical example, assume 99 exchangeable calibration cases. The chosen rank is 90. If the 90th smallest absolute error is $12,000, a new point forecast of $100,000 receives the interval $88,000 to $112,000. All figures here are invented to illustrate the arithmetic; they describe neither observed bank performance nor a commercial system.
The example separates two tasks. The point model decides the center. The calibration procedure determines the error allowance. A worse point model can still produce an interval with valid coverage, but the interval may need to become much wider. The coverage certificate alone cannot rank the economic quality of the forecasts.
What 90% coverage does and does not mean
The standard guarantee is marginal: it averages over the randomness in calibration and the next case under the stated assumptions. It is not a promise that exactly 90 of the next 100 realized forecasts will succeed. It also does not establish a 90% success rate for every cash-flow profile, every borrower or every industry. Research on the limits of distribution-free conditional inference shows why unrestricted exact individual-conditional guarantees cannot generally be obtained with informative intervals without additional assumptions. [2]
Consider 1,000 hypothetical predictions containing 900 stable businesses and 100 volatile businesses. Suppose intervals cover 855 outcomes in the first group and 45 in the second. Overall coverage is 900 out of 1,000, or 90%. Coverage is 95% for stable businesses and 45% for volatile businesses. An aggregate score would hide a severe concentration of misses.
This example is not a proof that conformal methods necessarily behave this way. It demonstrates that the aggregate metric does not rule the pattern out. Even a large calibration sample cannot make an average statement mean an individual one. Group definitions, their sample sizes and the specific guarantee all matter.
Width can adapt without solving every problem
A fixed dollar error allowance may be awkward when forecasts range from small accounts to large businesses. Conformalized quantile regression combines lower and upper quantile estimates with a calibration adjustment. Unlike a constant-width residual interval, the starting bands can reflect differences in the expected spread of outcomes. Romano, Patterson and Candès developed this approach in a 2019 research paper. [3]
In an illustrative pair of forecasts, a stable business might receive a $95,000–$105,000 band, while a seasonal business with the same central forecast receives $65,000–$135,000. Both ranges could be economically sensible if their variability differs. Equal width would not automatically mean equal quality, just as equal point predictions do not establish equal uncertainty.
Adaptive width is still distinct from exact conditional coverage. A model can learn a useful pattern in volatility without learning every relevant difference. Rare customer types may remain poorly represented. An interval that looks individualized is not, by appearance alone, a calibrated probability statement for that exact individual.
Time changes the statistical problem
Cash flows, defaults and market prices are frequently serially dependent. A rate shock, altered underwriting policy or new payment product can change how future observations relate to historical calibration cases. Randomly mixing dates may conceal that change, and repeated observations from one account may be much less informative than the same number of independent accounts.
Gibbs and Candès’ adaptive conformal inference research studies sequential prediction under distribution shift. Its central result concerns coverage frequency over long time intervals, using feedback to adjust a parameter. That is a different guarantee from standard exchangeable finite-sample marginal coverage, and neither statement promises immediate recovery after every abrupt financial shock. [4]
Imagine a forecast interval that misses sharply for three months after a business loses a major customer, then widens enough to cover nearly everything. A long-run metric may eventually recover. The cash shortage during those three months remains economically important. Reporting the recovered average alone would blur the distinction between eventual calibration and timely warning.
Delayed outcomes constrain adaptation
Feedback cannot arrive before the outcome is observed. A next-day balance forecast can be assessed quickly. A one-year default outcome may remain unresolved for months, and early is a different target. Updating intervals based on immature labels can change the question without making that change obvious.
Suppose a hypothetical model is recalibrated monthly using only loans whose one-year outcomes are complete. Its feedback is necessarily drawn from older originations. If a new product changes customer behavior, the apparently latest calibration dataset can still lag the relevant change by a year. That lag follows from the outcome definition, not merely from slow computing.
Forecast horizons also matter. Coverage for each individual month does not imply coverage for all twelve months simultaneously. If a model needs a path that remains within bounds throughout a year, that is a joint-event problem. Adding twelve separate 90% labels does not create a 90% annual-path guarantee.
Coverage and usefulness are separate evidence
A range wide enough to include nearly every possible outcome can satisfy a coverage objective while contributing little to a financing discussion. Conversely, narrow intervals may look helpful while missing precisely during stress. Coverage, width and the timing and concentration of errors describe different dimensions of performance.
Conformal prediction is therefore best understood as a disciplined statistical wrapper with a specified target and assumptions. It does not replace the point model, establish causality or turn historic errors into protection from every regime change. Its contribution is to make an uncertainty claim more explicit and testable, while leaving visible the financial questions that a coverage percentage cannot answer.
Sources
- Angelopoulos and Bates, A Gentle Introduction, arXiv v6, December 7, 2022Technical reportBack to text: ↑1↑2
- Barber, Candès, Ramdas and Tibshirani, The Limits of Distribution-Free Conditional Predictive Inference, 2019 research paperTechnical reportBack to text: ↑
- Romano, Patterson and Candès, Conformalized Quantile Regression, 2019Technical reportBack to text: ↑
- Gibbs and Candès, Adaptive Conformal Inference Under Distribution Shift, v3 October 28, 2021Technical reportBack to text: ↑