FINANCE, POLICY & MARKETSPublished by Paul Ivinskas
fc.The Financial CurrentDAILY INTELLIGENCEWhat matters across finance
Deep-dive library

HMDA mortgage data: what application records reveal and what denial rates cannot prove

6 min read · estimatedAI-generated analysis · Methodology
Current version · 1 version · Publication details

First published . This version published .

Initial full research explaining the mechanism, current regulatory context, customer and business consequences, worked hypothetical examples, competing interpretations and limitations. Primary sources checked October 3, 2026 (America/Denver).

Related research, policy & entities ↓

At a glance

Excerpts from this version
What it covers
HMDA makes mortgage applications and outcomes visible, but a public record is not a complete underwriting file. Denial-rate comparisons depend on product mix, denominators, reporting coverage and privacy modifications.
Product mix can change the apparent comparison
Geography, occupancy, loan purpose and product structure can all motivate meaningful comparisons. Overly narrow segmentation creates another problem: small cells become unstable and some combinations disappear entirely. A detailed chart is not necessarily a better comparison if it selects a tiny, unrepresentative remainder or conceals how many records were excluded.Read in context
0% through article

Tap a dotted-underlined term for a definition; terms are highlighted once per section. Use Aa in the navigation for reading preferences.

In this article

A public map of lending activity

The Home Mortgage Disclosure Act creates a structured view of mortgage activity that would otherwise remain dispersed among lenders. The CFPB describes its uses as showing whether financial institutions serve housing needs, informing public investment and helping identify possible discriminatory patterns. These are investigative and descriptive purposes; an observed difference is the beginning of an explanation, rather than a complete explanation in itself. [1]

The distinction matters because a large dataset can look more conclusive than it is. Millions of records improve precision for many descriptive questions, but size does not supply variables that were never collected, restore details intentionally withheld or establish why a particular decision occurred. Statistical confidence about a measured gap and confidence about its cause are different claims.

HMDA is especially valuable when the question is carefully matched to the records. It can illuminate the location and composition of reported mortgage activity. It cannot directly observe every household that considered borrowing but never applied, every marketing impression or every conversation before a reportable application existed.

Applications, originations and purchases are different events

Regulation C distinguishes action outcomes including originated loans, applications approved but not accepted, denials, withdrawals and files closed for incompleteness. Purchased loans have their own reporting treatment, as do specified preapproval requests. Reported variables also describe purposes, products, property characteristics and applicants. [2]

An origination is a completed extension of credit; a purchase reflects acquisition of a loan already made. Adding them together might answer a question about reported business activity, but it would not count newly financed households without potential double counting. One loan can appear in different institutions’ records at different stages, while one household can generate multiple applications during shopping.

Withdrawals are not automatically denials in disguise. A household may find a different lender, abandon the property purchase or change its financing plan. Equally, a withdrawal does not prove an absence of friction. The record’s category describes the reported event. Establishing the circumstances behind that event may require a file-level investigation or other information.

The denominator can change the headline

Consider a hypothetical lender with 700 originations, 100 approvals not accepted, 200 denials, 150 withdrawals and 50 incomplete files. A denial rate among applications receiving an approval or denial decision is 200 divided by 1,000, or 20%. Dividing the same denials by all 1,200 records instead produces 16.7%.

Neither arithmetic operation is inherently impossible, but they answer different questions. The first focuses on decided applications under the stated convention. The second includes applicants whose records ended without one of those decisions. A comparison becomes misleading when one institution is measured by the first convention and another by the second while both numbers carry the same label.

Now suppose 200 purchased loans are added to the file. They do not represent 200 additional approval decisions by this lender on new applications. Including them mechanically in the denominator would lower the apparent rejection share again without any borrower’s original decision changing. A well-specified metric therefore contains a population definition as well as a percentage.

Product mix can change the apparent comparison

Imagine two hypothetical lenders serving two product segments. In a lower-denial segment, each denies 5% of decided applications. In a higher-denial segment, each denies 25%. Lender A receives 800 lower-denial and 200 higher-denial applications, producing 40 plus 50 denials, or 9%. Lender B receives 200 and 800, producing 10 plus 200 denials, or 21%.

The 12-percentage-point aggregate difference exists even though both lenders have identical rates inside each segment. This is a composition effect. It does not prove that the lenders are equally situated in every relevant respect; the simplified example deliberately holds within-segment rates constant to isolate the arithmetic.

Geography, occupancy, loan purpose and product structure can all motivate meaningful comparisons. Overly narrow segmentation creates another problem: small cells become unstable and some combinations disappear entirely. A detailed chart is not necessarily a better comparison if it selects a tiny, unrepresentative remainder or conceals how many records were excluded.

The public file is deliberately different from the reported file

Regulation C assigns disclosure and public-access responsibilities and provides for privacy modification. The CFPB’s disclosure guidance excludes fields such as property address and applicant credit score from public loan-level data and reduces precision for selected other fields. The official publication platform identifies released datasets as modified for privacy. [3][4][5]

This design balances transparency against the risk that a public record could be linked to an identifiable person. A mortgage involves a property, approximate transaction amount, geography and applicant characteristics; combinations can be revealing even without a name. Privacy protection therefore affects the analytical resolution of the dataset, not merely the removal of a single identifier.

For example, two applicants who appear similar in public data may differ materially in a withheld credit measure. That limitation does not establish that an observed gap is justified. It means the public data alone cannot conclusively resolve the competing explanations. Regulators’ access to more detailed information is also not the same as an automatic finding that any disparity violates law.

Missing values carry information about the reporting process

The CFPB’s HMDA FAQs explain that specified underwriting measures must be reported when relied upon in a decision, even if they were not the decisive factor. [6] Consequently, a field’s availability can depend on the transaction and decision process, not just on whether a data engineer successfully collected it.

In a hypothetical model, replacing every absent ratio with zero would make an unavailable measurement look like a particularly favorable one. Deleting every incomplete row could instead concentrate the sample in institutions or products with fuller reporting. Either transformation changes the meaning of the analyzed population. Missingness is an analytical issue before it is a software-cleaning task.

Identifiers and reporting periods create similar challenges. A lender’s name can change, organizations can combine, and a current-year file can reflect a different business perimeter from the previous year. The apparent growth of one label may combine organic growth with institutional change. Linking records requires a view of the reporting entity and time period, rather than a simple text match.

Data vintages and plausible explanations

A public dataset is also a dated release. The publication system distinguishes available releases, and a later version can incorporate corrections. [5] A reproducible comparison therefore depends on a fixed and stated filters. Two researchers can obtain different counts without either making an arithmetic error if they retrieved different versions or applied different inclusion rules.

The broader economic interpretation requires similar care. A low denial rate could reflect accommodating underwriting, applicants already screened through another channel, a narrow product offering or a customer base with stronger measured profiles. A high rate could reflect broader access to applications, a weak applicant pool, restrictive decisions or process problems. The aggregate statistic does not select among those explanations.

HMDA’s contribution is substantial precisely because it makes such questions testable. It supplies a common starting point for comparing mortgage access and activity, while preserving the distinction between a reported outcome, an analytical association and an established causal or legal conclusion. The examples here are invented and do not characterize any institution or demographic group.

Sources

  1. CFPB, Home Mortgage Disclosure Act data hub and purposesOfficial sourceBack to text: ↑
  2. CFPB, Regulation C §1003.4, reportable data and official interpretationsOfficial textBack to text: ↑
  3. CFPB, Regulation C §1003.5, disclosure and reportingOfficial textBack to text: ↑
  4. CFPB, Disclosure of Loan-Level HMDA Data, final policy guidance, December 2018Official source · PDFBack to text: ↑
  5. FFIEC/CFPB, Snapshot National Loan-Level Dataset publication pageOfficial sourceBack to text: ↑1↑2
  6. CFPB, Home Mortgage Disclosure Act FAQs, multiple-data-point reportingOfficial sourceBack to text: ↑

Flag an error or suggest a correction →Public corrections log →