Status: terminated order, continuing data lesson
On November 28, 2023, the CFPB issued a against Bank of America, N.A. concerning mortgage demographic information reported under the Home Mortgage Disclosure Act. It imposed a $12 million penalty. The CFPB’s public case page reports that the order was terminated on June 5, 2025 after the bank fulfilled the order’s obligations. [1]
The termination document identifies the penalty, compliance plan, annual report and improvements to HMDA management. [2] This article treats the matter as a historical, terminated supervisory case. It does not describe the bank as presently subject to that order or infer a confidential supervisory rating from either the original action or its termination.
Mortgage data serve people beyond the reporting institution
Mortgage-market information can help public agencies, researchers and businesses understand where lending occurs and how activity differs across applicants or places. That usefulness depends on the meaning of each reported field. A code describing a customer’s choice is not interchangeable with a record that the question was never properly asked.
The customer’s interaction and the published data therefore belong to the same information chain. A reporting error can weaken later analysis even when it does not change the original loan balance. The terminated Bank of America case offers a concrete illustration without establishing that every missing field reflects discrimination or a deliberate reporting failure.
A syntactically valid answer can be false
The order describes loan officers recording that applicants did not wish to provide demographic information when the required inquiry had not been properly made. It also describes the discontinuation of monitoring and later discovery of highly unusual information-not-provided patterns. The bank consented without admitting or denying the findings except as stated for jurisdiction. [3]
The analytical distinction is between an allowed code and a truthful representation of the event. A file can pass format edits while recording an applicant refusal that never occurred. Downstream software cannot reliably recover that distinction from the code alone. The strongest control must therefore sit near the original interaction and preserve evidence of what was asked and answered.
Regulation C’s Appendix B gives instructions for collecting ethnicity, race and sex information, including differences across application methods and treatment when applicants decline to provide it. [4] The consumer’s ability to decline is not permission for staff to skip the inquiry. Nor should staff pressure a consumer to answer simply to improve a missing-data metric.
The unit of monitoring changes what becomes visible
An overall missing-information percentage averages across channels, offices, loan officers and application types. A small problematic group can disappear inside an acceptable enterprise average. Conversely, a legitimate channel difference can look suspicious if compared against an inappropriate benchmark. Monitoring should therefore compare relevant peers and examine changes over time before drawing conclusions.
The historical order gives concrete examples of concentrated information-not-provided patterns and the consequences of discontinuing a monitoring report. [3] The broader lesson is not that one fixed percentage proves misconduct. It is that unusual distributions can identify where to test the actual process. Statistical outliers are investigative leads; recordings, forms and system histories provide the more direct evidence.
Segmentation by application channel, employee tenure, product, office and entry method can reveal differences hidden by an aggregate measure. Automated feeds and manually entered records can have different error mechanisms. A sudden improvement in completion may reflect a new process, inappropriate default values or pressure on staff; each explanation has different implications.
Channel comparisons need comparable collection processes
A mortgage business might compare branch, telephone and online applications to understand service and market reach. If those channels collect information differently, apparent customer differences can partly reflect process differences. The analysis should establish which fields were requested, how responses were recorded and whether the same rules were applied.
That does not mean all channels must have identical customer populations. Different applicants may choose different routes for legitimate reasons. The useful task is to separate real differences in demand and access from artifacts created by the reporting process before drawing a commercial or policy conclusion.
A hypothetical quality-control exercise
Assume two teams each process 1,000 applications. Team A records information-not-provided for 5%; Team B records it for 45%. This difference alone does not establish a violation. Differences in application methods and applicant choices, along with source-interaction samples from both teams, help explain whether the gap reflects collection practices.
Suppose source review shows that some Team B staff never asked the demographic questions while selecting the refusal code. The remediation should address scripts, training, user-interface defaults, supervision and affected records. Simply requiring Team B to reduce its percentage can create a new incentive to fabricate answers. Quality is truthful collection and reporting, not maximum completion at any cost.
Sampling apparently complete records can reveal fabricated demographic entries that look better than honestly recorded refusals. Observed information, self-reported information and required coding conventions have distinct meanings under the applicable instructions. Retaining the original record alongside authorized corrections preserves the audit trail.
Why this matters beyond HMDA
Mortgage demographic data support public transparency and analysis of lending patterns. [1][4] If missingness is related to employee behavior or channel selection, comparisons can be distorted. A model trained on such data may learn artifacts of collection rather than meaningful borrower or market differences.
Analysis: the same control problem appears in reasons, complaint dispositions and fraud labels. A permitted reason code is only useful if it accurately reflects the decision. How the label was generated, who could override it and what incentives affected entry determine its value for automated analysis. An AI classifier that predicts unreliable labels with high accuracy can institutionalize the original weakness.
Better data support better questions, not automatic verdicts
Reliable records can identify patterns requiring more detailed review and help a lender evaluate changes to its application process. They cannot by themselves explain every pricing or approval difference. Additional information about products, applicants and decision rules may be necessary to understand the mechanism.
A stronger assessment would show that collection practices match the reported codes and remain consistent after system changes. The June 2025 termination is relevant evidence about the specific order’s status. It does not make historical reporting errors irrelevant or establish the quality of every current mortgage dataset.
Remediation should prove persistence
A source-to-report review follows data from the interaction to the loan-origination system, intermediate transformations and final regulatory submission. Reconciliation can reveal count or field changes at each stage, whether corrections propagate through the reporting process and whether a later batch feed overwrites a valid local change.
Sampled error rates, unresolved exceptions, repeat findings and the outcomes of corrective training describe different dimensions of remediation. Completion of a training course does not establish behavioral change. A later independent review can show whether the process remains accurate after heightened attention fades.
Costs and evidence that would change the view
Source-level sampling is more expensive than running automated file edits. It can require secure access to recordings, specialist reviewers and careful handling of sensitive information. That cost is justified where a field’s truth depends on a human interaction the database cannot reconstruct. Automation remains valuable for population screening and reconciliation, provided it is not mistaken for proof of truth.
New authoritative findings could change the bank-specific assessment. For process quality, sustained source-to-report accuracy would be stronger evidence than a favorable aggregate missing-data percentage. The verified legal status remains termination of the 2023 order; the practical lesson is to validate the event behind the data, not just the data’s format.
Sources
- CFPB, Bank of America HMDA action, November 28, 2023; termination status June 5, 2025Official sourceBack to text: ↑1↑2↑3
- CFPB, order terminating Bank of America consent order, filed June 5, 2025Official source · PDFBack to text: ↑
- CFPB, Bank of America consent order 2023-CFPB-0016, November 28, 2023Official source · PDFBack to text: ↑1↑2
- CFPB, Regulation C Appendix B, demographic collection instructions; reviewed September 27, 2026Official textBack to text: ↑1↑2↑3