Invented rows can carry real information
Synthetic financial data consists of generated records designed to resemble selected features of another dataset or an assumed population. A row might contain an invented income, balance, repayment history and loan status. No actual customer needs to match the row exactly for it to reveal information learned from real customers. Conversely, a dataset can be useful for testing an interface even when its values have no statistical relationship to any real portfolio.
NIST's explanation distinguishes ordinary synthetic generation from differentially private synthetic generation. A generator typically learns a statistical model and samples new records from it; methods range from simple distributions to neural networks. Merely replacing identifying information or producing fabricated-looking rows does not establish a rigorous privacy guarantee. Differential privacy requires an appropriate mechanism rather than the synthetic label. [1]
The economic attraction is clear in an illustrative software project. Developers need realistic column names, date formats, balances and payment events to test a servicing screen. Giving them direct customer records creates an exposure that fabricated test accounts can avoid. But the same file might be unsuitable for estimating next year's losses. The word realistic can refer to formatting, plausible behavior, statistical distributions or predictive relationships. Those are separate achievements.
Matching averages can miss the risk relationship
Consider a fully hypothetical source population of 10,000 loans. Half are classified as low income and half as high income. Assume 400 defaults occur among the low-income loans and 100 among the high-income loans. The total default rate is 5%, while the subgroup rates are 8% and 2%. These categories and numbers are invented solely to show a statistical issue; they do not describe actual borrowers or imply a lending rule.
Now imagine a generator that independently assigns income category and default status. It preserves the 50–50 income split and 5% overall default rate. Its expected output has 250 defaults in each income category, giving each a 5% rate. The two headline distributions are exactly right in expectation, but their relationship has disappeared. A chart of average defaults would look convincing while a model trained to distinguish the two categories would learn the wrong pattern.
A joint distribution can preserve that particular relationship. Yet adding loan term, geography, , product and payment history creates many more combinations. The amount of evidence for any one combination can become small. NIST discusses the difficulty of preserving correlations and the accuracy tradeoff when measuring more detailed combinations under differential privacy. The choice of relationships to retain is therefore part of dataset design, not an automatic consequence of creating many rows. [1]
The same logic applies to time. A generated file could reproduce average monthly balances while failing to preserve the sequence from missed payment to arrears to default. A collections simulator needs transitions, not just snapshots. A payment reconciliation test instead may care mainly that each transaction has a unique identifier and balances reconcile. A dataset can succeed at one of those purposes and fail at the other without being internally contradictory.
Rare events expose the limits of imitation
Suppose an invented dataset contains one million accounts, but only 40 examples of a particular combination of fraud, product and recovery outcome. Generating ten million synthetic rows does not create ten times as much observed evidence about that combination. It can produce more samples from the generator's assumptions, including any errors or omissions in those assumptions. Numerical volume and independent information are different things.
If the generator averages the rare cases into a broader group, it may erase an economically important failure mechanism. If it closely reproduces them, it may expose unusual characteristics of the original records. This is a tension between usefulness and disclosure risk, not proof that every rare record is identifiable. Its significance depends on what an outside observer already knows and what the output reveals.
An original thought experiment illustrates the stakes. A file omits names but contains a uniquely large loan, an unusual origination day and a rare property category. Someone already familiar with that transaction may recognize its pattern. Changing the last digits of the amount might make the row look synthetic without defeating that inference. A privacy evaluation has to concern the information conveyed, not the aesthetic appearance of randomness.
What a differential-privacy claim actually covers
NIST SP 800-226, published in March 2025, describes how to evaluate differential-privacy guarantees, including the protected unit, privacy parameters and the surrounding trust model. Broadly, the guarantee limits how much the output distribution can change when a protected contribution changes. Smaller epsilon generally means a stronger guarantee when the other definitions and parameters are comparable. An epsilon value alone cannot be compared meaningfully across systems with different protected units or assumptions. [2]
For example, one invented system might protect a single transaction while another protects all transactions belonging to a person. A customer with 300 transactions is not represented in the same way by those definitions. Calling both systems private at the same headline parameter would conceal a material difference. This example does not calculate a formal conversion; it explains why the object of protection matters before a number can be interpreted.
NIST also distinguishes protection for an intentional release from security of the underlying raw data. Differential privacy does not make a database breach harmless, and faulty implementation can undermine a mathematical design. Repeated access to private data requires appropriate accounting across releases; post-processing an already protected output is a different situation from returning to the raw records for fresh measurements. [2]
The practical implication is bounded rather than absolute. A properly protected synthetic file can support many downstream calculations without each calculation accessing original records again. But an analyst who repeatedly requests new versions tuned to increasingly specific findings may create a different release process. The privacy properties attach to the actual mechanism and its total interactions, not to the filename or an assurance that every individual export looked innocuous.
Training on invented defaults is not validation on real outcomes
A hypothetical modeling team can use synthetic data to exercise feature pipelines, detect incompatible formats and compare whether algorithms train successfully. Those are valuable engineering results. If it then evaluates a model on another sample generated by the same generator, the evaluation mainly tests performance inside that synthetic world. Shared assumptions can make the model and test agree even when both differ from the real portfolio.
Imagine a generator that never produces a borrower with volatile income and a large emergency expense at the same time. A model may look exceptionally accurate on its synthetic test set because a difficult interaction is absent. A held-out real sample could expose that failure immediately. The example explains why evaluation on independently observed outcomes provides different evidence from evaluation on regenerated rows.
Even real holdout data need a defined population and date. A synthetic file trained on an earlier portfolio cannot silently establish performance after a change in product terms or economic conditions. Historical consistency, forward performance and privacy protection answer different questions. A strong answer to one does not substitute for evidence about the others.
Utility is specific to the intended use
Synthetic data is most intelligible when its purpose is narrow enough to test. A sample used to demonstrate a dashboard may only need plausible values and correct accounting identities. Research into loss concentration needs reliable dependence among exposures. A model-development benchmark needs relevant out-of-sample outcomes. Public release adds disclosure considerations that an isolated internal demonstration may not share.
The central tradeoff is therefore not real versus fake. It is which information the synthetic system preserves, which uncertainty it introduces and what it allows others to infer. A carefully bounded dataset can reduce unnecessary access and make experimentation easier. It becomes misleading when a convincing imitation is presented as new empirical evidence, or when the existence of invented records is treated as proof that privacy has been secured.
Sources
- NIST, Differentially Private Synthetic Data, May 3, 2021Official sourceBack to text: ↑1↑2
- NIST SP 800-226, Guidelines for Evaluating Differential Privacy Guarantees, March 2025Official source · PDFBack to text: ↑1↑2