A payment decision combines data, prediction and execution
Sardine describes device and behavioral intelligence, issuing-fraud tools and research on transaction-sequence models. Signals about a session, patterns across a cardholder’s history and a policy that turns a score into an action are separate layers. Their value depends on the actual payment workflow, not on the model architecture alone. [1][2][3]
Its AI Labs page reports held-out issuer improvements using AUC-PR, a precision-recall summary. Those are vendor research results under stated comparisons, not an equivalent percentage reduction in fraud losses or a guaranteed result for another issuer. The retained sections below explain this distinction and the proposed transfer test. [3][4]
What the foundation-model material says
Sardine AI Labs describes a payments foundation model trained on unlabeled transaction sequences and used to supply additional features to an existing issuing-fraud model. Its public page outlines tokenized transaction fields and an eight-layer transformer, with issuer-held-out evaluations. [3] This is a sequence-learning approach, not evidence that a general-purpose chatbot directly decides every authorization.
The page reports a 68% relative AUC-PR improvement for a held-out consumer-card issuer and 41% for a held-out business-card issuer. [3] These are vendor-reported experimental results. They should not be translated into equivalent percentage reductions in fraud losses, fewer false declines or a guaranteed result for another issuer.
The linked whitepaper landing page establishes that the vendor offers additional research material. [4] The public review here does not claim independent replication of the full experiment. A bank evaluating the result should obtain the precise baseline, test population, outcome definitions and implementation version rather than rely only on the headline.
AUC-PR is not ordinary accuracy
Area under the precision–recall curve measures performance across thresholds. Precision is the share of selected cases that are truly positive; recall is the share of true positives selected. In rare-event fraud, overall classification accuracy can look excellent even when most fraud is missed. AUC-PR provides a more informative view of ranking, but it still does not select the economically appropriate operating point.
Its baseline depends on event prevalence, so comparisons across datasets require care. A relative improvement from 0.10 to 0.168 is 68%; it is not an increase of 68 percentage points and does not mean 68% of fraud was newly detected. This numerical illustration explains the arithmetic and is not a reconstruction of Sardine’s undisclosed baseline values.
For an issuing decision, examine recall at a fixed false-decline rate, fraud dollars detected at a fixed review capacity and the stability of results across transaction types. AUC-PR can improve while the threshold actually used in production delivers little benefit. The bank needs both the curve and the operating-point economics.
Test transfer to the bank’s own population
Holding an issuer out of training is useful evidence about transfer, but it does not eliminate all leakage or representativeness questions. Shared customers, merchants, devices or networks can connect training and test populations. That overlap may be legitimate in a network product, but it should be disclosed so the bank understands what generalization is being demonstrated.
Recommended evaluation also holds out a later time period and tests changes in fraud patterns. Confirm that every feature was available at decision time and that labels do not depend on future information inadvertently included in the input. Separate confirmed fraud from disputes, and suspected events whose status remains unresolved.
Data from declined transactions create a further limitation: the bank may never observe whether a prevented transaction would actually have become fraud. Treat such outcomes carefully and use lawful review or experimental methods to reduce uncertainty. Otherwise, the existing policy can create labels that make a challenger appear better or worse for the wrong reason.
A hypothetical authorization tradeoff
Assume a strategy prevents an additional $100,000 of confirmed fraud in a month but also creates 2,000 additional legitimate declines. If the combined service, lost-transaction and customer-retention cost averages an assumed $30 per affected event, that cost is $60,000 before vendor and implementation expense. These are hypothetical inputs, not Sardine results or a recommended valuation of customer harm.
The remaining $40,000 is not automatically net benefit. Include review workload, repeat attempts, payment rerouting and any losses shifted to another channel. Conversely, customer protection may create benefits not captured by direct bank loss. A transparent business case shows which effects are measured, estimated or omitted.
Authorization latency also matters. A high-performing model that cannot respond reliably within the transaction path’s timing constraints may require a different architecture or use case. Test tail latency, outage behavior and data gaps under realistic traffic rather than reporting only average response time.
Governance and data boundaries
Device and network intelligence can improve context while expanding sensitive data flows. Inventory collected attributes, permitted uses, retention and cross-customer sharing. Validate that the implemented configuration matches contractual and consumer-facing representations. A consortium’s scale does not by itself establish that every contributing signal is accurate, lawful to use or relevant to the bank’s purpose.
Model and rule versions should be traceable to individual actions. The current SR 26-2 guidance supplies a risk-based supervisory reference for material models. [5] A bank should separately review predictive scoring, any generative investigation tools and deterministic authorization rules rather than treat AI as one indivisible product.
One intended purchase may generate several events
Analysis: a declined purchase can produce repeated attempts, a switch in payment method or abandonment. Evaluating every retry as an independent lost sale overstates the commercial effect; treating the eventual approval as eliminating all friction understates customer effort. Connect attempts to a defensible purchase or customer unit when measuring completion.
Distinguish risk declines from insufficient funds, technical failures and other reasons. A fraud model cannot be credited with fixing every authorization problem. Compare confirmed net loss, genuine completion and contact burden at the same intervention level, keeping recovery and dispute timing consistent.
Sequence information must arrive in time to be useful
A longer history may add context while increasing the work needed to assemble and score an event. Measure end-to-end decision time and missing-history behavior in the relevant channel, not only model inference in isolation. New customers and new issuers can supply much less context than established ones.
For the commercial case, use the fixed-friction comparison in the earlier hypothetical example and include implementation, data and operations cost. An experimental gain can justify a measured trial without establishing production readiness across every geography, card program or payment type.
What would turn the research into a business result
Confidence rises with reproducible performance on later, locally relevant data, followed by stable payment outcomes and transparent intervention costs. It weakens if the baseline, labels or unit of measurement changes between tests.
The broader opportunity is using transaction context to protect activity that customers want to complete. The evidence needs to connect the research metric to that service and its net economics.
Sources
- Sardine, Device and Behavior Intelligence; reviewed September 27, 2026; vendor claimsSourceBack to text: ↑
- Sardine, Card Issuing Fraud Intelligence; reviewed September 27, 2026; vendor claimsSourceBack to text: ↑
- Sardine AI Labs, model approach and reported evaluation results; undated page reviewed September 27, 2026SourceBack to text: ↑1↑2↑3↑4
- Sardine, Building a Foundation Risk Model, research landing page; reviewed September 27, 2026SourceBack to text: ↑1↑2
- Federal Reserve, SR 26-2, April 17, 2026Official sourceBack to text: ↑