Datadory notebook

Where can I get labeled credit card fraud data? Three benchmarks, delivered as rows

Datadory delivers transaction & payment processing data covering every shelf that carries a row-level fraud label: the ULB benchmark's 284,807 European card authorizations holding 492 frauds at a 0.172% positive rate, IEEE-CIS Fraud Detection's 590,540 Vesta e-commerce transactions across 394 features plus a device-identity join, and PaySim's million-row simulated mobile-money ledger with both parties' balances - typed, keyed, field-dictionaried, and delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What does “labeled credit card fraud data” actually point at?

The query reads like an errand: find the file, take the file, train the model. The honest answer is smaller and more useful - a shelf of exactly three records carries a per-transaction fraud label, and choosing among them matters more than locating them. Requirements that brutal narrow fast: genuine outcomes, extreme scarcity, documentation solid enough to defend a result.

Around them, the Transaction & Payment Processing Services slice holds six primary records averaging 8.33 out of 10 - three 9s, two 8s, one 7 - with two pooled consumer-finance neighbors sitting one join away.

Which dataset fits which job?

Each fraud workflow maps to a specific record rather than a generic category:

What do the rows look like once delivered?

Three records, three theories of what a fraud label is - printed as they arrive:

# ULB card authorization - subset of the 31 columns, values exactly as delivered
Time : 0      V1 : -1.3598071336738      V28 : -0.0210530534538215
Amount : 149.62      Class : 0

# IEEE-CIS identity record, joined on TransactionID 2987004
id_30 : Android 7.0      id_31 : samsung browser 6.2      id_33 : 2220x1080
DeviceType : mobile      DeviceInfo : SAMSUNG SM-G892A Build/NRD90M

# PaySim simulated transfer, hour 1 of 744
step : 1      type : TRANSFER      amount : 181.0
nameOrig : C1305486145      oldbalanceOrg : 181.0      newbalanceOrig : 0.0
nameDest : C553264065      oldbalanceDest : 21182.0     newbalanceDest : 0.0
isFraud : 1      isFlaggedFraud : 0

Read the anatomy rather than the digits. The ULB row is an observed outcome behind anonymity: coordinates in a compressed feature space plus a bare verdict, with nothing recoverable about merchant, country or cardholder - which is exactly why every published method faces identical conditions. The IEEE-CIS row is a real outcome plus the adversary: a Samsung phone on Android 7.0 at a 2220x1080 resolution is the composite a rule engine cannot score but a model can. The PaySim row is scripted behavior with a full ledger, and the arithmetic closes to the cent - 181.0 leaves, the origin balance drops to zero, the destination absorbs it - so every row carries its own reconciliation test.

See what a fraud label commits you to before treating any of the three as ground truth.

How imbalanced is labeled fraud data in practice?

The number everyone quotes belongs to the ULB file: 492 positives among 284,807 rows, 0.172%. A model that predicts legitimate on every single row scores roughly 99.8 percent accuracy while catching zero fraud, which is why resampling and anomaly-detection methods keep getting benched against this specific corpus rather than a generic one.

The other two records stretch or shrink the problem deliberately. IEEE-CIS runs near 3.5% across six months ordered by TransactionDT: train on earlier periods, test on later ones, measure drift the way production actually experiences it - the test a random shuffle would erase. Set the two base rates side by side and the gap is the lesson: 0.172% against roughly 3.5% is a twenty-fold difference, and any threshold tuned on one misprices the other.

PaySim removes privacy exposure entirely because no row describes a real customer, but its labels characterize simulated agents rather than confirmed cases. Results earned there deserve the adjective when they reach risk committees or regulators - mechanism findings transfer; market-specific rates do not exist to quote.

Which fields carry the analytical weight?

ULB - four groups, 31 columns. The clock (Time, seconds elapsed since the first transaction), the survivor (Amount, the only original input left untransformed), the component block (V1 through V28, principal components of a confidentiality-preserving transformation) and the verdict (Class, 1 for fraud, 0 otherwise). Feature engineering tops out at recombinations of those groups - which is the point of a comparable benchmark.

IEEE-CIS - two tables, sixteen field families. The money block opens each transaction: amount, product code, card1-card6 issuer metadata, masked billing region and country, distance measures, purchaser versus recipient email domains. Three engineered families follow - C1-C14 counters tallying addresses, cards and phones tied to an instrument, D1-D15 timedeltas measuring recency, M1-M9 match flags testing whether party details agree. Then the famous block: V1-V339, Vesta's own ranking and counting features, documented as opaque because they are. The identity table adds id_01-id_38 plus DeviceType and DeviceInfo, on about 24% of transactions - a join rate that is itself a modeling decision.

PaySim - eleven columns, three groups. Movement (step, type, amount), the ledger group almost every extract omits (both parties' balances before and after) and the label pair: isFraud for observed fraud, isFlaggedFraud for what the threshold control caught - the rule you are evaluating beside the answer key.

Where does coverage reach - and where does it stop?

  • Geographic: ULB covers European cardholders with issuer and acquirer countries undisclosed - mechanism-level findings, never a country's fraud rate. IEEE-CIS spans global e-commerce with countries surviving only as masked codes. PaySim states no geography at all, by design: amounts behave like real ones and attach to no carrier or currency.
  • Temporal: ULB holds a fixed 48-hour window in September 2013, expressed as elapsed seconds - intra-day rhythm is recoverable, seasonality is out of scope by construction. IEEE-CIS orders six months chronologically, the drift test built in. PaySim runs 30 simulated days as 744 hourly steps, sequence intact, calendar dates absent.
  • Granularity: one row per card authorization (ULB), one row per online payment optionally enriched with device attributes (IEEE-CIS), one row per simulated transfer with dual-side balances (PaySim).

Two scope notes belong in every plan. Quoting the 0.172% rate as an industry-wide figure overstates what a two-day sample asserts. And terms live at the source level for all three originals, while Datadory delivers derived, normalized rows into your environment under our own terms - spelled out in the sample paperwork before anything ships.

What pairs with the labels when the job is measurement, not modeling?

Three records fill the edges the benchmarks leave open - none carries row labels, which is precisely their job.

Reserve Bank of India – Payment System Indicators publishes a five-part monthly edition whose Part V carries the domestic payment fraud series from September 2022 onward: monthly volume, value, a one-in-every-X-transactions rate and a basis-point ratio, beside an archive reaching to 1990. The June 2026 edition reads 3.53 lakh frauds worth Rs. 489 crore - 0.150 basis points of payment value, roughly one fraudulent transaction per 74,000 completed digital payments.

UK Finance / Pay.UK – UK Payment Markets and Payments Statistics counts processing volume and value for Bacs, CHAPS, Faster Payments and the Image Clearing System, monthly from 1990 - the closest national analog when a fraud rate needs a denominator that moves.

CFPB Consumer Complaint Database, pooled in from consumer finance, contributes roughly 17.24 million individual complaints received since December 2011 - the dispute tail that follows the fraudulent transaction, which no benchmark here covers.

For the adoption context behind any fraud headline, World Bank DataBank – Global Financial Inclusion (Global Findex Query Interface) serves 3,313 series across roughly 170 economies over survey waves 2011 through 2024; India's made-a-digital-payment share climbs 19.97% (2017) → 24.28% (2021) → 29.33% (2024), the expanding surface the fraud rates sit on top of.

Who builds on labeled fraud data?

Data scientists and ML engineers hold the shortest honest path to a defensible fraud baseline: real labels, real scarcity and a literature behind nearly every technique; the workflow ranks the slice at data scientists use cases. Payments and risk analysts read genuine card flow in which expensive mistakes separate cleanly from cheap ones, and drain chains they can follow end to end; see e-commerce operators use cases. Competitive-intelligence teams treat the benchmarks as the shared ruler - see competitive intel product teams use cases. Developers and builders stress ingestion, feature stores and monitoring on fixed schemas that never change shape mid-project; see developers builders use cases. Investors and quants benchmark a portfolio company's loss ratios against what the public labels make reproducible; see investors quants use cases. Educators, journalists and academics cite the corpus behind a large share of published fraud-detection work, every column documented.

How is labeled credit card fraud data delivered?

Name the corpus, the segment and the cut when you request a sample and the extract arrives shaped to that scope; the standing feed follows the same shape, so anything prototyped on the sample survives delivery intact. Derived conveniences used in published walkthroughs - balance-delta error columns for PaySim, elapsed-time buckets for ULB - ship with the sample rather than getting rebuilt in your notebook.

Why get the fraud benchmarks through Datadory?

Because the hard part was never the first extract - it is the tenth. Thirty-one columns of anonymity invite recombinations that need documenting. Five files and a partial join invite null patterns that surface mid-model instead of mid-contract. A simulated ledger invites consistency checks nobody wants to hand-write twice.

Datadory normalizes before delivery: field definitions verified against the files themselves, joins shipped as explicit coverage flags, derived columns included where published walkthroughs rely on them, and new versions arriving as rows under the same dictionary - no re-integration project when the next vintage lands. Supplied alongside are the sample rows and coverage profiles that make a benchmark claim auditable, which is the difference between citing a number and defending one.

Where to go next

Start with the product pages: the ULB card benchmark, the IEEE-CIS transaction and identity tables and PaySim's simulated mobile-money ledger, each with its field dictionary and a sample request shaped to your segment. For the head-to-heads, read ULB vs PaySim and IEEE-CIS fraud data vs RBI payment indicators.

Adjacent deep-dives: the sibling briefing on RBI payment system indicators Excel and ecommerce fraud detection training data. The best transaction-payment-processing-services datasets ranking puts all eight working records side by side, the transaction & payment processing data hub holds the pooled scorecard, and the Kaggle source profile shows the wider mirroring shelf these benchmarks sit on.

The measurement rail that pairs with the labels (Datadory catalog, August 2026)
DatasetWhat it countsDepthQuestion it settles
Reserve Bank of India – Payment System Indicators (Monthly Volume and Value)National volume and value per payment rail, plus a Part V domestic fraud series with one-in-X rates and basis-point ratiosMonthly editions archived to 1990; fraud series from September 2022How much fraud moves through a national rail, month by month
UK Finance / Pay.UK – UK Payment Markets and Payments StatisticsProcessing counts and values for Bacs, CHAPS, Faster Payments and the Image Clearing SystemMonthly series from 1990 to the presentHow UK clearing volumes trend beneath the fraud headlines
CFPB Consumer Complaint DatabaseIndividual complaints about U.S. financial products, with company responses and opt-in narrativesAbout 17.24 million complaints since December 2011What disputes look like after the fraudulent transaction settles
World Bank DataBank – Global Financial Inclusion (Global Findex Query Interface)Financial-inclusion percentages: digital payments, merchant payments, account ownership3,313 series across roughly 170 economies, survey waves 2011-2024Who is actually transacting digitally - the denominator under every fraud rate

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Transaction & Payment Processing Services European cardholders

Credit Card Fraud Detection (ULB) – 284,807 European Card Transactions

Time · Amount · Class

Transaction & Payment Processing Services Global e-commerce transactions with countries masked as codes

IEEE-CIS Fraud Detection – Vesta Transaction & Identity Tables

Transaction & Payment Processing Services None stated - synthetic data scaled from an aggregated sample…

Kaggle PaySim - Synthetic Mobile Money Transactions (1M+ Rows)

step · type · amount …+8 more

Transaction & Payment Processing Services India (domestic financial transactions

Reserve Bank of India - Payment System Indicators (Monthly Volume and Value)

Transaction & Payment Processing Services United Kingdom - domestic clearing across the four interbank…

UK Finance / Pay.UK – UK Payment Markets and Payments Statistics

Consumer Finance United States - consumer mailing state (50 states plus DC and…

CFPB Consumer Complaint Database

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

How imbalanced is labeled credit card fraud data?

Extremely, and unevenly. The ULB benchmark sits at 492 positives among 284,807 rows, where predicting legitimate on everything scores about 99.8% accuracy and catches no fraud at all. IEEE-CIS runs near 3.5% across six months, so a threshold tuned on one corpus misprices the other by roughly twenty-fold.

Can PaySim substitute for real transaction data?

For mechanisms, yes; for market rates, no. Calibrated on aggregated, anonymized mobile-money logs, it produces balances that reconcile arithmetically and drain chains that behave realistically - but no row describes a real customer, carrier or country. Findings transfer as mechanisms; quoting its percentages as industry rates overstates what the simulation claims.

What comes with an IEEE-CIS fraud data delivery?

Two tables keyed on TransactionID: a 590,540-row transaction table with amount, product code, card metadata, masked geographies, C1-C14 counters, D1-D15 timedeltas, M1-M9 match flags and the opaque V1-V339 block, plus a 144,233-row identity table of device and browser attributes. The roughly 24% identity join rate arrives as an explicit coverage column.