Genory
Testing

Synthetic Data vs Fake Data: What Your Tests Actually Need

Choose test inputs by the behavior they need to expose. Compare placeholders, synthetic records, mocks and fixtures, with practical examples and clear limits.

A customer card can render perfectly with a name like “Test User.” The same record can fail immediately when a checkout service tries to calculate delivery costs from its country and postal code. Neither result tells you whether “fake data” is good or bad. It tells you that two tests depend on different properties of their input.

The useful question in the synthetic data vs fake data debate is not which label sounds more advanced. It is which parts of the data must behave like the real system, which can remain simple, and how you will know that the generated input satisfies those requirements. This comparison focuses on application testing, rather than training machine-learning models.

Fake, dummy, mock and synthetic data are not four competing products

A developer might call a Faker-generated profile fake data, synthetic test data or mock data in different conversations. Establish your team’s meaning before choosing a tool. In particular, separate the contents of a record from the mechanism that supplies it to a test.

Term Useful working meaning Example What the label does not guarantee
Fake data Invented values used instead of operational records Test User in a preview Valid formats or representative cases
Dummy data A placeholder whose detailed value is irrelevant to this test A name passed to a function that only reads an ID Fitness for a complete user journey
Mock data Informal shorthand for substitute responses or records A fixed JSON response from a stubbed endpoint The behavior of the real dependency
Synthetic data Artificially generated values with chosen rules or modeled properties Profiles generated with linked location fields Verified identities or statistical fidelity
Test fixture The controlled data and setup a test starts from A customer plus two orders loaded before a test That the data was generated rather than handwritten

A mock is also a specific testing mechanism. For example, Python’s unittest.mock can replace a dependency and record how it was called. The JSON it returns is only one part of that behavior. A fake implementation, such as an in-memory repository, is likewise different from a fake name in a database row.

A fixture can contain handwritten placeholders, saved generated records, or a deliberately invalid case. It can also include setup and cleanup, such as a database state or temporary service. pytest’s fixture documentation describes this broader role. Treating “fixture” as a synonym for a JSON file misses the environment that makes the file usable.

Compare the properties your test needs, not the marketing labels

For this article, simple placeholders mean values selected mainly to occupy a field. Rule-based synthetic samples mean generated values constrained by an explicit model. These are two useful approaches within a much wider vocabulary; they are not an official quality ranking. A carefully written fixture can be more trustworthy than thousands of loosely generated records.

Property Simple placeholder approach Rule-based synthetic approach
Setup effort Often a few explicit values Requires a schema and generation rules
Variation Add selected cases by hand Generate variations within defined boundaries
Field relationships Write the relevant relationship explicitly Encode and validate relationships in the generator
Reproducibility Preserve the exact fixture Preserve the seed, configuration, versions and time inputs
Debugging Small examples are easy to inspect Save failing outputs and reduce them to a small example
Best reason to choose it Extra realism cannot affect the assertion Variation or relationships can expose additional failures

When simple fake data is enough

Suppose a function computes a total from quantity and unit price. If it never reads the customer’s name, a richly generated identity adds no evidence. Use explicit amounts that exercise rounding, zero quantities or your chosen boundary conditions. The same principle applies to a loading indicator, an empty table state or a button disabled by one boolean value.

For a layout test, “Test User” may be a starting point, but a single short string cannot prove that long names wrap correctly. Add a reviewed long string or Unicode case when that is the behavior under test. You still do not need an entire realistic customer history. Small, intentional examples often communicate a failure more clearly.

When realistic synthetic data adds useful coverage

More structured data becomes useful when a workflow combines several fields, displays varied records, imports many rows or joins related entities. A shipping form might choose fields by country; an account screen might combine names and email addresses; an order import might reject missing customers. Realism should model those dependencies instead of merely making individual strings look plausible.

There are also statistical synthetic datasets generated from models of source data. Those require separate evaluation of retained distributions, relationships and privacy. A rule-based profile generator does not become statistically representative merely because its outputs resemble people. Decide whether you need readable examples, domain consistency, measured distributions, or a combination.

Field relationships: four checks that attractive sample data can miss

Three example constraints connect country to a reviewed location, a fixture customer to an order reference, and quantity and unit price to a total.
Choose shared inputs once, derive related values and check the rules that matter. Plausible-looking fields alone do not establish a coherent record.

Country, city and postal code

Selecting a German country code, a French city and a US-style ZIP independently can create a record that passes three superficial string checks but fails the actual shipping flow. For a positive case, use a reviewed location tuple or a generator that explicitly supports the relationship. For a negative case, break one relationship intentionally and record the expected rejection.

Even a plausible combination such as DE, Berlin and 10115 does not establish that a particular street address exists or receives deliveries. Preserve postal codes as strings so leading zeros survive storage and CSV imports. The Address Generator can supply sample fields; its quality notes and your application’s requirements determine what you may infer from them.

Country and phone number

Use the country or numbering region required by the scenario when checking a phone’s format. Do not assume a shipping country must always match the phone’s country code: people travel and keep foreign numbers. Encode that relationship only if it is an actual product requirement. Syntactic validity does not prove that a number is assigned, reachable or safe to call.

Profile and email address

Deriving an email-shaped value from a generated profile can make a demo easier to follow. For example, Mira Example and [email protected] form a readable editorial pair. A person’s real email need not resemble their name, so keep that correlation out of business validation unless the application truly requires it.

Use domains reserved for documentation, such as example.com, for illustrative addresses, and prevent test code from sending real messages. IANA’s example-domain guidance explains their documentation purpose. That reservation is not a replacement for a mail capture service or mocked transport, and it does not promise a usable mailbox.

Structured identifiers and references

A valid UUID has a defined representation; an order’s customer ID must additionally refer to a customer that exists in the fixture. An IBAN may pass a checksum without identifying an open account. A card-shaped test number cannot establish payment authorization. Keep shape, checksum, reference integrity and real-world verification as separate checks.

Use the UUID Generator for identifier examples and the validator overview to choose the appropriate format checks. Use your payment provider’s documented sandbox fixtures when testing payment behavior. Do not expect a general-purpose generator to supply working financial credentials.

Worked example: a useful customer-and-order fixture

The following standalone Node.js example is an editorial fixture, not a Genory API response. It invents a customer, keeps one location tuple, and derives the order reference from that customer. The amounts deliberately omit tax, shipping and discounts so that the example tests one simple arithmetic rule.

import assert from 'node:assert/strict';

const customer = {
  id: 'customer-fixture-1',
  name: 'Mira Example',
  email: '[email protected]',
  shipping: { country: 'DE', city: 'Berlin', postalCode: '10115' }
};
const order = {
  id: 'order-fixture-1',
  customerId: customer.id,
  quantity: 2,
  unitPriceCents: 1250,
  totalCents: 2500,
  createdAt: '2026-01-15T12:00:00.000Z'
};

function assertOrderFixture(customer, order) {
  assert.equal(order.customerId, customer.id);
  assert.ok(Number.isInteger(order.quantity) && order.quantity > 0);
  assert.equal(order.totalCents, order.quantity * order.unitPriceCents);
}

assertOrderFixture(customer, order);

// A deliberate negative case changes only the customer reference.
const orphanOrder = { ...order, customerId: 'customer-missing' };
assert.throws(() => assertOrderFixture(customer, orphanOrder));

The check detects a broken reference rather than silently repairing it. In an integration test, submit the valid order to your application and assert the expected stored result. Submit the orphan order separately and assert your application’s defined validation response. Checking a fixture generator alone does not prove that the application enforces the same rules.

Which data should you use at each testing layer?

Choose explicit values for narrow assertions, linked records for workflows, and controlled distributions for load tests; save replay context for each approach.
Select the minimum data needed by the assertion. Add relationships or controlled volume when those properties matter; preserve replay information across every approach.
Testing goal Useful starting point Add when needed Watch out for
Unit tests Small explicit fixtures Carefully chosen boundaries or generated properties Random values that obscure the assertion
Integration tests Known related entities Constrained generated combinations Broken references mistaken for application bugs
Manual QA A named scenario set Varied languages, states and longer fields A demo that covers only the happy path
Performance tests Controlled dataset and workload Realistic cardinality, skew, sizes and concurrency Uniform rows that produce misleading results

These are starting points, not rules that ban synthetic data from unit tests or fixtures from integration tests. A property-based unit test may generate many values. A difficult integration regression may need exactly two explicit rows. Choose based on the property you want to observe, then make the failure reproducible.

For performance testing, one million distinct names are not necessarily more representative than a small set. Query plans and cache behavior can depend on repeated values, row sizes, selectivity and how many orders belong to each customer. Choose a workload from documented requirements or appropriately reviewed observations. State your assumptions and measure the system; do not treat a generator’s default distribution as production evidence.

Synthetic data vs anonymized production data

Rule-based synthetic records can be created without copying customer records at all. That can reduce the amount of sensitive operational material spread across development machines, screenshots and CI artifacts. It also makes deliberate edge cases possible when production happens not to contain them.

Anonymization starts with existing information and aims to remove the ability to identify people under the relevant conditions. Replacing names or hashing IDs does not automatically achieve that outcome: unusual combinations, retained free text or links to other information may still identify someone. Pseudonymized data should not simply be treated as anonymous.

Model-based synthetic data can also retain information about its source records. The ICO’s guidance on effective anonymisation discusses synthetic generation and identifiability assessment. A “synthetic” label is not evidence that privacy risk has disappeared. Assess the generation method, source access, output and intended recipients; do not turn this comparison into a blanket compliance claim.

Seeds help replay generation; they do not preserve everything

A seed initializes a deterministic random sequence. Repeating it can reproduce generated values only while the rest of the generation process remains compatible. Library versions, locale data, configuration, call order and time-dependent inputs can all affect the result. Extra random calls introduced by a new field may change later values.

Faker’s reproducibility documentation specifically notes version changes and relative dates. Fix the reference date for time-based generation, preserve the configuration and pin dependencies. For a regression that must survive generator upgrades, keep the reviewed output itself rather than relying only on regeneration.

fixture: checkout-customer-reference
seed: 42
generator: your factory and pinned library version
schema: checkout-fixture-v3
locale: explicitly recorded
referenceTime: 2026-01-15T12:00:00.000Z
expected: orphan customer reference is rejected
artifact: reviewed fixture JSON saved with the test

This manifest is a suggested record of your own test setup, not a Genory request format. A fixed seed also does not create new coverage on every run. Keep a stable regression suite, then explore additional seeds separately and save failures as small named cases. The API fixture guide covers the wider generation-and-replay workflow.

A practical workflow with Genory, without assuming extra capabilities

Use the Test Data Builder when you need supported field types under your own column names, or the Profile Generator when an existing profile shape fits the test. Review a small sample before expanding it. Choose targeted helpers from the tool directory instead of collecting fields your test never reads.

The builder produces flat datasets, not a relational database or automatic customer-to-order links. Prepare cross-table references in your fixture code, as the example above does. Its seed-based replay depends on matching settings and generator/data versions. Keep the exported artifact for important regressions and consult the API documentation for actual endpoints, authentication and supported options.

Best practices to apply before the next test run

  1. Name the behavior. Write what the test should prove before generating input.
  2. Choose required fidelity. Specify which formats, relationships or distributions matter. Keep irrelevant values simple.
  3. Validate the sample. Check constraints independently rather than assuming that attractive output is correct.
  4. Separate positive and negative cases. Label intentional faults and expected outcomes.
  5. Preserve replay context. Save versions, settings, time inputs and important output files.
  6. Control side effects. Keep mail, messaging, payment and external API activity in an isolated test setup.
  7. Review coverage after changes. Revisit fixtures when schemas, rules or supported locales change.

Frequently asked questions

Is synthetic data always better than fake data?

No. The terms overlap, and more elaborate generation can make a narrow test harder to understand. Choose additional structure when it exercises a requirement that simpler inputs cannot demonstrate.

Does realistic test data represent real customers?

Not necessarily. Plausible formats and readable profiles do not prove that the dataset matches customer behavior or statistical distributions. They also do not verify a person, address, phone or financial account.

Can dummy data become a test fixture?

Yes. Once you package it with a controlled setup and expected behavior, simple dummy values can be part of a fixture. The fixture’s purpose and lifecycle matter more than how its values were first created.

Can I use a fake data generator for performance tests?

It can provide inputs, but you still need to design volume, cardinality, relationships and the workload. Verify that the load generator itself is not the bottleneck, and describe the assumptions behind any reported result.

Choose the smallest dataset that can expose the failure

Start with explicit inputs when only a few values determine the assertion. Add generated structure when a workflow depends on variation or relationships. Preserve both the constraints and the failing example. The best test dataset is the one that makes the behavior observable and the failure explainable—not the one with the most realistic-looking names.