Quality is not optional. It's our standard. Free QA tools for testers and developers.

AI Test Data Generation from Production Patterns

AI Test Data Generation from Production Patterns

Turning real‑world traffic into reliable, privacy‑safe test data – without the manual grind.


Why Production‑Derived Data Matters

Pain pointTypical workaroundWhy it falls short
Schema drift – new columns, enum values, or nested structures appear in prodHand‑crafted CSV/JSON fixturesFixtures become stale the moment a migration lands
Edge‑case coverage – rare error codes, locale‑specific formats, high‑cardinality IDsRandom generators with static rangesRandomness rarely reproduces the exact distribution that triggers bugs
Data‑privacy compliance – GDPR, CCPA, HIPAAFull production dumps in lower environmentsLegal risk, audit findings, and costly anonymisation pipelines
Test flakiness – tests depend on “magic” IDs that disappear after a deployHard‑coded IDs in test scriptsEvery deploy breaks a handful of tests, eroding confidence

Generating synthetic data from production patterns solves all four at once: the schema stays current, the statistical shape mirrors reality, privacy is baked in, and tests get stable, representative inputs.


High‑Level Workflow

flowchart TD
    A[Capture production traffic] --> B[Profile & model distributions]
    B --> C[Apply privacy transforms]
    C --> D[Generate synthetic datasets]
    D --> E[Validate against contracts]
    E --> F[Publish to test environments]
    F --> G[Continuous feedback loop]
  1. Capture – Export a representative sample (e.g., 1 % of requests over a 24 h window) from API gateways, message brokers, or DB change‑data‑capture streams.
  2. Profile – Infer column types, cardinalities, correlation matrices, and temporal patterns.
  3. Model – Fit a generative model (tabular GAN, CTGAN, TVAE, or a lightweight conditional VAE).
  4. Transform – Apply differential privacy, k‑anonymity, or tokenisation to any PII fields.
  5. Generate – Produce N rows per test‑run, optionally conditioned on test‑case parameters (e.g., “user = premium”, “region = EU”).
  6. Validate – Run schema contracts, statistical distance checks (KS‑test, Jensen‑Shannon), and business‑rule assertions.
  7. Publish – Load into test DBs, seed message queues, or feed into contract‑testing harnesses.
  8. Feedback – Compare test‑run outcomes (flakiness, coverage) with production metrics; retrain monthly.

Decision Criteria: Build vs. Buy vs. Open‑Source

CriterionBuild‑in‑houseCommercial SaaSOpen‑Source (e.g., SDV, Faker, Great Expectations)
Time to first usable dataset4‑8 weeks (data‑engineers + ML)1‑2 weeks (onboarding)1‑3 weeks (setup + tuning)
Model quality for high‑cardinality IDsCustom architectures possiblePre‑trained on similar domainsRequires manual feature engineering
Privacy guaranteesFull control, but you must implementBuilt‑in DP/k‑anonymity modulesYou add the privacy layer yourself
Integration with CI/CDTailored pipelinesWebhooks, CLI, GitHub ActionsCLI + Python API – easy to script
Cost (annual)Salary + infra$15k‑$120k depending on volumeFree (community) / support contracts
Team skill‑set fitStrong ML + data‑engLow – mostly configModerate – Python, pandas, pytest

Rule of thumb – If you have a dedicated data‑engineering squad and unique domain constraints (e.g., medical claim hierarchies), invest in a custom model. Otherwise, start with an open‑source stack and graduate to a managed service only when volume or compliance demands it.


Worked Example: E‑Commerce Order Service

1. Capture a Representative Sample



# Export 24 h of Kafka “order.created” events (≈ 250 k msgs)


kafka-console-consumer \
  --topic order.created \
  --bootstrap-server prod-broker:9092 \
  --from-beginning \
  --max-messages 250000 \
  --property print.key=true \
  --property print.value=true \
  > /tmp/order_created_24h.jsonl

Result: 250 k JSON lines, each ≈ 1.2 KB → ~300 MB raw.

2. Profile with pandas‑profiling (or ydata‑profiling)

import pandas as pd
from ydata_profiling import ProfileReport


df = pd.read_json("/tmp/order_created_24h.jsonl", lines=True)
profile = ProfileReport(df, title="Order Created – 24h Profile", explorative=True)
profile.to_file("/tmp/order_profile.html")

Key findings (excerpt):

ColumnTypeDistinctMissingTop‑5 values (freq)
order_idstring (UUID)250 k0 %–
customer_idstring (UUID)12 k0 %–
statuscategory50 %NEW 45 %, PAID 30 %, SHIPPED 15 %
total_amountfloat250 k0 %–
currencycategory30 %USD 70 %, EUR 20 %, GBP 10 %
itemsarray[object]–0 %–
created_atdatetime250 k0 %–

items is a nested array – each element carries sku, qty, unit_price. Flatten for modelling or treat as a separate relational table.

3. Choose a Generative Model

ModelStrengthWeaknessWhen to pick
CTGAN (Conditional Tabular GAN)Handles mixed types, learns correlationsLonger training (GPU‑recommended)Tabular data with strong cross‑column dependencies
TVAE (Variational Auto‑Encoder)Faster, stable on CPUSlightly lower fidelity on high‑cardinality categoricalQuick iteration, limited GPU
GReaT (LLM‑based)Captures free‑text, nested JSONHeavy compute, newerWhen payloads contain long descriptions or logs

Decision: Start with TVAE on CPU (≈ 10 min for 250 k rows). If KS‑test on total_amount > 0.05, upgrade to CTGAN.

from sdv.tabular import TVAE


model = TVAE(
    epochs=300,
    batch_size=5000,
    cuda=False,               # CPU run
    enforce_min_max_values=True,
    enforce_rounding=True,
)
model.fit(df.drop(columns=["items"]))   # flatten later
synthetic = model.sample(num_rows=100_000)
synthetic.to_parquet("/tmp/synthetic_orders.parquet")

4. Privacy Transform

from sdv.metadata import SingleTableMetadata
from sdv.single_table import GaussianCopulaSynthesizer   # for DP demo


metadata = SingleTableMetadata()
metadata.detect_from_dataframe(synthetic)


# Apply differential privacy (ε=1.0) on PII columns


synth_dp = GaussianCopulaSynthesizer(metadata, epsilon=1.0)
synth_dp.fit(synthetic)
private = synth_dp.sample(num_rows=100_000)
private.to_parquet("/tmp/synthetic_orders_dp.parquet")

Result: customer_id and order_id become synthetic UUIDs; total_amount distribution stays within 1 % KS distance.

5. Conditional Generation for Targeted Tests



# Generate 5 k premium‑user orders in EUR, status = PAID


cond = private.sample_remaining_columns(
    known_columns={
        "customer_tier": "PREMIUM",
        "currency": "EUR",
        "status": "PAID",
    },
    num_rows=5_000,
)
cond.to_parquet("/tmp/premium_eur_paid.parquet")

6. Validation Checklist

CheckToolPass criteria
Schema conformancegreat_expectations / jsonschema0 violations
Statistical fidelityscipy.stats.ks_2samp on numeric cols; chisquare on categoricalKS p‑value > 0.05, χ² p‑value > 0.05
Referential integrityCustom SQL / pandas.mergeEvery customer_id exists in synthetic customers table
Business rulesgreat_expectations expectationstotal_amount == sum(items.qty * items.unit_price)
Privacyopendp auditε ≤ 1.0, no exact matches to prod IDs
PerformanceLoad test (e.g., locust)≤ 5 % latency increase vs. prod‑size fixture

Automate the checklist in CI:



# .github/workflows/validate-synthetic.yml


name: Validate Synthetic Data
on:
  schedule: [cron: "0 3 * * MON"]   # weekly retrain
jobs:
  validate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Set up Python
        uses: actions/setup-python@v5
        with: { python-version: "3.11" }
      - name: Install deps
        run: pip install -r requirements.txt
      - name: Run validation suite
        run: pytest tests/validate_synthetic.py -q

7. Publish to Test Environments

TargetMethodExample
PostgreSQL (integration DB)COPY FROM via psqlpsql -c "\copy orders from '/tmp/synthetic_orders_dp.parquet' (FORMAT parquet);"
Kafka (contract tests)kafka-producer-perf-testkafka-producer-perf-test --topic order.created --num-records 100000 --record-size 1200 --throughput -1 --producer-props bootstrap.servers=test-broker:9092
S3 (data‑lake snapshots)aws s3 cpaws s3 cp /tmp/synthetic_orders_dp.parquet s3://qa-test-data/orders/2024-07-15/
QA3 free test‑data generatorUI / APIcurl -X POST https://qa3.io/tools/test-data-generator/generate -d @spec.json

Pitfalls & Mitigations

PitfallSymptomMitigation
Mode collapse – generator only emits a few high‑frequency valuesKS‑test passes but distinct count drops 80 %Increase model capacity, add min‑max constraints, or switch to CTGAN
Temporal leakage – synthetic created_at clusters around training windowTime‑series tests fail (e.g., “orders per hour”)Model created_at as a separate cyclic feature (hour‑of‑day, day‑of‑week) and sample from a uniform calendar
Nested structure loss – items array flattened incorrectlyReferential integrity errors on order_items tableTreat items as a child table; generate parent‑child pairs with a foreign‑key aware synthesizer (SDV RelationalSynthesizer)
Privacy over‑masking – differential privacy destroys rare‑event signalsFraud‑detection tests never see status=FRAUDUse conditional DP: lower ε for high‑risk columns, higher ε for bulk columns
Schema drift unnoticed – new column appears in prod, synthetic still oldCI passes but production bugs slip throughAdd a schema‑watch job that diffs information_schema nightly and triggers retrain
Resource contention – training on shared GPU starves other workloadsNightly jobs miss SLASchedule training on spot instances or dedicated CPU‑only TVAE runs

Tooling Landscape (2024‑2025 Snapshot)

CategoryNotable OptionsLicensingTypical Use‑Case
Tabular synthesisSDV (TVAE, CTGAN, CopulaGAN), YData Synthetic, MOSTLY AIApache‑2.0 / CommercialCore data generation
Relational / multi‑tableSDV RelationalSynthesizer, Synthesized.io, Tonic.aiApache‑2.0 / SaaSOrder‑line, user‑profile graphs
Privacy‑enhancingOpenDP, Google Differential Privacy, ARXApache‑2.0 / MITε‑DP, k‑anonymity, synthetic‑data‑release
ValidationGreat Expectations, Deequ, PanderaApache‑2.0Contract testing, CI gates
OrchestrationAirflow, Prefect, Dagster, GitHub ActionsApache‑2.0 / MITEnd‑to‑end pipeline
Free quick‑startQA3 Test Data Generator – /tools/test-data-generatorFree (no account)One‑off CSV/JSON/Parquet for prototyping

Tip: Keep the generation pipeline declarative (YAML/JSON spec) so you can swap the underlying engine without rewriting tests.


Scaling the Practice Across Teams

  1. Centralised “Data‑Factory” repo – single source of truth for specs, models, and validation suites.
  2. Self‑service API – POST /synthetic?spec=order&rows=5000&cond=premium_eur_paid returns a signed URL to a Parquet file in object storage.
  3. Versioned datasets – Tag each generation run (v2024.07.15‑01) and store metadata (model hash, ε, row count) in a lightweight catalog (e.g., DataHub).
  4. Governance board – Quarterly review of privacy budget consumption, model drift metrics, and test‑flakiness trends.
  5. Enable developers – Provide a thin wrapper (qa3-testdata npm/py package) so unit tests can request a fresh slice on the fly.

Checklist: From Zero to Production‑Ready Synthetic Data

  • Inventory all data sources that feed test environments (DB, Kafka, S3, third‑party APIs).
  • Define a sampling policy (percentage, time window, stratification).
  • Select a baseline generator (TVAE for speed, CTGAN for fidelity).
  • Implement privacy transforms (DP, tokenisation) before any data leaves the secure zone.
  • Automate schema‑contract validation in CI (Great Expectations + JSON Schema).
  • Measure statistical distance on a hold‑out production slice each retrain.
  • Publish artifacts to a versioned, immutable store (S3 + Glue catalog).
  • Integrate with test harnesses (pytest fixtures, JUnit @ParameterizedTest, Cypress cy.task).
  • Monitor flakiness & coverage dashboards; correlate with synthetic‑data version.
  • Document the end‑to‑end flow in the team wiki; include rollback steps.

Next Action

  1. **Pick a single, high‑impact

Read more

Using AI to Expand a Small Test Dataset

A practical guide to “Using AI to Expand a Small Test Dataset,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.

AI-Generated Addresses, Names, and Profiles: Realism vs Safety

A buyer-focused guide to “AI-Generated Addresses, Names, and Profiles: Realism vs Safety,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.

Multilingual Test Data Generation with AI

A practical guide to “Multilingual Test Data Generation with AI,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.