How to Validate AI-Generated Test Data Before Use
How to Validate AI‑Generated Test Data Before Use
AI‑driven test data generators can spin up thousands of rows in seconds, but speed does not guarantee correctness. A single malformed record can corrupt a test run, hide a bug, or—worse—leak production‑like patterns into a non‑production environment. This guide walks through a practical validation workflow that you can embed in any CI/CD pipeline, whether you’re using a commercial platform or the free QA3 test data generator at /tools/test-data-generator.
1. Why Validation Is a Separate Concern
| Symptom | Root cause | What validation catches |
|---|---|---|
| Flaky UI tests that fail on “invalid email” | Generator produced user@ without a domain | Format / regex checks |
| Load test reports “connection refused” on a service that works locally | Synthetic data referenced a non‑existent foreign key | Referential integrity |
| Security audit flags PII in staging | Generator copied real‑world name/address patterns | Privacy / masking rules |
| Performance baseline drifts after each run | Data volume or distribution changes silently | Statistical distribution thresholds |
Treating validation as a first‑class step—rather than an after‑thought—turns these symptoms into deterministic gate checks.
2. Prerequisites Before You Generate
| Item | Description | How to verify |
|---|---|---|
| Schema contract | JSON Schema, Avro, Protobuf, or DB DDL that defines every column, type, nullability, and constraints. | Store in version control; run jsonschema -i or equivalent on a sample. |
| Business rules catalog | Cross‑field rules (e.g., order_date ≤ ship_date), enum lists, checksum algorithms. | Document as executable specs (Cucumber, SpecFlow, or plain Python functions). |
| Reference data sets | Lookup tables (country codes, currency codes, product catalog IDs) that generated rows must join against. | Load into a test‑only DB or in‑memory map for fast look‑ups. |
| Privacy policy | List of fields that must never contain real PII, plus masking requirements (e.g., last‑4 of SSN). | Encode as tag‑based rules (@pii, @mask:last4). |
| Statistical baselines | Expected distributions (e.g., 70 % “standard” users, 20 % “premium”, 10 % “admin”). | Capture from production snapshots (anonymized) or from domain experts. |
If any of these artifacts are missing, generate a minimum viable contract first—validation cannot be stricter than the specification you own.
3. Validation Layers
Think of validation as a pipeline of increasingly expensive checks. Fast, cheap checks run on every record; heavy statistical checks run on a sample or on the full set once per build.
3.1 Structural Validation
| Check | Tool / Technique | Failure handling |
|---|---|---|
| Schema conformity | jsonschema, avro-tools, protobuf validation | Reject record, log path |
| Required fields present | Simple presence test | Reject |
| Type coercion safety | Try‑cast to target language type (e.g., int(), datetime.fromisoformat()) | Reject |
| Enum / allow‑list membership | Set lookup (value in ALLOWED_VALUES) | Reject |
Implementation tip: Wrap these in a pure function validate_structure(record) -> ValidationResult so you can unit‑test the validator independently of the generator.
3.2 Business‑Rule Validation
def validate_business_rules(rec: dict) -> list[str]:
errors = []
if rec["order_date"] > rec["ship_date"]:
errors.append("order_date after ship_date")
if rec["account_type"] == "premium" and rec["credit_limit"] < 5000:
errors.append("premium account requires credit_limit ≥ 5000")
return errors
- Keep rules declarative (data‑driven) when possible: a CSV of
rule_id, expression, error_msglets non‑coders add constraints. - Use a sandboxed expression evaluator (e.g.,
asteval,jsonlogic) to avoid arbitrary code execution.
3.3 Referential Integrity
| Approach | When to use |
|---|---|
| In‑memory hash set of primary keys | Small reference tables (< 100 k rows) |
| Temp DB (SQLite, DuckDB) with foreign‑key constraints | Larger sets, need SQL semantics |
| Streaming join (Kafka Streams, Flink) | Continuous generation pipelines |
Validate both directions: every generated FK must exist in the reference set, and every reference key that should be covered by tests must appear at least once (coverage check).
3.4 Distribution & Statistical Validation
| Metric | Acceptance criterion | Tool |
|---|---|---|
Categorical frequencies (e.g., account_type) | χ² p‑value > 0.05 vs. baseline | scipy.stats.chisquare |
| Numeric moments (mean, std, quantiles) | Within ±5 % of baseline | numpy.percentile |
Correlation matrix (e.g., age vs. income) | Max absolute diff < 0.1 | pandas.DataFrame.corr |
| Null‑rate per column | ≤ defined threshold | Simple aggregation |
Run these on a representative sample (e.g., 10 % of rows, stratified by primary key) to keep CI time reasonable. Store baseline snapshots as JSON artifacts in the repo for reproducibility.
3.5 Privacy & Masking Validation
| Rule | Check |
|---|---|
No real email domains (e.g., gmail.com) | Regex `^[^@]+@example.(com |
SSN masked to ***-**-1234 | value.endswith(last4) and value.startswith("***-**-") |
| No geographic coordinates within 1 km of real addresses | Haversine distance against a sanitized address list |
Automate with a policy engine (OPA, custom rule runner) so new privacy rules are just data, not code changes.
4. Embedding Validation in CI/CD
# .github/workflows/test-data.yml
jobs:
generate-and-validate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Generate data
run: |
python -m qa3.datagen --spec spec.yaml --out data/out.parquet
- name: Structural validation
run: python -m validate.structural data/out.parquet
- name: Business rule validation
run: python -m validate.business data/out.parquet
- name: Referential integrity
run: python -m validate.referential data/out.parquet ref/
- name: Distribution checks (sample)
run: python -m validate.distribution data/out.parquet baselines/
- name: Privacy scan
run: python -m validate.privacy data/out.parquet
- name: Publish artifact
uses: actions/upload-artifact@v4
with:
name: validated-test-data
path: data/out.parquet
Key points
- Fail fast – structural and business‑rule steps run on the full set; they are cheap and catch the majority of defects.
- Gate on distribution – treat a statistical failure as a warning first; promote to error after you have a stable baseline.
- Artifact promotion – only the validated Parquet/CSV moves to downstream test jobs (integration, performance, contract).
5. Worked Example: E‑Commerce Order Dataset
5.1 Specification (excerpt)
# spec.yaml
tables:
- name: customers
rows: 5000
columns:
- {name: id, type: uuid, pk: true}
- {name: email, type: string, format: email, pii: true}
- {name: tier, type: enum, values: [standard, premium, vip]}
- {name: created_at, type: datetime, range: ["2020-01-01", "2024-12-31"]}
- name: orders
rows: 20000
columns:
- {name: id, type: uuid, pk: true}
- {name: customer_id, type: uuid, fk: customers.id}
- {name: total_cents, type: int, min: 0}
- {name: status, type: enum, values: [placed, paid, shipped, cancelled]}
- {name: placed_at, type: datetime}
- {name: shipped_at, type: datetime, nullable: true}
5.2 Generation (using QA3 free generator)
qa3-datagen --spec spec.yaml --out /tmp/ecom.parquet
5.3 Validation Scripts (Python snippets)
Structural
import pyarrow.parquet as pq
from jsonschema import validate, ValidationError
schema = {
"type": "object",
"properties": {
"id": {"type": "string", "format": "uuid"},
"email": {"type": "string", "format": "email"},
"tier": {"enum": ["standard", "premium", "vip"]},
"created_at": {"type": "string", "format": "date-time"},
},
"required": ["id", "email", "tier", "created_at"],
"additionalProperties": False,
}
def validate_structural(table_path):
table = pq.read_table(table_path)
for batch in table.to_batches():
for row in batch.to_pylist():
try:
validate(instance=row, schema=schema)
except ValidationError as e:
yield f"{row['id']}: {e.message}"
Business rule – order dates
def validate_order_dates(orders_path, customers_path):
orders = pq.read_table(orders_path).to_pandas()
customers = pq.read_table(customers_path).to_pandas()
merged = orders.merge(customers, left_on="customer_id", right_on="id", suffixes=("_o", "_c"))
bad = merged[
(merged["placed_at_o"] > merged["shipped_at_o"]) &
merged["shipped_at_o"].notna()
]
for _, r in bad.iterrows():
yield f"order {r['id_o']}: placed_at after shipped_at"
Referential integrity
def validate_fk(orders_path, customers_path):
order_ids = set(pq.read_table(orders_path)["customer_id"].to_pylist())
cust_ids = set(pq.read_table(customers_path)["id"].to_pylist())
missing = order_ids - cust_ids
for mid in missing:
yield f"order references missing customer {mid}"
Distribution (χ² on tier)
from scipy.stats import chisquare
import numpy as np
EXPECTED_TIER = {"standard": 0.70, "premium": 0.20, "vip": 0.10}
def validate_tier_distribution(customers_path):
df = pq.read_table(customers_path).to_pandas()
observed = df["tier"].value_counts().reindex(EXPECTED_TIER.keys(), fill_value=0)
total = observed.sum()
expected = np.array([EXPECTED_TIER[k] * total for k in EXPECTED_TIER])
stat, p = chisquare(observed, expected)
if p < 0.05:
yield f"Tier distribution drift (χ²={stat:.2f}, p={p:.4f})"
Privacy – email domain
import re
ALLOWED_DOMAIN = re.compile(r"^[^@]+@example\.(com|org|net)$")
def validate_email_privacy(customers_path):
for row in pq.read_table(customers_path).to_pylist():
if not ALLOWED_DOMAIN.match(row["email"]):
yield f"customer {row['id']}: non‑example email {row['email']}"
5.4 Running the Full Suite
python -m validate.all /tmp/ecom.parquet /tmp/ecom_customers.parquet
# Exit code 0 → artifact promoted
# Exit code 1 → CI fails, logs show exact rows
The output is a concise list of offending primary keys, making triage trivial.
6. Common Failure Modes & Mitigations
| Failure mode | Symptom | Root cause | Mitigation |
|---|---|---|---|
| Schema drift | New column appears in generator output but not in validator | Generator version upgraded without updating contract | Pin generator version in CI; run qa3-datagen --version as a step |
| Enum expansion | Validation rejects tier: "gold" | Business added a tier, spec not updated | Treat enum list as source of truth in a shared repo; generate both spec and validator from it |
| Statistical flakiness | χ² test fails on 1 % of runs | Small sample size, natural variance | Increase sample size or switch to sequential probability ratio test (SPRT) for streaming validation |
| Reference data stale | FK errors after a product catalog refresh | Reference snapshot not refreshed | Automate reference snapshot refresh as a nightly job; version the snapshot artifact |
| Privacy leakage | Real email domain appears in logs | Generator fell back to a default faker provider | Explicitly disable default providers; whitelist only safe providers in generator config |
| Performance bottleneck | Validation step > 10 min on 5 M rows | Full‑scan distribution checks on every column | Run heavy stats on a stratified sample; cache baseline histograms |
7. Checklist – Before You Ship Generated Data
- Schema contract committed and versioned.
- Business‑rule catalog expressed as executable tests.
- Reference data snapshot available and immutable for the run.
- Privacy tags applied to every PII column.
- Statistical baselines stored as JSON artifacts.
- Validation pipeline runs in CI before any consumer job.
- Failure policy defined: hard fail for structural/business/referential; warning → error for distribution after baseline stabilisation.
- Artifact promotion only for fully‑validated data.
- Observability: validation metrics (rows checked, error counts, χ² p‑values) exported to monitoring.
- Rollback plan: keep the last known‑good dataset artifact for quick revert.
8. Next Steps
- Add a validation stage to your existing generation job—start with structural checks only.
- Export your current schema (DB DDL → JSON Schema) and commit it alongside the generator config.
- Run the free QA3 test data generator at
/tools/test-data-generatoron a small spec, then feed the output through the validation snippets above. - Measure CI time; if the full suite exceeds your budget, move distribution checks to a nightly job and keep fast checks on every PR.
- Iterate: each new business rule or privacy requirement becomes a new validator test, not a manual review.
By treating validation as code—versioned, testable, and gated—you turn AI‑generated data from a “black‑box convenience” into a reliable, auditable asset that your test suites can trust.
Read more
Local LLM vs Hosted AI for Test Data Generation
A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.
AI Test Data Hallucinations: Detection and Guardrails
A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.
Seeded AI Test Data Generation for Stable Automation
A practical guide to “Seeded AI Test Data Generation for Stable Automation,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.