Quality is not optional. It's our standard. Free QA tools for testers and developers.

How to Validate AI-Generated Test Data Before Use

QTQA3 Team

How to Validate AI‑Generated Test Data Before Use

AI‑driven test data generators can spin up thousands of rows in seconds, but speed does not guarantee correctness. A single malformed record can corrupt a test run, hide a bug, or—worse—leak production‑like patterns into a non‑production environment. This guide walks through a practical validation workflow that you can embed in any CI/CD pipeline, whether you’re using a commercial platform or the free QA3 test data generator at /tools/test-data-generator.


1. Why Validation Is a Separate Concern

SymptomRoot causeWhat validation catches
Flaky UI tests that fail on “invalid email”Generator produced user@ without a domainFormat / regex checks
Load test reports “connection refused” on a service that works locallySynthetic data referenced a non‑existent foreign keyReferential integrity
Security audit flags PII in stagingGenerator copied real‑world name/address patternsPrivacy / masking rules
Performance baseline drifts after each runData volume or distribution changes silentlyStatistical distribution thresholds

Treating validation as a first‑class step—rather than an after‑thought—turns these symptoms into deterministic gate checks.


2. Prerequisites Before You Generate

ItemDescriptionHow to verify
Schema contractJSON Schema, Avro, Protobuf, or DB DDL that defines every column, type, nullability, and constraints.Store in version control; run jsonschema -i or equivalent on a sample.
Business rules catalogCross‑field rules (e.g., order_date ≤ ship_date), enum lists, checksum algorithms.Document as executable specs (Cucumber, SpecFlow, or plain Python functions).
Reference data setsLookup tables (country codes, currency codes, product catalog IDs) that generated rows must join against.Load into a test‑only DB or in‑memory map for fast look‑ups.
Privacy policyList of fields that must never contain real PII, plus masking requirements (e.g., last‑4 of SSN).Encode as tag‑based rules (@pii, @mask:last4).
Statistical baselinesExpected distributions (e.g., 70 % “standard” users, 20 % “premium”, 10 % “admin”).Capture from production snapshots (anonymized) or from domain experts.

If any of these artifacts are missing, generate a minimum viable contract first—validation cannot be stricter than the specification you own.


3. Validation Layers

Think of validation as a pipeline of increasingly expensive checks. Fast, cheap checks run on every record; heavy statistical checks run on a sample or on the full set once per build.

3.1 Structural Validation

CheckTool / TechniqueFailure handling
Schema conformityjsonschema, avro-tools, protobuf validationReject record, log path
Required fields presentSimple presence testReject
Type coercion safetyTry‑cast to target language type (e.g., int(), datetime.fromisoformat())Reject
Enum / allow‑list membershipSet lookup (value in ALLOWED_VALUES)Reject

Implementation tip: Wrap these in a pure function validate_structure(record) -> ValidationResult so you can unit‑test the validator independently of the generator.

3.2 Business‑Rule Validation

def validate_business_rules(rec: dict) -> list[str]:
    errors = []
    if rec["order_date"] > rec["ship_date"]:
        errors.append("order_date after ship_date")
    if rec["account_type"] == "premium" and rec["credit_limit"] < 5000:
        errors.append("premium account requires credit_limit ≥ 5000")
    return errors
  • Keep rules declarative (data‑driven) when possible: a CSV of rule_id, expression, error_msg lets non‑coders add constraints.
  • Use a sandboxed expression evaluator (e.g., asteval, jsonlogic) to avoid arbitrary code execution.

3.3 Referential Integrity

ApproachWhen to use
In‑memory hash set of primary keysSmall reference tables (< 100 k rows)
Temp DB (SQLite, DuckDB) with foreign‑key constraintsLarger sets, need SQL semantics
Streaming join (Kafka Streams, Flink)Continuous generation pipelines

Validate both directions: every generated FK must exist in the reference set, and every reference key that should be covered by tests must appear at least once (coverage check).

3.4 Distribution & Statistical Validation

MetricAcceptance criterionTool
Categorical frequencies (e.g., account_type)χ² p‑value > 0.05 vs. baselinescipy.stats.chisquare
Numeric moments (mean, std, quantiles)Within ±5 % of baselinenumpy.percentile
Correlation matrix (e.g., age vs. income)Max absolute diff < 0.1pandas.DataFrame.corr
Null‑rate per column≤ defined thresholdSimple aggregation

Run these on a representative sample (e.g., 10 % of rows, stratified by primary key) to keep CI time reasonable. Store baseline snapshots as JSON artifacts in the repo for reproducibility.

3.5 Privacy & Masking Validation

RuleCheck
No real email domains (e.g., gmail.com)Regex `^[^@]+@example.(com
SSN masked to ***-**-1234value.endswith(last4) and value.startswith("***-**-")
No geographic coordinates within 1 km of real addressesHaversine distance against a sanitized address list

Automate with a policy engine (OPA, custom rule runner) so new privacy rules are just data, not code changes.


4. Embedding Validation in CI/CD



# .github/workflows/test-data.yml


jobs:
  generate-and-validate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Generate data
        run: |
          python -m qa3.datagen --spec spec.yaml --out data/out.parquet
      - name: Structural validation
        run: python -m validate.structural data/out.parquet
      - name: Business rule validation
        run: python -m validate.business data/out.parquet
      - name: Referential integrity
        run: python -m validate.referential data/out.parquet ref/
      - name: Distribution checks (sample)
        run: python -m validate.distribution data/out.parquet baselines/
      - name: Privacy scan
        run: python -m validate.privacy data/out.parquet
      - name: Publish artifact
        uses: actions/upload-artifact@v4
        with:
          name: validated-test-data
          path: data/out.parquet

Key points

  • Fail fast – structural and business‑rule steps run on the full set; they are cheap and catch the majority of defects.
  • Gate on distribution – treat a statistical failure as a warning first; promote to error after you have a stable baseline.
  • Artifact promotion – only the validated Parquet/CSV moves to downstream test jobs (integration, performance, contract).

5. Worked Example: E‑Commerce Order Dataset

5.1 Specification (excerpt)



# spec.yaml


tables:
  - name: customers
    rows: 5000
    columns:
      - {name: id, type: uuid, pk: true}
      - {name: email, type: string, format: email, pii: true}
      - {name: tier, type: enum, values: [standard, premium, vip]}
      - {name: created_at, type: datetime, range: ["2020-01-01", "2024-12-31"]}
  - name: orders
    rows: 20000
    columns:
      - {name: id, type: uuid, pk: true}
      - {name: customer_id, type: uuid, fk: customers.id}
      - {name: total_cents, type: int, min: 0}
      - {name: status, type: enum, values: [placed, paid, shipped, cancelled]}
      - {name: placed_at, type: datetime}
      - {name: shipped_at, type: datetime, nullable: true}

5.2 Generation (using QA3 free generator)

qa3-datagen --spec spec.yaml --out /tmp/ecom.parquet

5.3 Validation Scripts (Python snippets)

Structural

import pyarrow.parquet as pq
from jsonschema import validate, ValidationError


schema = {
    "type": "object",
    "properties": {
        "id": {"type": "string", "format": "uuid"},
        "email": {"type": "string", "format": "email"},
        "tier": {"enum": ["standard", "premium", "vip"]},
        "created_at": {"type": "string", "format": "date-time"},
    },
    "required": ["id", "email", "tier", "created_at"],
    "additionalProperties": False,
}


def validate_structural(table_path):
    table = pq.read_table(table_path)
    for batch in table.to_batches():
        for row in batch.to_pylist():
            try:
                validate(instance=row, schema=schema)
            except ValidationError as e:
                yield f"{row['id']}: {e.message}"

Business rule – order dates

def validate_order_dates(orders_path, customers_path):
    orders = pq.read_table(orders_path).to_pandas()
    customers = pq.read_table(customers_path).to_pandas()
    merged = orders.merge(customers, left_on="customer_id", right_on="id", suffixes=("_o", "_c"))
    bad = merged[
        (merged["placed_at_o"] > merged["shipped_at_o"]) &
        merged["shipped_at_o"].notna()
    ]
    for _, r in bad.iterrows():
        yield f"order {r['id_o']}: placed_at after shipped_at"

Referential integrity

def validate_fk(orders_path, customers_path):
    order_ids = set(pq.read_table(orders_path)["customer_id"].to_pylist())
    cust_ids = set(pq.read_table(customers_path)["id"].to_pylist())
    missing = order_ids - cust_ids
    for mid in missing:
        yield f"order references missing customer {mid}"

Distribution (χ² on tier)

from scipy.stats import chisquare
import numpy as np


EXPECTED_TIER = {"standard": 0.70, "premium": 0.20, "vip": 0.10}


def validate_tier_distribution(customers_path):
    df = pq.read_table(customers_path).to_pandas()
    observed = df["tier"].value_counts().reindex(EXPECTED_TIER.keys(), fill_value=0)
    total = observed.sum()
    expected = np.array([EXPECTED_TIER[k] * total for k in EXPECTED_TIER])
    stat, p = chisquare(observed, expected)
    if p < 0.05:
        yield f"Tier distribution drift (χ²={stat:.2f}, p={p:.4f})"

Privacy – email domain

import re


ALLOWED_DOMAIN = re.compile(r"^[^@]+@example\.(com|org|net)$")


def validate_email_privacy(customers_path):
    for row in pq.read_table(customers_path).to_pylist():
        if not ALLOWED_DOMAIN.match(row["email"]):
            yield f"customer {row['id']}: non‑example email {row['email']}"

5.4 Running the Full Suite

python -m validate.all /tmp/ecom.parquet /tmp/ecom_customers.parquet


# Exit code 0 → artifact promoted


# Exit code 1 → CI fails, logs show exact rows


The output is a concise list of offending primary keys, making triage trivial.


6. Common Failure Modes & Mitigations

Failure modeSymptomRoot causeMitigation
Schema driftNew column appears in generator output but not in validatorGenerator version upgraded without updating contractPin generator version in CI; run qa3-datagen --version as a step
Enum expansionValidation rejects tier: "gold"Business added a tier, spec not updatedTreat enum list as source of truth in a shared repo; generate both spec and validator from it
Statistical flakinessχ² test fails on 1 % of runsSmall sample size, natural varianceIncrease sample size or switch to sequential probability ratio test (SPRT) for streaming validation
Reference data staleFK errors after a product catalog refreshReference snapshot not refreshedAutomate reference snapshot refresh as a nightly job; version the snapshot artifact
Privacy leakageReal email domain appears in logsGenerator fell back to a default faker providerExplicitly disable default providers; whitelist only safe providers in generator config
Performance bottleneckValidation step > 10 min on 5 M rowsFull‑scan distribution checks on every columnRun heavy stats on a stratified sample; cache baseline histograms

7. Checklist – Before You Ship Generated Data

  • Schema contract committed and versioned.
  • Business‑rule catalog expressed as executable tests.
  • Reference data snapshot available and immutable for the run.
  • Privacy tags applied to every PII column.
  • Statistical baselines stored as JSON artifacts.
  • Validation pipeline runs in CI before any consumer job.
  • Failure policy defined: hard fail for structural/business/referential; warning → error for distribution after baseline stabilisation.
  • Artifact promotion only for fully‑validated data.
  • Observability: validation metrics (rows checked, error counts, χ² p‑values) exported to monitoring.
  • Rollback plan: keep the last known‑good dataset artifact for quick revert.

8. Next Steps

  1. Add a validation stage to your existing generation job—start with structural checks only.
  2. Export your current schema (DB DDL → JSON Schema) and commit it alongside the generator config.
  3. Run the free QA3 test data generator at /tools/test-data-generator on a small spec, then feed the output through the validation snippets above.
  4. Measure CI time; if the full suite exceeds your budget, move distribution checks to a nightly job and keep fast checks on every PR.
  5. Iterate: each new business rule or privacy requirement becomes a new validator test, not a manual review.

By treating validation as code—versioned, testable, and gated—you turn AI‑generated data from a “black‑box convenience” into a reliable, auditable asset that your test suites can trust.

Read more

Local LLM vs Hosted AI for Test Data Generation

A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.

AI Test Data Hallucinations: Detection and Guardrails

A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.

Seeded AI Test Data Generation for Stable Automation

A practical guide to “Seeded AI Test Data Generation for Stable Automation,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.