Quality is not optional. It's our standard. Free QA tools for testers and developers.

How to Make AI-Generated Test Data Reproducible

QTQA3 Team

How to Make AI‑Generated Test Data Reproducible

Reproducibility is the backbone of any reliable test suite. When test data comes from an AI model, the same prompt can yield different outputs on each run, breaking deterministic pipelines and making flaky failures hard to diagnose. Below is a practical, evidence‑driven workflow that turns stochastic generation into a repeatable artifact you can version, share, and audit.


1. Why Reproducibility Matters for AI‑Generated Data

SymptomRoot CauseImpact
Test passes locally but fails in CIDifferent random seeds or model versionsWasted debugging time, loss of confidence
Data‑driven tests produce divergent edge casesNon‑deterministic samplingInconsistent coverage, hidden bugs
Auditors ask for “exact input that triggered the bug”No snapshot of generated payloadsCompliance gaps, inability to reproduce incidents

If you can answer “exactly what data did the model emit for this run?” you close the loop between generation, execution, and analysis.


2. Prerequisites

ItemReasonMinimum Viable Setup
Fixed model endpointGuarantees same weights & tokenizerSelf‑hosted model (e.g., Llama‑2‑7B) or a pinned API version (model=gpt‑4‑0613)
Explicit seed / temperature = 0Removes stochastic samplingtemperature=0, top_p=1, seed=42 (if supported)
Prompt template stored in source controlPrompt drift = data driftprompts/generate_user_profile.tmpl
Schema / contract for outputEnables validation & diffingJSON Schema, Protobuf, or OpenAPI
Artifact storeImmutable snapshot of each generation runS3 bucket with versioned keys, Git LFS, or an internal artifact registry
CI/CD integrationAutomates generation, validation, and publishingGitHub Actions, GitLab CI, Jenkins, etc.

Tip: If you cannot control the model (e.g., a third‑party SaaS), treat the generated payload as an external dependency and lock it down exactly like you would a binary artifact.


3. Decision Matrix: Generation Strategies

StrategyDeterminismMaintenance CostWhen to Use
Static prompt + temperature = 0High (provided model version fixed)LowSimple schemas, low‑volume data
Prompt + few‑shot examples + seedMedium (seed may not be honored by all APIs)MediumNeed diversity but still want repeatable sets
Model fine‑tuned on a seed datasetHigh (model weights frozen)HighDomain‑specific language, strict compliance
Hybrid: AI generates template, deterministic filler fills valuesVery highLow‑MediumLarge volume, strict schema, need for data‑privacy guarantees

Choose the simplest strategy that satisfies your determinism requirement. Over‑engineering (e.g., fine‑tuning) adds latency and model‑drift risk without proportional benefit for most QA pipelines.


4. End‑to‑End Workflow

Below is a concrete, copy‑paste‑ready pipeline you can drop into a repo. It uses a self‑hosted Llama‑2‑7B container, but the same steps apply to any pinned API.

4.1. Repository Layout

repo/
├─ prompts/
│   └─ generate_user_profile.tmpl
├─ schemas/
│   └─ user_profile.json
├─ scripts/
│   ├─ generate.py
│   ├─ validate.py
│   └─ publish.py
├─ .github/
│   └─ workflows/
│       └─ generate-test-data.yml
└─ data/
    └─ snapshots/          # committed artifacts (Git LFS)

4.2. Prompt Template (prompts/generate_user_profile.tmpl)

{% set schema = load_json("../schemas/user_profile.json") %}
You are a test‑data generator. Output **only** a JSON object that conforms to the following JSON Schema:


{{ schema | tojson(indent=2) }}


Constraints:
- All string fields must be realistic but synthetic.
- Numeric ranges must respect the `minimum` / `maximum` keywords.
- Arrays must contain 1‑5 items unless `minItems`/`maxItems` say otherwise.
- Do **not** include any commentary, markdown, or extra keys.


Generate a single instance.

Why Jinja? It lets you embed the schema at render time, guaranteeing the prompt always matches the contract.

4.3. Generation Script (scripts/generate.py)

#!/usr/bin/env python3
"""
Deterministic AI test‑data generation.
Requires: llama-cpp-python (or any client that honors seed & temperature).
"""


import json
import os
import subprocess
import sys
from pathlib import Path
from jinja2 import Environment, FileSystemLoader


ROOT = Path(__file__).resolve().parents[1]
PROMPT_DIR = ROOT / "prompts"
SCHEMA_PATH = ROOT / "schemas" / "user_profile.json"
OUT_DIR = ROOT / "data" / "snapshots"
OUT_DIR.mkdir(parents=True, exist_ok=True)


# ----------------------------------------------------------------------


# 1. Render prompt with current schema


# ----------------------------------------------------------------------


env = Environment(loader=FileSystemLoader(PROMPT_DIR))
template = env.get_template("generate_user_profile.tmpl")
schema = json.loads(SCHEMA_PATH.read_text())
prompt = template.render(schema=schema)


# ----------------------------------------------------------------------


# 2. Call model with deterministic settings


# ----------------------------------------------------------------------


# Example using llama.cpp CLI (adjust for your runtime)


cmd = [
    "llama-cli",
    "-m", os.getenv("LLAMA_MODEL_PATH", "/models/llama-2-7b.Q4_K_M.gguf"),
    "-p", prompt,
    "--temp", "0",
    "--top-p", "1",
    "--seed", "42",
    "--n-predict", "512",
    "--json-schema", str(SCHEMA_PATH),   # optional: enforce schema at inference time
]


result = subprocess.run(cmd, capture_output=True, text=True, check=True)
raw_output = result.stdout.strip()


# ----------------------------------------------------------------------


# 3. Parse & validate


# ----------------------------------------------------------------------


try:
    data = json.loads(raw_output)
except json.JSONDecodeError as exc:
    sys.exit(f"Model returned invalid JSON: {exc}")


# Validate against schema (using jsonschema)


from jsonschema import validate, ValidationError
try:
    validate(instance=data, schema=schema)
except ValidationError as exc:
    sys.exit(f"Schema validation failed: {exc.message}")


# ----------------------------------------------------------------------


# 4. Write immutable snapshot (filename = hash of prompt + model version)


# ----------------------------------------------------------------------


import hashlib
model_version = os.getenv("LLAMA_MODEL_VERSION", "unknown")
hash_input = f"{prompt}|{model_version}|{data}".encode()
snap_name = hashlib.sha256(hash_input).hexdigest()[:12] + ".json"
snap_path = OUT_DIR / snap_name
snap_path.write_text(json.dumps(data, indent=2))
print(f"Snapshot written to {snap_path.relative_to(ROOT)}")

Key deterministic knobs

ParameterValueEffect
temperature0Greedy decoding → single most‑likely token each step
top_p1Disables nucleus sampling
seed42Fixed RNG state (if the backend respects it)
model_versionpinned in CI envGuarantees same weights

4.4. Validation Script (scripts/validate.py)

#!/usr/bin/env python3
"""Standalone validator for CI gate."""
import json, sys
from pathlib import Path
from jsonschema import validate, ValidationError


SCHEMA = Path(__file__).resolve().parents[1] / "schemas" / "user_profile.json"
schema = json.loads(SCHEMA.read_text())


for snap in Path(sys.argv[1]).rglob("*.json"):
    try:
        validate(json.loads(snap.read_text()), schema)
    except ValidationError as e:
        sys.exit(f"{snap}: {e.message}")
print("All snapshots valid")

4.5. Publish Script (scripts/publish.py)

#!/usr/bin/env python3
"""Copy validated snapshots to an artifact store (S3, GCS, etc.)."""
import os, shutil, subprocess, sys
from pathlib import Path


SRC = Path(__file__).resolve().parents[1] / "data" / "snapshots"
DEST = os.getenv("ARTIFACT_DEST")   # e.g. s3://my-bucket/test-data/user_profile/


if not DEST:
    sys.exit("ARTIFACT_DEST not set")


# Using AWS CLI as example; replace with gsutil, azcopy, etc.


cmd = ["aws", "s3", "sync", str(SRC), DEST, "--delete"]
subprocess.run(cmd, check=True)
print(f"Published {len(list(SRC.glob('*.json')))} snapshots to {DEST}")

4.6. CI Pipeline (.github/workflows/generate-test-data.yml)

name: Generate & Publish Test Data


on:
  workflow_dispatch:
  schedule:
    - cron: "0 2 * * MON"   # weekly refresh
  push:
    paths:
      - "prompts/**"
      - "schemas/**"
      - "scripts/**"


env:
  LLAMA_MODEL_PATH: /models/llama-2-7b.Q4_K_M.gguf
  LLAMA_MODEL_VERSION: "llama-2-7b-q4km-2024-03-15"
  ARTIFACT_DEST: ${{ secrets.TEST_DATA_BUCKET }}


jobs:
  generate:
    runs-on: ubuntu-latest
    container:
      image: ghcr.io/yourorg/llama-cpp:latest   # includes llama-cli, python, jsonschema
    steps:
      - uses: actions/checkout@v4
        with:
          lfs: true


- name: Generate snapshots
        run: python scripts/generate.py


- name: Validate snapshots
        run: python scripts/validate.py data/snapshots


- name: Commit snapshots (Git LFS)
        if: github.event_name == 'push'
        run: |
          git config user.name "ci-bot"
          git config user.email "ci-bot@users.noreply.github.com"
          git add data/snapshots/*.json
          git diff --staged --quiet || git commit -m "chore: update test‑data snapshots"
          git push


- name: Publish to artifact store
        run: python scripts/publish.py

What this pipeline guarantees

  1. Same prompt + same model version + same seed → identical JSON every run.
  2. Schema validation fails the build before any bad data lands in the artifact store.
  3. Git LFS snapshots give you a full history you can git diff to see exactly what changed.
  4. Artifact store provides a stable, versioned location for downstream test jobs (e.g., aws s3 cp s3://bucket/.../abc123.json ./test-input.json).

5. Worked Example: Generating a User‑Profile Dataset

Assume the schema (schemas/user_profile.json) defines:

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "title": "UserProfile",
  "type": "object",
  "required": ["id", "email", "age", "preferences"],
  "properties": {
    "id": {"type": "string", "format": "uuid"},
    "email": {"type": "string", "format": "email"},
    "age": {"type": "integer", "minimum": 13, "maximum": 120},
    "preferences": {
      "type": "array",
      "minItems": 1,
      "maxItems": 5,
      "items": {"type": "string", "enum": ["newsletter", "sms", "push", "dark_mode"]}
    }
  }
}

Running the pipeline once produces a snapshot like data/snapshots/3f2a1c9e4b7d.json:

{
  "id": "a1b2c3d4-5678-90ab-cdef-1234567890ab",
  "email": "jane.doe.8421@example.com",
  "age": 34,
  "preferences": ["newsletter", "dark_mode"]
}

Re‑run (same commit, same CI environment) → exact same file (byte‑for‑byte).
Change the schema (e.g., add phone field) → new hash → new snapshot 7e9f...json. The diff shows precisely what the model added.


6. Common Failure Modes & Mitigations

Failure ModeSymptomDetectionMitigation
Model version drift (provider upgrades)Snapshots change without code changeCI hash mismatch, diff shows new valuesPin model version in CI env; store model artifact (GGUF, ONNX) in your own registry
Seed not honored (API ignores seed)Non‑deterministic output despite temperature=0Run generation twice locally; compareSwitch to a backend that respects seed (llama.cpp, vLLM) or fall back to hybrid template + deterministic filler
Prompt leakage (model adds commentary)JSON parse error or extra keysvalidate.py failsAdd “Output only JSON” instruction; post‑process with a strict JSON extractor (regex ^\s*\{.*\}\s*$)
Schema evolution without migrationOld snapshots invalid for new testsValidation step fails on historic artifactsVersion schemas (user_profile.v1.json, v2.json); keep a migration script that upgrades old snapshots
Large‑scale generation timeoutCI job exceeds 6 h limitJob logs show SIGTERMParallelize per‑record generation; use batch inference endpoint; store intermediate checkpoints
Data‑privacy leakage (real PII appears)Audit flags real emails/phonesAutomated PII scanner (e.g., presidio) in CIEnforce synthetic‑only constraints in prompt; run a PII detector as a gate before publish

7. Checklist for a Reproducible AI‑Data Pipeline

  • Model pinned – exact weight file or API version recorded in ENV/Dockerfile.
  • Deterministic decoding – temperature=0, top_p=1, seed fixed.
  • Prompt under version control – templated, schema‑injected, no manual edits at runtime.
  • Schema as contract – JSON Schema (or Protobuf) stored beside prompt.
  • Generation script – pure function: prompt + model → JSON; no hidden state.
  • Validation gate – runs in CI, fails fast on schema violation.
  • Immutable snapshot – content‑addressed filename (hash of prompt+model+output).
  • Artifact store – versioned bucket / registry, accessible to downstream test jobs.
  • Change detection – git diff on snapshots or hash comparison in CI to trigger downstream rebuilds.
  • Rollback procedure – documented steps to revert to a previous snapshot hash.

8. Scaling Considerations

ScaleChallengePractical Adjustment
< 1 k records / runNegligibleSingle‑process script fine
10 k – 100 kCPU/GPU time, memoryBatch inference (vLLM, TGI); write snapshots in parallel shards
> 1 MStorage, CI minutesGenerate once, store in columnar format (Parquet) + manifest; use incremental generation (only new IDs)
Multi‑teamOwnership, schema conflictsCentral “test‑data platform” repo with PR‑based schema review; each team consumes via versioned artifact path

9. When Not to Use AI Generation

SituationReasonAlternative
Strict regulatory data (e.g., HIPAA‑covered PHI)Even synthetic data may be deemed “derived”Curated hand‑crafted fixtures, anonymized production snapshots
Performance‑critical load testsAI latency dominates test runtime

Read more

Cost Model for AI Test Data Generation at Scale

A buyer-focused guide to “Cost Model for AI Test Data Generation at Scale,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.

Local LLM vs Hosted AI for Test Data Generation

A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.

AI Test Data Hallucinations: Detection and Guardrails

A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.