Seeded AI Test Data Generation for Stable Automation
Seeded AI Test Data Generation for Stable Automation
Why “just generate data” isn’t enough, and how a seeded approach gives you repeatable, realistic test inputs without the flakiness of pure randomness.
The Problem: Random Data ≠ Reliable Tests
Most teams start with a simple script that calls faker or a cloud‑based “AI data generator” and feeds the output straight into their test suite. The first few runs look green, then:
| Symptom | Root cause |
|---|---|
| Intermittent failures on CI | Different data shapes each run (e.g., missing required fields, out‑of‑range values) |
| Tests pass locally but fail on staging | Environment‑specific constraints (unique indexes, foreign‑key limits) not respected |
| Hard to reproduce a bug | No deterministic seed → you cannot re‑run the exact data set that triggered the failure |
| Data‑privacy complaints | Real‑world PII leaks into generated payloads because the model was trained on production snapshots |
Randomness is useful for exploratory testing, but automation pipelines need stability. A seeded AI test data generator gives you the best of both worlds: realistic, domain‑aware values and a single source of truth you can version‑control.
What “Seeded” Means in This Context
| Term | Definition |
|---|---|
| Seed | A deterministic input (integer, UUID, or hash) that drives the pseudo‑random number generator (PRNG) inside the AI model or the downstream data‑shaping code. |
| Prompt template | A version‑controlled prompt (or system message) that describes the domain, constraints, and output format. |
| Deterministic pipeline | Seed → Prompt → Model → Post‑processing → Serialized test data artifact (JSON, CSV, SQL dump). |
| Artifact versioning | Store the generated artifact (or the seed + prompt) alongside the test code so any CI run can reconstruct the exact same data set. |
When you commit the seed (or the artifact) you get repeatability without sacrificing the richness that a large language model (LLM) can provide.
Decision Criteria: When to Use a Seeded AI Generator
| Situation | Seeded AI generator | Pure scripted generators (Faker, FactoryBot) | Hand‑crafted fixtures |
|---|---|---|---|
| Need realistic‑looking names, addresses, medical codes, financial transactions | ✅ | ❌ (requires large custom dictionaries) | ✅ (but maintenance heavy) |
Data must satisfy complex cross‑field rules (e.g., start_date < end_date, total = sum(line_items)) | ✅ (prompt can encode rules) | ✅ (code can enforce) | ✅ (manual) |
| Test suite runs > 10× per day on multiple agents | ✅ (artifact cached) | ✅ (fast) | ❌ (slow to load large fixtures) |
| Regulatory audit requires proof of data provenance | ✅ (seed + prompt stored) | ✅ (code is source) | ✅ (fixture files) |
| Team has no ML ops expertise | ✅ (use hosted API with seed param) | ✅ | ✅ |
| Latency budget < 50 ms per test case | ❌ (model call) | ✅ | ✅ |
Rule of thumb: If you need domain realism and determinism for CI, a seeded AI approach pays off. If you only need simple scalar values (ints, enums, UUIDs), stick with a lightweight library.
Workflow Overview
┌─────────────────────┐
│ 1️⃣ Define prompt │ (version‑controlled, markdown or JSON)
└───────┬─────────────┘
│
▼
┌─────────────────────┐
│ 2️⃣ Choose seed │ (CI build number, git SHA, or static constant)
└───────┬─────────────┘
│
▼
┌─────────────────────┐
│ 3️⃣ Call model API │ (OpenAI, Anthropic, local Llama, etc.)
│ with seed param │
└───────┬─────────────┘
│
▼
┌─────────────────────┐
│ 4️⃣ Post‑process │ (schema validation, referential integrity,
│ & serialize │ PII scrubbing, format conversion)
└───────┬─────────────┘
│
▼
┌─────────────────────┐
│ 5️⃣ Store artifact │ (git‑LFS, S3, artifact repo, or commit JSON)
└───────┬─────────────┘
│
▼
┌─────────────────────┐
│ 6️⃣ Consume in tests│ (load once per suite, or per test class)
└─────────────────────┘
Key invariants
- Prompt + seed = deterministic output (provided the model version is pinned).
- Artifact is the source of truth for the test run; the model is only a build‑time step.
- Validation gates (schema, business rules) run after generation, not inside the prompt.
Worked Example: E‑Commerce Order Flow
1. Prompt Template (committed as prompts/order_v1.md)
# Order Generation Prompt v1.0
You are a test data generator for an e‑commerce platform.
Produce a JSON array of **10** order objects.
Each order must contain:
- `order_id`: UUID v4
- `customer`: object with `id` (UUID), `email` (valid format), `full_name` (realistic), `phone` (E.164)
- `items`: array of 1‑5 line items, each with
- `sku` (alphanumeric, 8‑12 chars)
- `quantity` (1‑10)
- `unit_price_cents` (100‑50000)
- `discount_pct` (0‑30, integer)
- `shipping_address`: object with `street`, `city`, `state`, `postal_code`, `country` (ISO‑3166‑1 alpha‑2)
- `billing_address`: same shape as shipping, may be identical
- `placed_at`: ISO‑8601 timestamp within the last 180 days
- `status`: one of `["pending","paid","shipped","delivered","cancelled"]`
- `total_cents`: **must equal** sum(`quantity * unit_price_cents * (100 - discount_pct) / 100`) rounded to nearest cent
Constraints:
- No PII from real people.
- `email` domain must be `example.com`.
- `postal_code` must match the `country` format (US 5‑digit, CA `A1A 1A1`, DE 5‑digit).
- `status` distribution: 10 % pending, 30 % paid, 30 % shipped, 20 % delivered, 10 % cancelled.
Output **only** the JSON array, no markdown fences.
Why this works: The prompt encodes cross‑field arithmetic (total_cents), format constraints (postal codes), and distribution requirements. All of that is expressed in natural language, which the model follows reliably when the seed is fixed.
2. Seed Selection
# CI uses the git short SHA so every commit gets a unique but reproducible data set
SEED=$(git rev-parse --short HEAD) # e.g. a1b2c3d
For local debugging you can export a constant:
export TEST_DATA_SEED=deadbeef
3. Model Call (Python snippet)
import os, json, subprocess, hashlib, pathlib
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
MODEL = "gpt-4o-2024-08-06" # pin the version
PROMPT_PATH = pathlib.Path("prompts/order_v1.md")
SEED = os.getenv("TEST_DATA_SEED", "deadbeef")
prompt = PROMPT_PATH.read_text()
# The OpenAI API does not have a native seed param yet (as of 2024‑08),
# so we emulate determinism by hashing the seed into the system message.
system_msg = f"Deterministic generation seed: {SEED}. Do not vary output."
resp = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content": system_msg},
{"role": "user", "content": prompt},
],
temperature=0.0, # deterministic sampling
max_tokens=4000,
)
raw = resp.choices[0].message.content
orders = json.loads(raw)
Note – If your provider supports a native
seedparameter (e.g., Azure OpenAI, Anthropic), use it instead of the system‑message hack.
4. Post‑Processing & Validation
from jsonschema import validate
import datetime, re
SCHEMA = {
"type": "array",
"items": {
"type": "object",
"required": ["order_id","customer","items","shipping_address",
"billing_address","placed_at","status","total_cents"],
"properties": {
"order_id": {"type":"string","format":"uuid"},
"customer": {"type":"object"},
"items": {"type":"array","minItems":1,"maxItems":5},
"shipping_address": {"type":"object"},
"billing_address": {"type":"object"},
"placed_at": {"type":"string","format":"date-time"},
"status": {"enum":["pending","paid","shipped","delivered","cancelled"]},
"total_cents": {"type":"integer","minimum":0}
}
},
"minItems": 10,
"maxItems": 10
}
validate(instance=orders, schema=SCHEMA)
# Business‑rule check: total_cents matches line items
def compute_total(order):
s = 0
for it in order["items"]:
price = it["unit_price_cents"] * it["quantity"]
discount = price * it["discount_pct"] // 100
s += price - discount
return s
for o in orders:
assert o["total_cents"] == compute_total(o), f"Total mismatch on {o['order_id']}"
# PII scrub – ensure email domain
for o in orders:
assert o["customer"]["email"].endswith("@example.com")
If any assertion fails, the CI job stops before the artifact is stored, preventing flaky data from polluting the test run.
5. Artifact Storage
# Write to a versioned location (git‑LFS recommended for >1 MB)
OUT_DIR="testdata/orders/${SEED}"
mkdir -p "$OUT_DIR"
echo "$raw" > "$OUT_DIR/orders.json"
git add "$OUT_DIR/orders.json"
git commit -m "chore(testdata): seed ${SEED} order dataset"
git push
Now every pipeline run can git checkout the exact same orders.json without calling the model again.
6. Consumption in Tests (pytest example)
import json, pathlib, pytest
DATA_ROOT = pathlib.Path(__file__).parents[1] / "testdata" / "orders"
@pytest.fixture(scope="session")
def order_dataset(request):
seed = request.config.getoption("--test-data-seed") or "deadbeef"
path = DATA_ROOT / seed / "orders.json"
return json.loads(path.read_text())
def test_order_total_matches_items(order_dataset):
for order in order_dataset:
assert order["total_cents"] == compute_total(order)
Run with:
pytest --test-data-seed=a1b2c3d
All tests now see identical data, regardless of which agent executes them.
Tool Considerations
| Category | Options | What to evaluate |
|---|---|---|
| Model hosting | OpenAI API, Azure OpenAI, Anthropic, self‑hosted Llama‑3‑70B, Mistral | Latency, cost per 1k tokens, seed support, data‑privacy (no logging), model version pinning |
| Prompt management | Plain markdown in repo, PromptLayer, LangSmith, custom JSON | Version control, diffability, CI integration |
| Validation | jsonschema, pydantic, cerberus, Great Expectations | Schema + business‑rule coverage, speed, error messages |
| Artifact store | Git‑LFS, S3 + versioned keys, Artifactory, Nexus, DVC | Size limits, immutable writes, access control |
| CI integration | GitHub Actions, GitLab CI, CircleCI, Buildkite | Ability to pass seed as env var, caching of artifacts |
| Local dev helper | qa3-test-data-gen CLI (wraps the above steps) | One‑command generate --seed=deadbeef --prompt=order_v1.md |
Quick win: If you already have a QA3 account, the free test data generator at /tools/test-data-generator lets you paste a prompt, set a seed, and download a JSON artifact—no code required for the first prototype.
Validation Checklist (Run on Every Generation)
- Schema compliance – JSON Schema / Pydantic passes
- Cross‑field invariants – totals, date ranges, referential IDs
- Distribution checks – status ratios, categorical spreads (Chi‑square if you need statistical proof)
- PII / compliance – no real emails, SSNs, credit‑card numbers; regex sweep for known patterns
- Determinism proof – re‑run with same seed & model version → byte‑identical artifact (hash compare)
- Size budget – artifact < agreed limit (e.g., 5 MB) to keep repo lean
- Documentation – seed, model version, prompt hash recorded in
metadata.jsonalongside artifact
Automate the checklist as a gate in CI; a single failure blocks the artifact promotion.
Common Pitfalls & Mitigations
| Pitfall | Symptom | Mitigation |
|---|---|---|
| Model drift (new version released) | Previously passing tests start failing because output format changes | Pin model version (gpt-4o-2024-08-06). Store the version string in metadata.json. |
| Non‑deterministic sampling (temperature > 0) | Same seed yields different JSON on successive runs | Force temperature=0 (or top_p=0). Use provider‑native seed if available. |
| Prompt ambiguity | Model occasionally omits required fields | Add explicit “Output only the JSON array” and a negative example (“Do not wrap in markdown”). |
| Large artifact bloat | Repo size grows > 500 MB after months of seeds | Keep only the last N seeds (e.g., 30) + a “golden” baseline; purge older via CI cleanup job. |
| Hidden PII leakage | Real‑looking names/addresses appear in logs | Run a PII detector (Microsoft Presidio, AWS Comprehend) as a post‑step; reject artifact if hits > 0. |
| Cross‑service referential integrity | Generated customer.id does not exist in the user service stub | Include a reference manifest in the prompt (list of valid IDs) or generate IDs from a deterministic hash of the seed. |
| Rate‑limit / cost spikes | CI fails because API quota exhausted | Cache artifact after first successful generation; only regenerate on seed change. |
| Schema evolution | New field added to order model but prompt not updated | Add a prompt‑version field; CI job that diffs prompt hash vs. stored hash and alerts on mismatch. |
Scaling the Pattern
- Domain‑specific prompt library – One prompt per aggregate root (Order, Shipment, Invoice).
- Seed hierarchy – Global seed for the whole suite, per‑module sub‑seeds derived via HKDF (
subseed = HKDF(global_seed, info="orders")). Guarantees independence while staying deterministic. - Parallel generation – Split a 10 k‑record dataset into 10 prompts each producing 1 k records; combine artifacts.
- Contract testing – Publish the generated JSON schema as a consumer‑driven contract (Pact, Spring
Read more
Cost Model for AI Test Data Generation at Scale
A buyer-focused guide to “Cost Model for AI Test Data Generation at Scale,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.
Local LLM vs Hosted AI for Test Data Generation
A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.
AI Test Data Hallucinations: Detection and Guardrails
A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.