Using AI to Generate Negative Test Data
Using AI to Generate Negative Test Data
Negative test data—inputs that should be rejected, cause errors, or expose edge‑case behaviour—are the backbone of resilient software. Yet most teams still hand‑craft them, which is slow, error‑prone, and rarely covers the combinatorial space that real users (or attackers) explore.
Large language models (LLMs) and specialised synthetic‑data services can now produce large, diverse negative datasets on demand. The challenge is not “can we generate it?” but “how do we generate useful negative data, validate it, and fold it into an existing CI/CD pipeline without creating noise?”
Below is a practical, evidence‑oriented workflow that QA engineers, test‑automation leads, and engineering managers can adopt today.
1. Why Negative Test Data Is Different
| Characteristic | Positive (happy‑path) data | Negative (error‑path) data |
|---|---|---|
| Goal | Verify that the system works as specified | Verify that the system fails safely and predictably |
| Volume | Usually a few representative records | Often needs thousands of permutations (boundary, type, length, encoding, semantic) |
| Maintenance | Low – schema changes rarely break happy paths | High – any schema or validation rule change can invalidate large swathes of negative cases |
| Risk of false positives | Low – a passing test is usually trustworthy | High – a test that expects a 400 but receives a 200 may hide a bug in the validation logic itself |
Because the risk profile is inverted, the generation process must be traceable (you need to know why a particular value was produced) and validatable (you must be able to confirm that the system really rejects it for the right reason).
2. Decision Criteria – When to Reach for an AI Generator
| Situation | AI‑generated negative data shines | Traditional hand‑crafted data is fine |
|---|---|---|
| Schema evolves weekly | ✔️ Regenerate in minutes | ✘ Manual updates become a bottleneck |
| Complex validation rules (regex, cross‑field, business logic) | ✔️ Prompt the model with the rule description | ✘ Hard to enumerate all violating combos manually |
| Need for adversarial / security‑oriented payloads (XSS, SQLi, command injection) | ✔️ Model can synthesize known attack patterns + variations | ✘ Requires deep security expertise per field |
| Small, static API (≤5 endpoints, <10 fields) | ✘ Overkill – a curated CSV is faster | ✔️ Simpler to maintain |
| Regulatory audit trail required | ✔️ If the generator logs prompts, seeds, and outputs | ✘ Hand‑crafted data often lacks provenance |
Rule of thumb: If you spend >30 min per sprint updating negative datasets, an AI generator will likely pay for itself in the first two sprints.
3. End‑to‑End Workflow
┌─────────────────────┐
│ 1. Capture rules │ (OpenAPI, JSON Schema, DB constraints, business docs)
└───────┬─────────────┘
▼
┌─────────────────────┐
│ 2. Prompt design │ (system + user prompts, few‑shot examples)
└───────┬─────────────┘
▼
┌─────────────────────┐
│ 3. Generate batch │ (LLM API or local model, temperature 0.2‑0.4)
└───────┬─────────────┘
▼
┌─────────────────────┐
│ 4. Validate & filter│ (schema check, rule‑engine, static analysis)
└───────┬─────────────┘
▼
┌─────────────────────┐
│ 5. Store & version │ (Git‑LFS, DVC, or artifact repo)
└───────┬─────────────┘
▼
┌─────────────────────┐
│ 6. Consume in CI │ (parameterised test runner, data‑driven framework)
└─────────────────────┘
3.1 Capture Rules – Make Them Machine‑Readable
- OpenAPI / Swagger –
type,format,maxLength,pattern,enum,minimum/maximum. - JSON Schema –
allOf,anyOf,dependencies,if/then/else. - Database constraints –
NOT NULL,CHECK, foreign‑key references. - Business rules – “discount % must be ≤ 100”, “start‑date < end‑date”, “email domain must be corporate”.
Export all of the above into a single rule bundle (JSON or YAML). This bundle becomes the source of truth for both the generator and the validator.
3.2 Prompt Design – The “Contract” Between You and the Model
A good prompt has three parts:
System: You are a test‑data engineer. Produce only JSON objects that violate at least one rule from the supplied rule bundle. Each object must include a "violation" field describing which rule is broken and why.
User:
Rule bundle:
{rule_bundle}
Examples:
[
{"field":"email","value":"not-an-email","violation":"email format regex"},
{"field":"age","value":-5,"violation":"minimum 0"}
]
Generate 200 distinct negative records for the "UserProfile" entity.
- Temperature – 0.2 – 0.4 keeps output deterministic while still exploring variations.
- Few‑shot count – 5‑10 examples usually enough; more can bias the model toward the examples.
- Output format – Enforce strict JSON (or CSV) via a schema validator in the next step.
3.3 Generation – Choose the Right Engine
| Engine | Pros | Cons | Typical Cost |
|---|---|---|---|
| OpenAI GPT‑4o / GPT‑4‑Turbo | Highest reasoning, good at cross‑field logic | Latency, per‑token cost, data‑privacy concerns | $0.01‑$0.03 / 1k tokens |
| Anthropic Claude 3.5 Sonnet | Strong instruction following, large context | Similar cost profile | $0.008‑$0.025 / 1k tokens |
| Local Llama‑3‑70B (via Ollama / vLLM) | No data leaves the network, flat cost | Requires GPU, slightly lower nuance on complex rules | Hardware amortisation |
| Specialised synthetic‑data SaaS (e.g., QA3 free test data generator) | Built‑in schema awareness, instant UI, export to CSV/JSON | Limited to supported rule types | Free tier up to 10k rows/month |
Tip: Start with the free QA3 test data generator at
/tools/test-data-generatorfor quick prototypes. It understands OpenAPI/JSON Schema out of the box and logs the prompt‑seed pair for auditability.
3.4 Validation & Filtering – Don’t Trust the Model Blindly
- Schema validation – Run each record through a JSON Schema validator (e.g.,
ajv). - Rule‑engine evaluation – Re‑apply the original rule bundle; keep only records that actually violate at least one rule.
- Deduplication – Hash the payload (excluding the
violationfield) and drop duplicates. - Semantic sanity check – For security payloads, run a lightweight static analyser (e.g.,
banditfor Python,eslint-plugin-securityfor JS) to confirm the payload is potentially malicious, not just syntactically wrong.
A typical pass‑rate after validation is 60‑80 %; the rest are either false positives (model thought it broke a rule but didn’t) or malformed JSON.
3.5 Storage & Versioning
- Git‑LFS for binary‑ish CSV/Parquet files.
- DVC if you want data‑pipeline reproducibility (
dvc add data/negative/). - Artifact repo (Nexus, Artifactory, GitHub Packages) for CI consumption.
Tag each dataset with:
negative-userprofile-v2024.03.15-<git‑sha>.parquet
Include a metadata.json alongside:
{
"generator": "gpt-4o-2024-05-13",
"prompt_sha": "a1f4c3e",
"rule_bundle_sha": "9f2b1d4",
"record_count": 1873,
"validation_pass_rate": 0.73
}
3.6 Consumption in CI
Most modern test runners support data‑driven tests:
- pytest –
@pytest.mark.parametrize("payload", load_negative("userprofile")) - JUnit 5 –
@ParameterizedTest @CsvFileSource(resources = "/negative/userprofile.csv") - Playwright / Cypress –
fixturesortestDataimports.
Wrap the test in a contract:
def test_userprofile_negative(payload, expected_violation):
resp = client.post("/api/userprofile", json=payload)
assert resp.status_code == 400
assert expected_violation in resp.json()["errors"]
If the system returns 200, the test fails fast and you know either the validation logic regressed or the generated payload was not truly negative (validation step missed it).
4. Worked Example – Generating Negative Data for a “Money Transfer” API
4.1 Rule Bundle (excerpt)
entity: MoneyTransfer
fields:
- name: sourceAccount
type: string
pattern: "^ACC-[0-9]{8}$"
required: true
- name: targetAccount
type: string
pattern: "^ACC-[0-9]{8}$"
required: true
- name: amount
type: number
minimum: 0.01
maximum: 1000000
required: true
- name: currency
type: string
enum: [USD, EUR, GBP, JPY]
required: true
- name: description
type: string
maxLength: 200
required: false
crossFieldRules:
- rule: "sourceAccount != targetAccount"
message: "Source and target must differ"
4.2 Prompt (sent to GPT‑4o)
System: You are a test‑data engineer. Output a JSON array of objects that violate at least one rule from the rule bundle. Each object must contain a "violation" field describing the broken rule.
User:
Rule bundle:
<the YAML above>
Examples:
[
{"sourceAccount":"ACC-12345678","targetAccount":"ACC-12345678","amount":100,"currency":"USD","violation":"sourceAccount != targetAccount"},
{"sourceAccount":"ACC-12345678","targetAccount":"ACC-87654321","amount":-5,"currency":"USD","violation":"amount minimum 0.01"},
{"sourceAccount":"ACC-12345678","targetAccount":"ACC-87654321","amount":100,"currency":"BTC","violation":"currency enum"}
]
Generate 300 distinct negative records for MoneyTransfer.
4.3 Generation Output (first 5 records)
[
{
"sourceAccount": "ACC-1234567",
"targetAccount": "ACC-87654321",
"amount": 500,
"currency": "USD",
"violation": "sourceAccount pattern ^ACC-[0-9]{8}$"
},
{
"sourceAccount": "ACC-12345678",
"targetAccount": "ACC-87654321",
"amount": 0,
"currency": "USD",
"violation": "amount minimum 0.01"
},
{
"sourceAccount": "ACC-12345678",
"targetAccount": "ACC-87654321",
"amount": 1500000,
"currency": "USD",
"violation": "amount maximum 1000000"
},
{
"sourceAccount": "ACC-12345678",
"targetAccount": "ACC-87654321",
"amount": 100,
"currency": "XYZ",
"violation": "currency enum"
},
{
"sourceAccount": "ACC-12345678",
"targetAccount": "ACC-12345678",
"amount": 100,
"currency": "USD",
"violation": "sourceAccount != targetAccount"
}
]
4.4 Validation Pipeline (Python snippet)
import json, jsonschema, hashlib
from pathlib import Path
RULE_BUNDLE = Path("rules/money_transfer.yaml")
SCHEMA = {
"type": "object",
"properties": {
"sourceAccount": {"type": "string", "pattern": "^ACC-[0-9]{8}$"},
"targetAccount": {"type": "string", "pattern": "^ACC-[0-9]{8}$"},
"amount": {"type": "number", "minimum": 0.01, "maximum": 1_000_000},
"currency": {"type": "string", "enum": ["USD","EUR","GBP","JPY"]},
"description": {"type": "string", "maxLength": 200},
"violation": {"type": "string"}
},
"required": ["sourceAccount","targetAccount","amount","currency","violation"],
"additionalProperties": False
}
def validate(record):
# 1. JSON Schema
jsonschema.validate(record, SCHEMA)
# 2. Cross‑field rule
if record["sourceAccount"] == record["targetAccount"]:
assert "sourceAccount != targetAccount" in record["violation"]
# 3. Return True if passes all checks
return True
raw = json.loads(Path("generated/money_transfer_raw.json").read_text())
valid = []
seen = set()
for rec in raw:
try:
if validate(rec):
key = hashlib.sha256(json.dumps({k:v for k,v in rec.items() if k!="violation"}, sort_keys=True).encode()).hexdigest()
if key not in seen:
seen.add(key)
valid.append(rec)
except Exception:
continue
Path("validated/money_transfer_negative.json").write_text(json.dumps(valid, indent=2))
print(f"Kept {len(valid)} / {len(raw)} records")
Result (typical run): Kept 212 / 300 records – ~70 % pass‑rate.
4.5 CI Test (pytest)
import pytest, json, requests
NEGATIVE_DATA = json.loads(Path("validated/money_transfer_negative.json").read_text())
@pytest.mark.parametrize("payload", NEGATIVE_DATA)
def test_money_transfer_negative(payload):
violation = payload.pop("violation")
resp = requests.post("http://api.test/money-transfer", json=payload, timeout=5)
assert resp.status_code == 400, f"Expected 400 for {violation}, got {resp.status_code}"
assert violation in resp.json()["errors"], f"Violation '{violation}' not reported"
Running this suite on every PR catches regressions such as a missing maximum check on amount or a new currency added without updating the enum.
5. Common Pitfalls & Mitigations
| Pitfall | Symptom | Mitigation |
|---|---|---|
| Model hallucinates a rule that doesn’t exist | Validation step discards > |
Read more
Local LLM vs Hosted AI for Test Data Generation
A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.
AI Test Data Hallucinations: Detection and Guardrails
A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.
Seeded AI Test Data Generation for Stable Automation
A practical guide to “Seeded AI Test Data Generation for Stable Automation,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.