Quality is not optional. It's our standard. Free QA tools for testers and developers.

Using AI to Generate Negative Test Data

QTQA3 Team

Using AI to Generate Negative Test Data

Negative test data—inputs that should be rejected, cause errors, or expose edge‑case behaviour—are the backbone of resilient software. Yet most teams still hand‑craft them, which is slow, error‑prone, and rarely covers the combinatorial space that real users (or attackers) explore.

Large language models (LLMs) and specialised synthetic‑data services can now produce large, diverse negative datasets on demand. The challenge is not “can we generate it?” but “how do we generate useful negative data, validate it, and fold it into an existing CI/CD pipeline without creating noise?”

Below is a practical, evidence‑oriented workflow that QA engineers, test‑automation leads, and engineering managers can adopt today.


1. Why Negative Test Data Is Different

CharacteristicPositive (happy‑path) dataNegative (error‑path) data
GoalVerify that the system works as specifiedVerify that the system fails safely and predictably
VolumeUsually a few representative recordsOften needs thousands of permutations (boundary, type, length, encoding, semantic)
MaintenanceLow – schema changes rarely break happy pathsHigh – any schema or validation rule change can invalidate large swathes of negative cases
Risk of false positivesLow – a passing test is usually trustworthyHigh – a test that expects a 400 but receives a 200 may hide a bug in the validation logic itself

Because the risk profile is inverted, the generation process must be traceable (you need to know why a particular value was produced) and validatable (you must be able to confirm that the system really rejects it for the right reason).


2. Decision Criteria – When to Reach for an AI Generator

SituationAI‑generated negative data shinesTraditional hand‑crafted data is fine
Schema evolves weekly✔️ Regenerate in minutes✘ Manual updates become a bottleneck
Complex validation rules (regex, cross‑field, business logic)✔️ Prompt the model with the rule description✘ Hard to enumerate all violating combos manually
Need for adversarial / security‑oriented payloads (XSS, SQLi, command injection)✔️ Model can synthesize known attack patterns + variations✘ Requires deep security expertise per field
Small, static API (≤5 endpoints, <10 fields)✘ Overkill – a curated CSV is faster✔️ Simpler to maintain
Regulatory audit trail required✔️ If the generator logs prompts, seeds, and outputs✘ Hand‑crafted data often lacks provenance

Rule of thumb: If you spend >30 min per sprint updating negative datasets, an AI generator will likely pay for itself in the first two sprints.


3. End‑to‑End Workflow

┌─────────────────────┐
│ 1. Capture rules    │  (OpenAPI, JSON Schema, DB constraints, business docs)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 2. Prompt design    │  (system + user prompts, few‑shot examples)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 3. Generate batch   │  (LLM API or local model, temperature 0.2‑0.4)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 4. Validate & filter│  (schema check, rule‑engine, static analysis)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 5. Store & version  │  (Git‑LFS, DVC, or artifact repo)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 6. Consume in CI    │  (parameterised test runner, data‑driven framework)
└─────────────────────┘

3.1 Capture Rules – Make Them Machine‑Readable

  • OpenAPI / Swagger – type, format, maxLength, pattern, enum, minimum/maximum.
  • JSON Schema – allOf, anyOf, dependencies, if/then/else.
  • Database constraints – NOT NULL, CHECK, foreign‑key references.
  • Business rules – “discount % must be ≤ 100”, “start‑date < end‑date”, “email domain must be corporate”.

Export all of the above into a single rule bundle (JSON or YAML). This bundle becomes the source of truth for both the generator and the validator.

3.2 Prompt Design – The “Contract” Between You and the Model

A good prompt has three parts:

System: You are a test‑data engineer. Produce only JSON objects that violate at least one rule from the supplied rule bundle. Each object must include a "violation" field describing which rule is broken and why.


User: 
Rule bundle:
{rule_bundle}


Examples:
[
  {"field":"email","value":"not-an-email","violation":"email format regex"},
  {"field":"age","value":-5,"violation":"minimum 0"}
]


Generate 200 distinct negative records for the "UserProfile" entity.
  • Temperature – 0.2 – 0.4 keeps output deterministic while still exploring variations.
  • Few‑shot count – 5‑10 examples usually enough; more can bias the model toward the examples.
  • Output format – Enforce strict JSON (or CSV) via a schema validator in the next step.

3.3 Generation – Choose the Right Engine

EngineProsConsTypical Cost
OpenAI GPT‑4o / GPT‑4‑TurboHighest reasoning, good at cross‑field logicLatency, per‑token cost, data‑privacy concerns$0.01‑$0.03 / 1k tokens
Anthropic Claude 3.5 SonnetStrong instruction following, large contextSimilar cost profile$0.008‑$0.025 / 1k tokens
Local Llama‑3‑70B (via Ollama / vLLM)No data leaves the network, flat costRequires GPU, slightly lower nuance on complex rulesHardware amortisation
Specialised synthetic‑data SaaS (e.g., QA3 free test data generator)Built‑in schema awareness, instant UI, export to CSV/JSONLimited to supported rule typesFree tier up to 10k rows/month

Tip: Start with the free QA3 test data generator at /tools/test-data-generator for quick prototypes. It understands OpenAPI/JSON Schema out of the box and logs the prompt‑seed pair for auditability.

3.4 Validation & Filtering – Don’t Trust the Model Blindly

  1. Schema validation – Run each record through a JSON Schema validator (e.g., ajv).
  2. Rule‑engine evaluation – Re‑apply the original rule bundle; keep only records that actually violate at least one rule.
  3. Deduplication – Hash the payload (excluding the violation field) and drop duplicates.
  4. Semantic sanity check – For security payloads, run a lightweight static analyser (e.g., bandit for Python, eslint-plugin-security for JS) to confirm the payload is potentially malicious, not just syntactically wrong.

A typical pass‑rate after validation is 60‑80 %; the rest are either false positives (model thought it broke a rule but didn’t) or malformed JSON.

3.5 Storage & Versioning

  • Git‑LFS for binary‑ish CSV/Parquet files.
  • DVC if you want data‑pipeline reproducibility (dvc add data/negative/).
  • Artifact repo (Nexus, Artifactory, GitHub Packages) for CI consumption.

Tag each dataset with:

negative-userprofile-v2024.03.15-<git‑sha>.parquet

Include a metadata.json alongside:

{
  "generator": "gpt-4o-2024-05-13",
  "prompt_sha": "a1f4c3e",
  "rule_bundle_sha": "9f2b1d4",
  "record_count": 1873,
  "validation_pass_rate": 0.73
}

3.6 Consumption in CI

Most modern test runners support data‑driven tests:

  • pytest – @pytest.mark.parametrize("payload", load_negative("userprofile"))
  • JUnit 5 – @ParameterizedTest @CsvFileSource(resources = "/negative/userprofile.csv")
  • Playwright / Cypress – fixtures or testData imports.

Wrap the test in a contract:

def test_userprofile_negative(payload, expected_violation):
    resp = client.post("/api/userprofile", json=payload)
    assert resp.status_code == 400
    assert expected_violation in resp.json()["errors"]

If the system returns 200, the test fails fast and you know either the validation logic regressed or the generated payload was not truly negative (validation step missed it).


4. Worked Example – Generating Negative Data for a “Money Transfer” API

4.1 Rule Bundle (excerpt)

entity: MoneyTransfer
fields:
  - name: sourceAccount
    type: string
    pattern: "^ACC-[0-9]{8}$"
    required: true
  - name: targetAccount
    type: string
    pattern: "^ACC-[0-9]{8}$"
    required: true
  - name: amount
    type: number
    minimum: 0.01
    maximum: 1000000
    required: true
  - name: currency
    type: string
    enum: [USD, EUR, GBP, JPY]
    required: true
  - name: description
    type: string
    maxLength: 200
    required: false
crossFieldRules:
  - rule: "sourceAccount != targetAccount"
    message: "Source and target must differ"

4.2 Prompt (sent to GPT‑4o)

System: You are a test‑data engineer. Output a JSON array of objects that violate at least one rule from the rule bundle. Each object must contain a "violation" field describing the broken rule.


User:
Rule bundle:
<the YAML above>


Examples:
[
  {"sourceAccount":"ACC-12345678","targetAccount":"ACC-12345678","amount":100,"currency":"USD","violation":"sourceAccount != targetAccount"},
  {"sourceAccount":"ACC-12345678","targetAccount":"ACC-87654321","amount":-5,"currency":"USD","violation":"amount minimum 0.01"},
  {"sourceAccount":"ACC-12345678","targetAccount":"ACC-87654321","amount":100,"currency":"BTC","violation":"currency enum"}
]


Generate 300 distinct negative records for MoneyTransfer.

4.3 Generation Output (first 5 records)

[
  {
    "sourceAccount": "ACC-1234567",
    "targetAccount": "ACC-87654321",
    "amount": 500,
    "currency": "USD",
    "violation": "sourceAccount pattern ^ACC-[0-9]{8}$"
  },
  {
    "sourceAccount": "ACC-12345678",
    "targetAccount": "ACC-87654321",
    "amount": 0,
    "currency": "USD",
    "violation": "amount minimum 0.01"
  },
  {
    "sourceAccount": "ACC-12345678",
    "targetAccount": "ACC-87654321",
    "amount": 1500000,
    "currency": "USD",
    "violation": "amount maximum 1000000"
  },
  {
    "sourceAccount": "ACC-12345678",
    "targetAccount": "ACC-87654321",
    "amount": 100,
    "currency": "XYZ",
    "violation": "currency enum"
  },
  {
    "sourceAccount": "ACC-12345678",
    "targetAccount": "ACC-12345678",
    "amount": 100,
    "currency": "USD",
    "violation": "sourceAccount != targetAccount"
  }
]

4.4 Validation Pipeline (Python snippet)

import json, jsonschema, hashlib
from pathlib import Path


RULE_BUNDLE = Path("rules/money_transfer.yaml")
SCHEMA = {
    "type": "object",
    "properties": {
        "sourceAccount": {"type": "string", "pattern": "^ACC-[0-9]{8}$"},
        "targetAccount": {"type": "string", "pattern": "^ACC-[0-9]{8}$"},
        "amount": {"type": "number", "minimum": 0.01, "maximum": 1_000_000},
        "currency": {"type": "string", "enum": ["USD","EUR","GBP","JPY"]},
        "description": {"type": "string", "maxLength": 200},
        "violation": {"type": "string"}
    },
    "required": ["sourceAccount","targetAccount","amount","currency","violation"],
    "additionalProperties": False
}


def validate(record):
    # 1. JSON Schema
    jsonschema.validate(record, SCHEMA)
    # 2. Cross‑field rule
    if record["sourceAccount"] == record["targetAccount"]:
        assert "sourceAccount != targetAccount" in record["violation"]
    # 3. Return True if passes all checks
    return True


raw = json.loads(Path("generated/money_transfer_raw.json").read_text())
valid = []
seen = set()
for rec in raw:
    try:
        if validate(rec):
            key = hashlib.sha256(json.dumps({k:v for k,v in rec.items() if k!="violation"}, sort_keys=True).encode()).hexdigest()
            if key not in seen:
                seen.add(key)
                valid.append(rec)
    except Exception:
        continue


Path("validated/money_transfer_negative.json").write_text(json.dumps(valid, indent=2))
print(f"Kept {len(valid)} / {len(raw)} records")

Result (typical run): Kept 212 / 300 records – ~70 % pass‑rate.

4.5 CI Test (pytest)

import pytest, json, requests


NEGATIVE_DATA = json.loads(Path("validated/money_transfer_negative.json").read_text())


@pytest.mark.parametrize("payload", NEGATIVE_DATA)
def test_money_transfer_negative(payload):
    violation = payload.pop("violation")
    resp = requests.post("http://api.test/money-transfer", json=payload, timeout=5)
    assert resp.status_code == 400, f"Expected 400 for {violation}, got {resp.status_code}"
    assert violation in resp.json()["errors"], f"Violation '{violation}' not reported"

Running this suite on every PR catches regressions such as a missing maximum check on amount or a new currency added without updating the enum.


5. Common Pitfalls & Mitigations

PitfallSymptomMitigation
Model hallucinates a rule that doesn’t existValidation step discards >

Read more

Local LLM vs Hosted AI for Test Data Generation

A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.

AI Test Data Hallucinations: Detection and Guardrails

A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.

Seeded AI Test Data Generation for Stable Automation

A practical guide to “Seeded AI Test Data Generation for Stable Automation,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.