Structured Output Prompts for JSON Test Data
Structured Output Prompts for JSON Test Data
When a test suite needs realistic payloads, the fastest way to get them is to ask a language model for JSON that matches a schema. The catch is that “just give me JSON” rarely produces something you can drop straight into a test harness. Missing fields, wrong types, or extra commentary break parsers and waste the time you tried to save.
This post shows how to write structured‑output prompts that reliably return valid JSON, how to validate the result, and where the approach fits in a QA workflow. You’ll see a complete worked example, a checklist you can copy into your repo, and a few pitfalls that trip up even experienced engineers.
1. Why Plain Prompts Fail
| Symptom | Typical Cause | Impact on Test Automation |
|---|---|---|
Extra markdown fences (```json) | Model defaults to code block formatting | Parser throws “unexpected token” |
| Trailing commas or comments | Model mimics JavaScript object literals | Strict JSON parsers reject the payload |
| Missing required fields | Prompt didn’t enumerate the schema | Tests fail with “field X is required” |
| Wrong data types (string vs. number) | Model guesses based on description | Schema validation errors, flaky tests |
| Hallucinated enum values | Model invents values not in the spec | Business‑logic tests exercise impossible paths |
A structured‑output prompt removes the ambiguity by giving the model a contract: a JSON Schema (or a concise equivalent) plus explicit formatting rules. The model then treats the task as “fill in the blanks” rather than “write a document”.
2. Decision Criteria – When to Use Structured Prompts
| Situation | Structured Prompt Good Fit? | Alternative |
|---|---|---|
| Need dozens of variations for a single API contract | ✅ | Hand‑crafted fixtures |
| Schema changes frequently (e.g., evolving GraphQL types) | ✅ | Code‑generation from OpenAPI |
Data must obey cross‑field constraints (e.g., endDate > startDate) | ⚠️ (add post‑generation validation) | Custom generator script |
| One‑off exploratory testing | ❌ | Quick manual JSON edit |
| High‑volume load‑test data (millions of rows) | ❌ | Dedicated data‑factory / DB seeding |
Rule of thumb: if you can express the shape in a JSON Schema (or a compact TypeScript interface) and you need hundreds of distinct samples, a structured prompt is the lowest‑effort path.
3. Prompt Anatomy
A reliable prompt contains four sections:
- Role & Goal – Tell the model who it is and what you expect.
- Schema Definition – Provide a machine‑readable description (JSON Schema, TypeScript, or a concise bullet list).
- Formatting Rules – Explicitly forbid markdown, comments, trailing commas, etc.
- Examples – One or two fully‑valid instances to anchor the model’s style.
3.1 Minimal Template
You are a test‑data generator. Output ONLY a JSON object that conforms to the schema below.
Do not include markdown fences, comments, or any explanatory text.
Schema (JSON Schema Draft‑07):
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["userId", "email", "roles"],
"properties": {
"userId": { "type": "string", "format": "uuid" },
"email": { "type": "string", "format": "email" },
"roles": { "type": "array", "items": { "type": "string", "enum": ["admin","editor","viewer"] }, "minItems": 1 }
},
"additionalProperties": false
}
Example:
{
"userId": "3f2a1c9e-7b4d-4a6e-9c1d-2f8e6b9a4c3d",
"email": "alice@example.com",
"roles": ["editor"]
}
Copy the template, replace the schema, and you have a ready‑to‑use prompt.
3.2 Adding Cross‑Field Rules
JSON Schema cannot express “endDate must be after startDate”. Two practical options:
| Option | How to Encode | Pros | Cons |
|---|---|---|---|
| Prompt‑level instruction | Add a line: Rule: endDate must be a date‑time strictly later than startDate. | No extra tooling | Model may still violate it |
| Post‑generation validator | Write a tiny script (Node, Python, Go) that loads the JSON and asserts the rule. | Guarantees correctness | Extra CI step |
Recommendation: keep the prompt simple, enforce complex constraints in a validator that runs immediately after generation.
4. Worked Example – Generating Order Payloads for an E‑Commerce API
4.1 Target Schema (simplified)
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["orderId", "customer", "items", "totalAmount", "currency", "placedAt"],
"properties": {
"orderId": { "type": "string", "format": "uuid" },
"customer": {
"type": "object",
"required": ["customerId", "email", "tier"],
"properties": {
"customerId": { "type": "string", "format": "uuid" },
"email": { "type": "string", "format": "email" },
"tier": { "type": "string", "enum": ["bronze","silver","gold"] }
},
"additionalProperties": false
},
"items": {
"type": "array",
"minItems": 1,
"maxItems": 10,
"items": {
"type": "object",
"required": ["sku", "quantity", "unitPrice"],
"properties": {
"sku": { "type": "string", "pattern": "^[A-Z]{3}-\\d{4}$" },
"quantity": { "type": "integer", "minimum": 1, "maximum": 99 },
"unitPrice": { "type": "number", "minimum": 0.01, "multipleOf": 0.01 }
},
"additionalProperties": false
}
},
"totalAmount": { "type": "number", "minimum": 0.01, "multipleOf": 0.01 },
"currency": { "type": "string", "enum": ["USD","EUR","GBP"] },
"placedAt": { "type": "string", "format": "date-time" }
},
"additionalProperties": false
}
4.2 Full Prompt
You are a test‑data generator for an e‑commerce order service.
Output ONLY a JSON object that conforms to the schema below.
Do not include markdown fences, comments, or any explanatory text.
Schema (JSON Schema Draft‑07):
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["orderId", "customer", "items", "totalAmount", "currency", "placedAt"],
"properties": {
"orderId": { "type": "string", "format": "uuid" },
"customer": {
"type": "object",
"required": ["customerId", "email", "tier"],
"properties": {
"customerId": { "type": "string", "format": "uuid" },
"email": { "type": "string", "format": "email" },
"tier": { "type": "string", "enum": ["bronze","silver","gold"] }
},
"additionalProperties": false
},
"items": {
"type": "array",
"minItems": 1,
"maxItems": 10,
"items": {
"type": "object",
"required": ["sku", "quantity", "unitPrice"],
"properties": {
"sku": { "type": "string", "pattern": "^[A-Z]{3}-\\d{4}$" },
"quantity": { "type": "integer", "minimum": 1, "maximum": 99 },
"unitPrice": { "type": "number", "minimum": 0.01, "multipleOf": 0.01 }
},
"additionalProperties": false
}
},
"totalAmount": { "type": "number", "minimum": 0.01, "multipleOf": 0.01 },
"currency": { "type": "string", "enum": ["USD","EUR","GBP"] },
"placedAt": { "type": "string", "format": "date-time" }
},
"additionalProperties": false
}
Rules:
- totalAmount must equal the sum of (quantity * unitPrice) for all items, rounded to two decimals.
- placedAt must be within the last 30 days.
- Use realistic looking SKUs (e.g., "ABC-1234") and email domains (example.com, test.org).
Example:
{
"orderId": "d2f1a3b4-5c6d-7e8f-9a0b-1c2d3e4f5a6b",
"customer": {
"customerId": "a1b2c3d4-e5f6-7a8b-9c0d-1e2f3a4b5c6d",
"email": "john.doe@example.com",
"tier": "silver"
},
"items": [
{ "sku": "ABC-1234", "quantity": 2, "unitPrice": 19.99 },
{ "sku": "XYZ-9876", "quantity": 1, "unitPrice": 149.50 }
],
"totalAmount": 189.48,
"currency": "USD",
"placedAt": "2024-02-15T14:32:10Z"
}
4.3 Generation Loop (pseudo‑code)
import json, subprocess, sys, datetime, math
from jsonschema import validate, ValidationError
PROMPT = open("order_prompt.txt").read()
def call_llm(prompt: str) -> str:
# Replace with your provider CLI / SDK
return subprocess.check_output(["llm", "generate", "--prompt", prompt], text=True)
def validate_order(obj: dict) -> bool:
schema = json.load(open("order_schema.json"))
try:
validate(instance=obj, schema=schema)
except ValidationError as e:
print("Schema error:", e.message)
return False
# Cross‑field checks
calc = sum(item["quantity"] * item["unitPrice"] for item in obj["items"])
if round(calc, 2) != round(obj["totalAmount"], 2):
print(f"Total mismatch: calculated {calc}, got {obj['totalAmount']}")
return False
placed = datetime.datetime.fromisoformat(obj["placedAt"].replace("Z", "+00:00"))
if placed < datetime.datetime.now(datetime.timezone.utc) - datetime.timedelta(days=30):
print("placedAt older than 30 days")
return False
return True
def generate_one() -> dict:
raw = call_llm(PROMPT)
# Strip accidental markdown fences if any
raw = raw.strip().strip("`")
if raw.startswith("json"):
raw = raw[4:].strip()
return json.loads(raw)
def main(count: int):
for i in range(count):
for attempt in range(3):
data = generate_one()
if validate_order(data):
print(json.dumps(data))
break
else:
print(f"Failed to produce valid order after 3 attempts (iteration {i})", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main(int(sys.argv[1]) if len(sys.argv) > 1 else 10)
What the script does
- Sends the prompt to the model (replace
llm generatewith your provider). - Strips any stray markdown fences.
- Validates against the JSON Schema and the two business rules.
- Retries up to three times per record.
Running python gen_orders.py 50 yields 50 ready‑to‑use order payloads.
5. Validation Checklist – Drop This Into Your Repo
# Test‑Data Generation Validation Checklist
- [ ] **Schema file** (`order_schema.json`) is version‑controlled and matches the API contract.
- [ ] **Prompt file** (`order_prompt.txt`) references the exact schema (copy‑paste or `$ref`).
- [ ] **Formatter** removes markdown fences, trailing commas, comments.
- [ ] **Validator** runs JSON Schema validation *and* all cross‑field rules.
- [ ] **Retry logic** (≥2 attempts) before failing the CI job.
- [ ] **Deterministic seed** (if your provider supports `seed`/`temperature=0`) for reproducibility.
- [ ] **Output directory** (`generated/orders/`) is git‑ignored; CI archives it as an artifact.
- [ ] **Sample count** parameterized (`make gen-orders COUNT=200`).
- [ ] **Documentation** in `README.md` explains how to regenerate and how to add new rules.
Copy the checklist into TEST_DATA_CHECKLIST.md and tick items during code review.
6. Tooling Landscape
| Category | Tool | Why It Helps |
|---|---|---|
| LLM CLI | llm (Simon Willison), openai CLI, ollama run | Scriptable, supports --prompt-file |
| Schema Validation | jsonschema (Python), ajv (Node), go-jsonschema | Fast, strict Draft‑07/2020‑12 support |
| Prompt Management | promptfoo, langchain prompt templates | Version‑control prompts, CI integration |
| Data Generation UI | QA3 free test data generator – /tools/test-data-generator | Quick one‑off generation without code |
| CI Integration | GitHub Actions, GitLab CI, CircleCI | Run generation + validation on every PR |
Practical tip: Keep the LLM call outside the CI critical path. Generate data locally or in a nightly job, commit the JSON fixtures, and let CI only run the validator. This avoids flaky network calls and token‑cost surprises.
7. Common Pitfalls & Mitigations
| Pitfall | Symptom | Mitigation |
|---|---|---|
Model ignores additionalProperties: false | Extra fields appear in output | Add explicit rule: Do not output any property not defined in the schema. |
| Enum values rendered as numbers | "tier": 2 instead of "silver" | In schema, set "type": "string" for enums; add rule All enum values must be quoted strings. |
| Date‑time format mismatch | 2024-02-15 14:32:10 (space) | Require ISO‑8601 with Z suffix; add example with Z. |
| Floating‑point rounding | totalAmount: 189.48000000000002 | Rule: Round totalAmount to two decimal places. + validator multipleOf: 0.01. |
| Token limit truncation | Incomplete JSON for large arrays | Split generation: ask for one order per call, or use a streaming API. |
| Non‑deterministic output | Same prompt yields different structures | Set temperature=0 and a fixed seed if provider supports it. |
| Schema drift | API adds a required field, generator still produces old shape | Add a CI step that diffs the schema against the OpenAPI spec; fail if out‑of‑sync. |
8. Scaling the Approach
| Scale | Strategy |
|---|---|
| < 100 records / run | Single‑prompt loop (as shown). |
| 1 000 – 10 000 | Parallelize prompt calls (e.g., xargs -P 8), keep validator lightweight. |
| > 10 000 | Switch to a template‑based generator (J |
Read more
Cost Model for AI Test Data Generation at Scale
A buyer-focused guide to “Cost Model for AI Test Data Generation at Scale,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.
Local LLM vs Hosted AI for Test Data Generation
A buyer-focused guide to “Local LLM vs Hosted AI for Test Data Generation,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.
AI Test Data Hallucinations: Detection and Guardrails
A practical risk review of “AI Test Data Hallucinations: Detection and Guardrails,” with warning signs, safeguards, and fixes for real QA workflows.