Quality is not optional. It's our standard. Free QA tools for testers and developers.

AI Test Data Generation from API Specifications

AI Test Data Generation from API Specifications

Why this matters
Modern services are described by machine‑readable contracts (OpenAPI, gRPC protobuf, AsyncAPI, GraphQL SDL). Those contracts already encode the shape, constraints, and examples of every request and response. When a test‑data generator can read the contract, it can produce data that is by construction valid for the API version under test. The alternative—hand‑crafted fixtures or random generators that ignore the spec—creates a maintenance burden every time the contract changes.


1. Problem‑aware hook

You have a CI pipeline that runs contract tests, integration tests, and end‑to‑end suites against a dozen micro‑services. Each service publishes an OpenAPI 3.1 document. Your test suite currently relies on a handful of JSON files checked into the repo. When a new required field is added, the tests break until someone updates the fixtures. The same story repeats for enum extensions, format changes (e.g., date-time → date), and deprecated fields that must still be accepted for backward compatibility.

Result: flaky tests, wasted debugging time, and a growing “test‑data debt” that slows releases.


2. Decision criteria – when to adopt AI‑driven generation

CriterionLow‑effort fitHigh‑effort fit
Spec maturityStable OpenAPI/AsyncAPI with rich examples and format annotationsSpecs are incomplete, missing required, or use many anyOf/oneOf without discriminators
Data variability needsHundreds of permutations per endpoint (boundary, negative, locale)Only a handful of happy‑path cases
Team skill setFamiliar with prompt engineering, JSON Schema, and CI scriptingNo experience with LLM APIs or schema‑aware tooling
Regulatory / PII constraintsSynthetic data is acceptable; no production data leakage riskMust guarantee no inference of real user data (e.g., GDPR “right to be forgotten”)
Tooling budgetFree/open‑source generators or QA3’s free test data generator at /tools/test-data-generatorEnterprise‑grade platforms with governance dashboards

If you tick ≥3 of the “Low‑effort fit” boxes, an AI‑driven generator will likely pay off within the first sprint.


3. End‑to‑end workflow

┌─────────────────────┐
│ 1. Pull spec (CI)   │
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 2. Validate spec    │  (lint, resolve $refs, check required)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 3. Feed to generator│  (prompt + schema + policy)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 4. Post‑process     │  (type coercion, enum expansion, PII masking)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 5. Store / version  │  (artifact repo, git LFS, or test‑data service)
└───────┬─────────────┘
        ▼
┌─────────────────────┐
│ 6. Consume in tests │  (parameterised test runners, contract test harness)
└─────────────────────┘

Key automation points

StepAutomation tip
1‑2Add spectral or vacuum linting as a pre‑commit hook.
3Use a deterministic seed (e.g., hash(spec + ticket-id)) so the same spec version always yields the same dataset.
4Run a JSON‑Schema validator (ajv, jsonschema) on every generated payload before committing.
5Tag artifacts with spec-sha256 and generator-version.
6Parameterise pytest / JUnit / Playwright tests via a tiny data‑loader utility.

4. Worked example – generating data for a User‑Service OpenAPI spec

4.1 Minimal spec fragment

openapi: 3.1.0
info:
  title: User Service
  version: 1.4.0
paths:
  /users:
    post:
      operationId: createUser
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CreateUserRequest'
      responses:
        '201':
          description: Created
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/User'
components:
  schemas:
    CreateUserRequest:
      type: object
      required: [email, password, profile]
      properties:
        email:
          type: string
          format: email
          example: "alice@example.com"
        password:
          type: string
          format: password
          minLength: 12
          maxLength: 64
          writeOnly: true
        profile:
          $ref: '#/components/schemas/Profile'
        marketingOptIn:
          type: boolean
          default: false
    Profile:
      type: object
      required: [firstName, lastName, locale]
      properties:
        firstName:
          type: string
          minLength: 1
          maxLength: 50
        lastName:
          type: string
          minLength: 1
          maxLength: 50
        locale:
          type: string
          enum: [en-US, de-DE, fr-FR, ja-JP]
          default: en-US
        birthDate:
          type: string
          format: date
          nullable: true
    User:
      allOf:
        - $ref: '#/components/schemas/CreateUserRequest'
        - type: object
          required: [id, createdAt]
          properties:
            id:
              type: string
              format: uuid
            createdAt:
              type: string
              format: date-time

4.2 Prompt sent to the generator

You are a test‑data generator.  
Input: the OpenAPI 3.1 document above (provided as JSON).  
Policy:
- Produce 200 distinct `CreateUserRequest` payloads.
- Cover each `locale` enum value at least 30 times.
- Include 10 payloads with `birthDate` set to a future date (negative test).
- Include 5 payloads with an invalid email format (negative test).
- Ensure `password` meets the length constraints.
- Output a JSON array, one object per line (NDJSON).
- Seed: "user-service-2024-06-15".

4.3 Sample output (first three lines)

{"email":"alice.smith@example.com","password":"K9!vL2#mPqRz","profile":{"firstName":"Alice","lastName":"Smith","locale":"en-US","birthDate":"1992-03-14"},"marketingOptIn":false}
{"email":"bob.müller@beispiel.de","password":"Xy7@wQ9!zLpA","profile":{"firstName":"Bob","lastName":"Müller","locale":"de-DE","birthDate":null},"marketingOptIn":true}
{"email":"invalid-email","password":"Short1","profile":{"firstName":"Carol","lastName":"Ng","locale":"fr-FR","birthDate":"2030-01-01"},"marketingOptIn":false}

4.4 Validation checklist (run automatically)

  • Every object validates against CreateUserRequest JSON Schema (ajv --strict).
  • locale distribution matches policy (≥30 per value).
  • Exactly 10 future birthDate entries.
  • Exactly 5 malformed email entries.
  • No password shorter than 12 or longer than 64 chars.
  • Deterministic: re‑run with same seed → byte‑identical NDJSON file (SHA‑256 match).

If any check fails, the CI job aborts and surfaces the offending line number.


5. Tool considerations

CategoryRepresentative toolsStrengthsGaps
Schema‑aware CLIopenapi-generator (template‑based), prism (mock + data)Zero‑cost, works offline, deterministicLimited to built‑in fakers; no LLM‑level semantic variety
LLM‑backed generatorsgpt‑engine (custom prompts), langchain + jsonschemaHandles anyOf/oneOf, natural language policies, cross‑field constraintsRequires API key, latency, non‑deterministic unless seeded
Hybrid platformsQA3 free test data generator (/tools/test-data-generator), Mockoon + faker.js pluginsUI for policy, versioned artifacts, CI‑ready API, free tier for OpenAPI up to 500 endpointsAdvanced governance (role‑based access, audit log) only in paid tier
Data‑privacy focusedSynthesized, Tonic.aiDifferential privacy guarantees, PII detectionCommercial licences, heavier onboarding

Choosing a tool – start with the free QA3 generator if you:

  1. Have an OpenAPI/AsyncAPI file ≤ 500 endpoints.
  2. Need a reproducible CI artifact without managing an LLM account.
  3. Want a UI to tweak policies (enum coverage, negative‑case ratios) and instantly download NDJSON/CSV.

When you outgrow the free tier (e.g., >10 k endpoints, need audit trails), evaluate the hybrid platforms.


6. Validation & quality gates

GateImplementationFailure mode
Schema conformanceajv compile(schema).validate(payload) in a GitHub Action stepInvalid payload → test flakiness
Policy complianceCustom script counting enum values, date ranges, negative‑case markersCoverage gaps → missed bugs
Determinismsha256sum generated.ndjson compared to stored baselineNon‑deterministic generator → flaky CI
PII leakageRegex scan for known patterns (SSN, credit‑card) + presidio detectorReal data slips into synthetic set
PerformanceMeasure generation time < 30 s for 10 k rowsSlow generator blocks pipeline

Add these gates to the same pipeline that publishes the test‑data artifact. Treat the artifact as a first‑class deliverable—version it, sign it, and store it in an artifact repository (Nexus, Artifactory, GitHub Packages).


7. Common pitfalls & mitigations

PitfallWhy it hurtsMitigation
Spec drift – generator runs against stale specTests pass against old contract, fail in productionEnforce spec-sha256 check at generation time; fail if mismatch
Over‑reliance on example fieldsExamples often cover only happy pathExplicitly request negative cases in policy; supplement with faker for edge values
LLM hallucination of optional fieldsGenerates fields not defined in schema → validation errorsPost‑process with JSON‑Schema additionalProperties: false filter
Seed reuse across unrelated specsCollides data sets, makes debugging harderDerive seed per spec: hash(spec-content + ticket-id)
Ignoring writeOnly / readOnlySends id or createdAt on create request → 400Strip writeOnly:false / readOnly:true properties before generation
Large anyOf without discriminatorGenerator picks random branch → uneven coverageAdd a policy rule: “exercise each anyOf branch at least N times”
No cleanup of generated artifactsStorage bloat, stale data used accidentallyRetention policy: keep last 30 generations per spec version

8. Scaling the approach

  1. Modular spec repositories – each micro‑service owns its OpenAPI file; a monorepo CI job discovers **/openapi*.yaml and spins a generator matrix.
  2. Shared policy library – store reusable policy snippets (e.g., locale-coverage, future-date-negative) in a central policies/ folder; reference them by name in each generator invocation.
  3. Test‑data service – expose generated NDJSON via a lightweight HTTP API (GET /test-data/{spec}/{version}) so test runners can stream data on‑demand instead of checking out large files.
  4. Feedback loop – capture production error payloads (sanitized) and feed them back as “real‑world negative cases” into the policy for the next generation cycle.

9. Next action

  1. Export your current OpenAPI document (or AsyncAPI/GraphQL SDL) to a single JSON file.
  2. Open the free QA3 test data generator at /tools/test-data-generator.
  3. Paste the spec, choose a seed (e.g., project‑2024‑06‑15), and define a minimal policy:
    • 200 rows
    • Cover every enum at least 20×
    • Add 5 % malformed email and 2 % future date values.
  4. Download the NDJSON artifact, commit it alongside the spec (tagged with the spec SHA).
  5. Add a CI step that runs ajv validate -s spec.json -d test-data.ndjson on every push.

You’ll have a versioned, schema‑guaranteed, policy‑driven test data set ready for contract, integration, and performance tests—without maintaining a single hand‑crafted fixture.


Happy generating.

Read more

Using AI to Expand a Small Test Dataset

A practical guide to “Using AI to Expand a Small Test Dataset,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.

AI-Generated Addresses, Names, and Profiles: Realism vs Safety

A buyer-focused guide to “AI-Generated Addresses, Names, and Profiles: Realism vs Safety,” with concrete selection criteria, trade-offs, and an evaluation path QA teams can use.

Multilingual Test Data Generation with AI

A practical guide to “Multilingual Test Data Generation with AI,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.