Quality is not optional. It's our standard. Free QA tools for testers and developers.

AI-Generated Addresses, Names, and Profiles: Realism vs Safety

AI‑Generated Addresses, Names, and Profiles: Realism vs Safety

When a test suite needs a thousand “customers” for a checkout flow, the fastest way to get them is to ask an LLM to spit out names, street addresses, phone numbers, and email addresses. The output looks convincing—proper capitalization, plausible ZIP codes, realistic‑looking phone formats.

But the moment those records hit a staging database that is shared with a downstream analytics pipeline, two problems surface:

ProblemWhy it mattersTypical symptom
PII leakageEven synthetic data can accidentally match a real person’s details (e.g., a common name + a real ZIP).Data‑privacy audit flags “potential PII” in non‑prod environments.
Domain‑specific invalidityAn address that passes a regex check may still be undeliverable (missing apartment number, wrong state‑ZIP pairing).Shipping‑service integration tests fail with “address not found” errors.

The tension is simple: realism helps tests exercise real‑world code paths; safety guarantees that no test run can expose or corrupt real data. The rest of this post gives you a repeatable way to decide where on that spectrum you need to land, how to evaluate a generator, and a concrete workflow you can copy into your next sprint.


1. Decision Criteria – What “Good Enough” Looks Like

Before you pick a tool or write a prompt, capture the constraints that matter for your context. Use the table below as a checklist during sprint planning; each row can be turned into a pass/fail gate in your CI pipeline.

#CriterionHow to measurePass threshold (example)
1Format complianceRegex / schema validation (e.g., US‑ZIP ^\d{5}(-\d{4})?$)100 % of generated rows pass
2Domain validityCross‑reference with authoritative lookup (USPS, GeoNames, libphonenumber)≥ 95 % of addresses resolve to a real‑world location
3UniquenessDuplicate detection across the generated set≤ 0.1 % duplicate full records
4PII‑collision riskRun a probabilistic match against a known PII corpus (e.g., HaveIBeenPwned email list)Zero exact matches; < 0.01 % fuzzy matches
5Statistical realismDistribution comparison (Chi‑square) against production histograms for name frequency, state population, etc.p‑value > 0.05 for each key attribute
6Determinism / seedabilitySame seed → identical outputRequired for reproducible test runs
7PerformanceGeneration time for 100 k rows< 30 s on a single CI node
8ExtensibilityAbility to add custom fields (loyalty tier, subscription plan) without code changesSupported via config or prompt templating

How to use the table

  1. Score each generator you evaluate (open‑source script, commercial SaaS, LLM prompt, QA3’s free test data generator at /tools/test-data-generator).
  2. Weight the criteria by project risk (e.g., a fintech KYC flow weights #4 and #2 higher than #7).
  3. Set a minimum total score that a generator must clear before it can be promoted to the shared test‑data library.

2. Evaluation Workflow – From Prompt to Pipeline

Below is a practical, repeatable workflow that turns the criteria above into an automated gate. It works whether you’re calling an LLM API, running a local Faker library, or using the QA3 generator.

2.1. Define a Data Contract

Create a JSON Schema (or Protobuf/Avro) that describes every column you need. Example for a “customer” entity:

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "title": "Customer",
  "type": "object",
  "required": ["id","firstName","lastName","email","phone","address"],
  "properties": {
    "id": {"type":"string","format":"uuid"},
    "firstName": {"type":"string","minLength":1,"maxLength":30},
    "lastName": {"type":"string","minLength":1,"maxLength":30},
    "email": {"type":"string","format":"email"},
    "phone": {"type":"string","pattern":"^\\+?[1-9]\\d{1,14}$"},
    "address": {
      "type":"object",
      "required":["street","city","state","zip","country"],
      "properties":{
        "street":{"type":"string"},
        "city":{"type":"string"},
        "state":{"type":"string","pattern":"^[A-Z]{2}$"},
        "zip":{"type":"string","pattern":"^\\d{5}(-\\d{4})?$"},
        "country":{"type":"string","enum":["US"]}
      }
    }
  }
}

Why a contract?

  • It gives you a single source of truth for validation.
  • It lets you swap generators without rewriting tests.
  • It makes the “format compliance” gate trivial (just run a schema validator).

2.2. Build a Generation Harness

A thin wrapper (Python, Node, Bash) that:

  1. Accepts a seed and row count.
  2. Calls the generator (LLM prompt, CLI, HTTP endpoint).
  3. Streams JSON lines to stdout.
  4. Returns a non‑zero exit code if any validation step fails.
#!/usr/bin/env bash


# generate_customers.sh  <seed> <count> <output-file>


SEED=$1
COUNT=$2
OUT=$3


# Example: call QA3 generator (replace with your own)


curl -s "https://qa3.io/tools/test-data-generator?seed=${SEED}&count=${COUNT}&schema=customer" \
  | jq -c . > "${OUT}.raw"


# 1️⃣ Schema validation


if ! python -m jsonschema -i "${OUT}.raw" customer_schema.json; then
  echo "❌ Schema validation failed"
  exit 1
fi


# 2️⃣ Domain validity (USPS address verification stub)


python verify_addresses.py "${OUT}.raw" > "${OUT}.verified" || exit 1


# 3️⃣ PII collision check


python pii_check.py "${OUT}.verified" || exit 1


# 4️⃣ Statistical sanity (compare to baseline histograms)


python stats_check.py "${OUT}.verified" || exit 1


mv "${OUT}.verified" "${OUT}"
echo "✅ Generation succeeded → ${OUT}"

Key points

  • Determinism: the same seed always yields the same file, making flaky tests impossible.
  • Fail‑fast: each gate stops the pipeline early, saving CI minutes.
  • Extensibility: add a new gate (e.g., “no profanity in names”) by dropping a new script into the chain.

2.3. CI Integration

Add a job to your pipeline (GitHub Actions, GitLab CI, Azure Pipelines) that runs the harness on every PR that touches test‑data definitions.



# .github/workflows/test-data.yml


name: Test Data Generation
on:
  pull_request:
    paths:
      - 'test-data/**'
jobs:
  generate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Set up Python
        uses: actions/setup-python@v5
        with: {python-version: '3.11'}
      - name: Install deps
        run: pip install -r requirements.txt
      - name: Generate & validate
        run: |
          ./generate_customers.sh ${{ github.sha }} 50000 artifacts/customers.jsonl
      - name: Upload artifact
        uses: actions/upload-artifact@v4
        with:
          name: customers
          path: artifacts/customers.jsonl

Now every PR that changes the schema, the prompt, or the generator version produces a verified, reproducible data set that downstream jobs can consume safely.


3. Worked Example – Generating 50 k US Customer Profiles

Below is a concrete end‑to‑end run using the QA3 free test data generator (the same steps apply to any generator you evaluate).

3.1. Prompt / Config



# qa3-config.yaml


schema: customer
count: 50000
seed: "2024-03-15-sprint-12"
locale: en_US
options:
  - name: "includeMiddleInitial"
    value: true
  - name: "phoneFormat"
    value: "E164"
  - name: "addressProvider"
    value: "usps"      # forces real‑world ZIP‑city‑state combos

3.2. Execution

$ ./generate_customers.sh 2024-03-15-sprint-12 50000 artifacts/customers.jsonl
✅ Generation succeeded → artifacts/customers.jsonl

3.3. Validation Results (sample)

GateMetricResult
Schema50 000 / 50 000 rows valid✅
USPS address lookup48 750 / 50 000 resolvable (97.5 %)✅ (threshold 95 %)
Duplicate full records3 duplicates (0.006 %)✅ (threshold 0.1 %)
PII collision (exact email)0 matches✅
Name‑frequency χ² vs prodp = 0.31✅
Generation time22 s✅ (threshold 30 s)

All gates pass, so the artifact is promoted to the shared test‑data bucket (s3://qa‑test‑data/customers/2024-03-15-sprint-12.jsonl). Downstream integration tests simply aws s3 cp the file at start‑up.

3.4. What If a Gate Fails?

Suppose the USPS lookup drops to 90 %. The harness exits non‑zero, the CI job fails, and the PR cannot be merged. The team then has three concrete levers:

  1. Switch provider – change addressProvider to geonames or a commercial address verification API.
  2. Relax the threshold – only if the product owner signs off that a 90 % deliverability rate is acceptable for the specific test (e.g., negative‑path testing).
  3. Post‑process – run a lightweight “address repair” script that fixes missing apartment numbers using a deterministic rule set, then re‑run the harness.

Because the harness is deterministic, you can iterate locally (./generate_customers.sh …) until the gates clear, then push the corrected config.


4. Common Pitfalls & How to Avoid Them

PitfallSymptomRoot CauseMitigation
Over‑reliance on LLM “creativity”Randomly formatted phone numbers, non‑existent ZIP‑city combosLLMs optimize for linguistic plausibility, not data‑quality rulesAlways wrap LLM output in a schema + domain validator (steps 1‑3 in the harness).
Hidden PII leakageAudit finds a real employee’s email in stagingGenerator seeded from a public dataset that contains real recordsUse synthetic‑only seed sources (e.g., Faker’s built‑in providers) and run a PII‑collision scan against known breach corpora.
Non‑deterministic outputFlaky tests that pass on one run, fail on anotherTemperature > 0, or generator uses system timePin temperature=0, pass explicit seed, and lock generator version (Docker image tag, npm package version).
Performance surpriseCI job exceeds 10 min for 200 k rowsSingle‑threaded Python Faker loop, no batchingBatch‑generate (e.g., 10 k rows per process) and parallelize; or use a compiled generator (Go, Rust) for high volume.
Schema driftNew column added to production DB but test data still oldNo automated contract testAdd a contract test that loads the latest production schema (via DB introspection) and asserts the test‑data schema is a superset.
Locale mismatchGerman street names appear in US address fieldGenerator defaulted to en_DE localeExplicitly set locale per entity; validate with a locale‑aware address library.
Regulatory blind spotGDPR‑style “right to be forgotten” test fails because synthetic data contains a real‑looking EU phone prefixNo region‑specific validationInclude regulatory rule packs (e.g., gdpr_phone_prefixes) in the validation stage.

Quick checklist for every new generator you onboard

  • Schema file committed alongside generator config
  • Deterministic seed documented in README
  • All eight decision‑criteria gates implemented in harness
  • CI job added with artifact upload
  • PII‑collision scan runs against at least two public breach lists
  • Performance baseline recorded (rows/sec)
  • Failure‑mode runbook linked (what to do when a gate fails)

5. Choosing the Right Tool for Your Context

Tool classTypical realismTypical safetyOperational costBest fit
LLM‑prompt (GPT‑4, Claude, local Llama)High (free‑form narrative)Low (needs heavy post‑validation)API cost + latencyExploratory data for UI copy, low‑volume prototypes
Faker / Factory‑Bot / Go‑fakerMedium (locale‑aware providers)High (deterministic, no external data)Zero (open source)Unit / integration tests, CI‑fast loops
Commercial synthetic‑data platforms (Tonic, Gretel, Mostly AI)Very high (statistical modeling)High (built‑in privacy guarantees)License + infraRegulated environments, large‑scale performance testing
QA3 free test data generator (/tools/test-data-generator)High (USPS‑validated addresses, realistic name distributions)High (seedable, schema‑driven, no PII sources)Free, no authTeams that need a quick, reproducible CSV/JSONL for CI without vendor lock‑in

Decision tip:
Start with the lowest‑cost option that satisfies your weighted criteria. If the free generator meets the thresholds for #1‑#5, you can stop there. Only move to a paid platform when you hit a hard ceiling (e.g., you need correlated multi‑table referential integrity at 10 M rows).


6. Next Steps – Put It Into Practice This Sprint

  1. Create a data contract for the most‑used entity in your test suite (customer, order, device, etc.). Commit the JSON Schema to test-data/schemas/.
  2. Spin up the generation harness (copy the Bash/Python skeleton above) and wire it to the QA3 free test data generator at /tools/test-data-generator.
  3. Add the CI job (GitHub Actions example) and run it on a feature branch. Verify that the artifact passes all eight gates.
  4. Document the seed you used (e.g., 2024-03-15-sprint-12) in the PR description so reviewers can reproduce the exact file locally.
  5. Run a retrospective after the first successful promotion: note any gate that felt too strict or too lax, adjust thresholds, and lock the new config as the baseline for the next sprint.

Action: Open /tools/test-data-generator, paste your schema, set a seed, generate 50 000 rows, and download the JSONL. Feed that file into the harness you just built. If the run finishes green, you have a reproducible, safety‑checked data set ready for every downstream test job.

Read more

AI Test Data Deduplication: Prompts and Post-Processing

A practical guide to “AI Test Data Deduplication: Prompts and Post-Processing,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.

Using AI to Expand a Small Test Dataset

A practical guide to “Using AI to Expand a Small Test Dataset,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.

Multilingual Test Data Generation with AI

A practical guide to “Multilingual Test Data Generation with AI,” with worked scenarios, tool considerations, validation checks, and actionable advice for QA teams.