The Synthetic Data Guide

How synthetic data is replacing risky real data for AI, analytics, testing, and compliance — without the privacy headaches.

Explore by topic

All Guides

13 guides

Fundamentals 10 min

How is synthetic data generated?

The main approaches to synthetic data generation — rule-based, model-based, simulation, and agentic — and how to choose the right one.

Fundamentals 10 min

Types of synthetic data: tabular, time-series, and unstructured

The main types of synthetic data — tabular, time-series, and unstructured — with how each is generated and where to use it.

Fundamentals 9 min

What is synthetic data?

Synthetic data is machine-generated data that mimics real data without using real records.

Use Cases 9 min

Synthetic data for financial services

See how banks, fintechs, and insurers use synthetic data for fraud detection, credit risk modeling, and compliant testing.

Use Cases 9 min

Synthetic data for machine learning and AI

How synthetic data trains machine learning and AI models — why teams use it, how it enters the pipeline, and where it works best.

Use Cases 9 min

Synthetic data for software testing and QA

How QA and engineering teams use synthetic test data to test apps safely, cover edge cases, and ship faster without production data.

Use Cases 9 min

Synthetic data for healthcare: HIPAA use cases and what's possible

How synthetic data supports HIPAA-context healthcare workflows — EHR testing, claims processing, interoperability, and AI model training.

Comparisons 11 min

The synthetic data tools landscape: approaches and categories

A map of the synthetic data tools landscape: the main generation approaches, tool categories, and how they differ from data de-identification.

Comparisons 10 min

Synthetic data vs. real (production) data

Compare synthetic and real production data on fidelity, privacy, and availability — and learn when to use each, or both, for your projects.

How it works 8 min

Mock data and synthetic APIs for development and testing

Learn what mock data and synthetic APIs are, how they work, and when to use them to build and test software faster.

How it works 10 min

Generating synthetic data from an existing database

How to generate synthetic data from an existing database by seeding net-new records from its schema and relationships, with referential integrity intact.

How it works 10 min

Synthetic data and privacy: re-identification risk and safeguards

Synthetic data isn't automatically private.

How it works 11 min

How to measure synthetic data quality and fidelity

Learn how to measure synthetic data quality across fidelity, utility, and privacy — which metrics matter and how to evaluate your generated output.