Agentic Synthetic Data

Ask for the dataset.

Describe what you need. An agent profiles, generates and scores it — a table you can share when the real one can't leave.

  • Free to start
  • 100K rows/month
  • Nothing to install
dataxid — synthetic data
Profile this dataset and generate 10k synthetic rows.
Profiling 24 columns… detected 3 correlations. Generating with schema-aware model.
dataxid_generate · running
Correlation matrix24×24
0.94Fidelity
24/24Columns
10kRows
Pipeline

Train, generate, preview — in one pass.

The same loop every run: train on the shape, generate the rows, preview before anything downstream sees them.

129.8529.85DSLMonth-to-monthNo
3456.951,889.5DSLOne yearNo
253.85108.15DSLMonth-to-monthYes
4542.31,840.75DSLOne yearNo
1129.6346.45DSLMonth-to-monthNo
474.4306.6Fiber opticMonth-to-monthYes
66105.656,844.5Fiber opticTwo yearNo
2289.11,980.2Fiber opticMonth-to-monthNo
1049.55475.3DSLMonth-to-monthNo
28104.83,046.05Fiber opticOne yearYes
5270.153,645.5DSLTwo yearNo
799.65684.35Fiber opticMonth-to-monthYes

Showing 12 of 7,043 rows

Model Trainingfrom customers.csv
Succeeded
Rows trained
7K
Features
6
Epochs
50
Duration
1m 36s
Stopped
Full
Final loss — train 0.21 · val 0.24
Model Training
Training
  1. Analyzing data
  2. 2Training
  3. 3Finalizing
Epoch 38/5076%
train 0.36 · val 0.40Running
Control

Steer the mix, not just the row count.

Rebalance a category, close a bias gap, hold known values, fill the gaps — in the same conversation that built the table.

Make this 50/50

Override a skewed category until the mix is the one you asked for.

Before

Asked

Close the gap

Reduce a statistical parity gap across a sensitive attribute.

Keep what you know

Fix the columns you already have; the model completes the rest.

Fill the holes

Nulls become model predictions. Every non-null cell stays put.

Multi-table

A customer isn't a row.

It's a profile, a year of transactions, a handful of support tickets and every order they ever placed — four tables, held together by keys. That's the thing we generate.

Children learn their parents

Child rows are generated conditioned on the parent, so per-customer counts, ordering and timing carry over. That's the default here, not a flag you switch on.

Not a copy

Same shape, different values. Keys are assigned fresh and remapped, so nothing in the synthetic database points back at a real one.

One customers row and the three tables that hang off it. On the real side, customer 4102Ankara, premium, joined 2019 — has 28 orders, 63 transactions, 3 tickets. On the generated side, customer 7 İzmir, premium, joined 2019 — has 31 orders, 58 transactions, 2 tickets. Every child row repeats the key it points at. The marks beside each count are illustrative, not measured data.

customers

real4102

Ankarapremiumjoined 2019

generated7

İzmirpremiumjoined 2019

orders

spread over 12 months

real4102

28

generated7

31

transactions

salary on the 5th

real4102

63

generated7

58

tickets

mostly billing

real4102

3

generated7

2

synthesize_tables() · order resolved from your foreign keys · circular schemas rejected before training

Evaluation

Every generation comes with a report.

Fidelity and privacy are measured, not asserted. The same scores you can defend to a reviewer ship with every run.

0.94Distribution fidelity
24/24Columns matched
dataxid — evaluate
Profile this dataset and generate 10k synthetic rows.
Profiling 24 columns… detected 3 correlations. Generating with schema-aware model.
dataxid_generate · running
Correlation matrix24×24
0.94Fidelity
24/24Columns
10kRows

Fidelity with breakdown

Univariate and bivariate fidelity, with a per-column breakdown.

Privacy you can audit

Distance to closest record, identical match share, and NNDR — then open the full instrument.

Open SynthEval
Your data

Bring the data you have.

Upload a file, connect a database, or start with the sample dataset. Everything in your workspace is ready the moment you ask.

CSV upload

Drop a file in; it is a table you can synthesize right away.

PostgreSQL / MySQL

Connect a database and work with its tables directly.

Sample dataset

Start with a sample table and see the whole loop in a minute.

Start with your own dataset.

Free to start. Upload a file or try the sample.