> ## Content Index
> Fetch the complete content index at: https://launch.ghostcms.templates.codememory.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Testing with realistic data, without copying real data
- URL: https://launch.ghostcms.templates.codememory.com/testing-with-realistic-data-without-copying-real-data/
- Published: 2026-06-11T14:00:00.000Z
- Updated: 2026-06-11T14:00:00.000Z
- Description: A generator built from the shape of customer data, never its content, and the eleven bugs it found.
- Author: Luis Ortega
- Tags: Engineering, #Import 2026-10-02 15:41

For a long time, we tested Tally with tiny, tidy data: ten users, a hundred events, everything in order. Real customers send millions of events, in the wrong order, with strange names and missing fields. Our tests passed; our customers found the bugs. This is how we fixed that without ever copying real customer data.

![Most of our worst bugs only appeared with data that looked like real life.](https://launch.ghostcms.templates.codememory.com/content/images/2026/10/test-desk.jpg)

Most of our worst bugs only appeared with data that looked like real life.

## The problem with tidy data

Small, clean test data hides three kinds of problems. Size problems: queries that are fast with a thousand rows and slow with a hundred million. Shape problems: events that arrive late, twice or with a timestamp in the future. Odd problems: names with emojis, properties with ten thousand different values, users who visit once a year.

The obvious answer is to test with a copy of production. We decided early never to do that. Our customers' data is theirs, and a copy in a test system is a copy that can leak.

## A generator instead

We built a data generator that produces realistic, fake data. It is based on statistics we collect about the shape of real data, never the content: how many events per user, how often events arrive late, how long names are, how many different values a property has.

```python
gen = Generator(seed=42)
gen.users(250_000, active_share=0.18)
gen.events(per_user='lognormal', late_share=0.03, duplicate_share=0.004)
gen.properties(cardinality={'plan': 4, 'page': 12_000, 'campaign': 900})
gen.write('test_dataset')
```

The seed makes every run the same, so a failing test fails every time. The numbers come from our monthly shape report.

> We do not need to know what our customers send. We need to know how it is shaped.  
>  
> **Luis Ortega**

![The large test dataset has 400 million events. It is rebuilt from scratch every week.](https://launch.ghostcms.templates.codememory.com/content/images/2026/10/test-db.jpg)

The large test dataset has 400 million events. It is rebuilt from scratch every week.

## Three sizes

We keep three datasets. A small one runs with every code change and finishes in a minute. A medium one runs every night. A large one, with four hundred million events, runs before every release and checks that no query got slower.

| Dataset | Events      | When it runs        | Time       |
| ------- | ----------- | ------------------- | ---------- |
| Small   | 50,000      | Every change        | 1 minute   |
| Medium  | 20 million  | Every night         | 25 minutes |
| Large   | 400 million | Before each release | 3 hours    |

## What it caught

In the first month, the generator found eleven bugs that customers had not reported yet, including a time-zone error that only appeared with events sent just after midnight on the last day of a month. It also found a query that took forty seconds on the large dataset; it now takes two.

💡

If you build a generator, make the oddness adjustable. Turning up the share of late or duplicate events is the fastest way to find the code that cannot handle them.

#### Do you ever look at real data to debug?

Only with the customer's written permission, for a specific problem, and never outside our production system.

#### Is the generator open source?

Not yet. We would like to share it once the code is less specific to Tally.