Testing with realistic data, without copying real data
A generator built from the shape of customer data, never its content, and the eleven bugs it found.
On this page
For a long time, we tested Tally with tiny, tidy data: ten users, a hundred events, everything in order. Real customers send millions of events, in the wrong order, with strange names and missing fields. Our tests passed; our customers found the bugs. This is how we fixed that without ever copying real customer data.

The problem with tidy data
Small, clean test data hides three kinds of problems. Size problems: queries that are fast with a thousand rows and slow with a hundred million. Shape problems: events that arrive late, twice or with a timestamp in the future. Odd problems: names with emojis, properties with ten thousand different values, users who visit once a year.
The obvious answer is to test with a copy of production. We decided early never to do that. Our customers' data is theirs, and a copy in a test system is a copy that can leak.
A generator instead
We built a data generator that produces realistic, fake data. It is based on statistics we collect about the shape of real data, never the content: how many events per user, how often events arrive late, how long names are, how many different values a property has.
gen = Generator(seed=42)
gen.users(250_000, active_share=0.18)
gen.events(per_user='lognormal', late_share=0.03, duplicate_share=0.004)
gen.properties(cardinality={'plan': 4, 'page': 12_000, 'campaign': 900})
gen.write('test_dataset')The seed makes every run the same, so a failing test fails every time. The numbers come from our monthly shape report.
We do not need to know what our customers send. We need to know how it is shaped.
Luis Ortega

Three sizes
We keep three datasets. A small one runs with every code change and finishes in a minute. A medium one runs every night. A large one, with four hundred million events, runs before every release and checks that no query got slower.
| Dataset | Events | When it runs | Time |
|---|---|---|---|
| Small | 50,000 | Every change | 1 minute |
| Medium | 20 million | Every night | 25 minutes |
| Large | 400 million | Before each release | 3 hours |
What it caught
In the first month, the generator found eleven bugs that customers had not reported yet, including a time-zone error that only appeared with events sent just after midnight on the last day of a month. It also found a query that took forty seconds on the large dataset; it now takes two.