Tally 2.4 is out: scheduled reports and a 3× faster API
Engineering

Testing with realistic data, without copying real data

A generator built from the shape of customer data, never its content, and the eleven bugs it found.

Testing with realistic data, without copying real data
On this page

For a long time, we tested Tally with tiny, tidy data: ten users, a hundred events, everything in order. Real customers send millions of events, in the wrong order, with strange names and missing fields. Our tests passed; our customers found the bugs. This is how we fixed that without ever copying real customer data.

Most of our worst bugs only appeared with data that looked like real life.
Most of our worst bugs only appeared with data that looked like real life.

The problem with tidy data

Small, clean test data hides three kinds of problems. Size problems: queries that are fast with a thousand rows and slow with a hundred million. Shape problems: events that arrive late, twice or with a timestamp in the future. Odd problems: names with emojis, properties with ten thousand different values, users who visit once a year.

The obvious answer is to test with a copy of production. We decided early never to do that. Our customers' data is theirs, and a copy in a test system is a copy that can leak.

A generator instead

We built a data generator that produces realistic, fake data. It is based on statistics we collect about the shape of real data, never the content: how many events per user, how often events arrive late, how long names are, how many different values a property has.

gen = Generator(seed=42)
gen.users(250_000, active_share=0.18)
gen.events(per_user='lognormal', late_share=0.03, duplicate_share=0.004)
gen.properties(cardinality={'plan': 4, 'page': 12_000, 'campaign': 900})
gen.write('test_dataset')

The seed makes every run the same, so a failing test fails every time. The numbers come from our monthly shape report.

We do not need to know what our customers send. We need to know how it is shaped.

Luis Ortega
The large test dataset has 400 million events. It is rebuilt from scratch every week.
The large test dataset has 400 million events. It is rebuilt from scratch every week.

Three sizes

We keep three datasets. A small one runs with every code change and finishes in a minute. A medium one runs every night. A large one, with four hundred million events, runs before every release and checks that no query got slower.

DatasetEventsWhen it runsTime
Small50,000Every change1 minute
Medium20 millionEvery night25 minutes
Large400 millionBefore each release3 hours

What it caught

In the first month, the generator found eleven bugs that customers had not reported yet, including a time-zone error that only appeared with events sent just after midnight on the last day of a month. It also found a query that took forty seconds on the large dataset; it now takes two.

💡
If you build a generator, make the oddness adjustable. Turning up the share of late or duplicate events is the fastest way to find the code that cannot handle them.

Do you ever look at real data to debug?

Only with the customer's written permission, for a specific problem, and never outside our production system.

Is the generator open source?

Not yet. We would like to share it once the code is less specific to Tally.

Great! Check your inbox and click the link to confirm.