Why we rebuilt our event pipeline in six weeks
We had a six-month plan and no time for it. So we asked a narrower question.
On this page
In June, the part of Tally that turns raw events into charts started to struggle. Queries that took a third of a second were taking two. Our largest customers were waiting for dashboards to load. We had a plan to fix it that would take six months, and we did not have six months.
This is the story of how we rebuilt it in six weeks instead, what we decided not to do, and what we would do differently.

The problem
Every chart in Tally is a question about events: how many people did X, how many then did Y, how many came back a week later. The old engine answered every question by reading the raw events for the whole time range and counting them.
That is simple and correct. It is also slow when a customer sends forty million events a month and asks about the last ninety days. We had been adding bigger machines for a year. In June, the machines stopped being the cheap part.
The plan we did not follow
The obvious plan was to move to a new database built for analytics. We made a proof of concept. It was fast. It would also have meant rewriting every query, migrating two years of data, and running two systems side by side for months.
We estimated six months. Everyone knows what software estimates mean, so call it nine.
The plan we followed
Instead we asked a narrower question: which questions are slow, and what do they have in common?
The answer was clear from our logs. More than 90% of slow queries asked about days or hours, not about individual events. They counted things by day, grouped by one or two properties. They almost never needed the exact second.
So we built summaries. Every hour, a small job counts the events of the last hour by event name and by the properties customers actually use in charts. Every night, a second job rolls the hours into days. When a question comes in, the engine reads the summaries for the past and the raw events only for the last few minutes.
-- One row per hour, event and property value
CREATE TABLE event_hours (
workspace_id BIGINT,
hour TIMESTAMP,
event TEXT,
prop_key TEXT,
prop_value TEXT,
count BIGINT,
users HLL, -- approximate distinct users
PRIMARY KEY (workspace_id, hour, event, prop_key, prop_value)
);The HLL column is a HyperLogLog sketch: a small structure that estimates how many different users did something, and that can be combined across hours without counting everyone again. It is accurate to within about 1%.

Six weeks, roughly
| Week | What we did |
|---|---|
| 1 | Measured every slow query and grouped them |
| 2 | Built the hourly summary job and backfilled one workspace |
| 3 | Ran old and new engines side by side, compared every answer |
| 4 | Fixed the differences (there were eleven) and switched 10% of workspaces |
| 5 | Switched everyone, kept the old engine as a fallback |
| 6 | Removed the fallback and deleted 4,000 lines of code |
The side-by-side week was the most important. For seven days, every query ran on both engines. Any difference bigger than 1% was logged. We found eleven, and every one was a real bug, in one engine or the other. Three were in the old engine and had been there for a year.
The best part of the project was deleting code. The second best was finding bugs in code we thought was finished.
Luis Ortega
What we gave up
Summaries are not free. Some questions still need raw events: "show me every person who did X", or properties we do not summarise. Those queries are exactly as fast as before, which is to say, not very. About 6% of queries fall into this group.
We also accepted approximate distinct counts. "12,408 users" might really be 12,390. For product analytics, we think that is fine. For billing, we still count exactly.
What we would do differently
We would build the comparison tool first. Running both engines side by side found more problems than all our tests together, and we only built it in week three because we were worried. Next time it will be week one.
We would also talk to customers earlier. Two customers used a property in a way we did not expect, and their charts were the last to switch. A short email in week one would have saved us a week of guessing.
How we kept customers in the loop
We told customers about the rebuild before it started. A short post on the status page explained what we were changing, why charts might be slow for a few more weeks, and what we would do if something went wrong. Surprisingly, this reduced support tickets about slow dashboards by almost half. People are patient when they know someone is working on the problem.
During the switch, every chart had a hidden "engine" label that support could see. When a customer reported a strange number, the first question was always which engine had answered it. That one label saved hours of debugging in week four.
The team
Three engineers built the new engine: Luis, plus two colleagues who had joined only four months earlier. We think the small team was an advantage. Every decision could be made at a whiteboard in ten minutes, and nobody had time to build anything that was not strictly needed. If we had started with the six-month plan and a team of eight, we suspect we would still be building it.