A few weeks ago I took two weeks of bug reports from our support inbox and looked for each one in Sentry.
For about half of them, Sentry had nothing. No error, no warning.
At the same time, our alert channel was so loud that people stopped reading it.
So we had two problems:
- The bugs that hurt users were invisible.
- The events that did not matter were loud.
Both lead to the same place. Users find the bugs first. They tell support. Support tells us. Days pass.
This post is about how I fixed that, and what it changed. Not only for developers, but for the product.
Why this is a product problem, not only a tech problem
It is easy to see error monitoring as a developer tool. I used to see it that way too.
But think about what a user feels when a file upload spins forever. They do not see an error. They see a product that does not work. Maybe they try again. Maybe they give up. If they are on a free trial, maybe they never come back.
And they rarely tell you. Most people do not write to support. They just leave.
So the real question was never "how many errors do we have?" It was:
When a user gets stuck, do we know? And do we know before they tell us?
Everything I did came from that one question.
What I did, in plain words
I will keep this short. There were four ideas.
1. Only wake people up for real problems. We split alerts into two channels. One is urgent: "this is happening again and again, right now". The other is for normal daily checks. A single failure for a single user no longer pages the team.
2. Give every problem a clear name. Our tool grouped very different failures into one big bucket. A failed login and a failed file save looked like the same bug. I changed the grouping, so each problem got its own clear name. Now you can read the list and understand it.
3. Count people, not events. A lot of errors come from a user's bad Wi-Fi. One person with a weak connection can create 100 errors in an hour. That is not an outage. So for connection problems, we now alert on how many different people are affected. Five people in five places do not lose their Wi-Fi at the same time. If they all fail together, something on our side is really broken.
4. Make silent problems loud. The worst bugs do not crash. The app just waits. An upload stays on "processing" forever. A download button does nothing. So I added small timers: "if this is not done after a few minutes, tell us". And when a request fails, we now keep the server's reason, not only "something went wrong".
That is it. No big new system. Just asking "does a user get stuck here?" at every step.
The results
I set up a small weekly report. Every Friday morning, it compares the last seven days with the seven days before.
The week the changes went live, the numbers went up. Errors rose about 17%, and new issues more than doubled.
That looked bad at first. But it was expected. The new signals were catching problems that were always there. We just could not see them before. For a monitoring project, the first week often looks worse. That is a good sign.
One week later, after we fixed what the new signals found, the numbers came back down.
The alert channel
This is the result my team felt most.
- Before: about 15 urgent alerts per week. Most were one user, one time.
- After: one urgent alert in a full week. And it was real.
When the urgent channel is quiet, people read it again. That is the whole point.
Faster answers
The day after the "stuck upload" timer went live, it fired. A file had been waiting in one processing step for 11 minutes.
Before, this would have been a support ticket, days later, with a vague "it doesn't work". This time we had the file, the step, and the time, within minutes. And we could see the problem was not in the browser. So we knew which team to ask.
What changed for the product side
This part surprised me the most.
Once the noise was gone, the data became useful for more than fixing bugs. People outside engineering started asking questions, and we could answer them honestly.
Our product team asked: "How many new users run into errors in their first days?"
The raw answer was scary: about one in three new users had at least one error recorded.
But when I split the errors into "the user saw something break" and "background noise the app recovered from on its own", the real answer was about one in eight. Most of the rest were short connection drops that nobody noticed.
That difference matters a lot. The first number could start a panic and a rushed rewrite. The second number started a calm talk about a short list of real problems, which ones to fix first, and why.
Without the cleanup, we could not have given either answer with confidence.
What I learned
A few lessons, if you want to do the same:
- Start from the user, not the error. Ask "does someone get stuck?" before "is this an error?"
- A quiet alert channel is a feature. If everything is urgent, nothing is.
- Count people for connection problems. Events tell you how noisy one person is. People tell you if something is really broken.
- The first week may look worse. You are seeing what was always there.
- Check production before you say "fixed". I once called a fix done too early. The next day's data showed it was only half fixed. Now I wait and look.
- Clean data helps the whole team. Good monitoring does not only help developers. It helps product people make calm decisions.
Where we are now
Some things are still open. Some problems cannot be seen from the browser at all, like slow or wrong answers from the server. Those need other tools, and that is our next step.
But the main goal is reached. A few weeks ago, our users told us about bugs before our tools did. Now it mostly works the other way around.
If you work on a frontend and your error tracker feels like noise, try this: take ten recent support tickets and search for each one in your tool. What you find will tell you where to start.