On-call
The first dashboard opened by whoever picks up an alert.
Observability dashboards for every audience: on-call, development teams and management.
Dashboards with dozens of charts nobody looks at, and yet during the incident the one that would have helped is missing. The information exists but isn't organised for whoever has to decide.
Five phases, always in the same order. Select each one to see what happens in it. In full projects they map onto the stages of our method.
We review what's monitored today, with which tools, which incidents went unnoticed and what it costs.
We design the collection layer with OpenTelemetry: agents, gateways, egress paths, common attributes and sampling.
We build dashboard groups by service and audience, with links to traces and logs.
We review each dashboard with its users and remove charts that do not help decision-making.
We tune cardinality, sampling and detectors based on real usage to contain cost and noise.
The first dashboard opened by whoever picks up an alert.
Each team with the view of its own services.
Service health in understandable indicators.
Yes, if they still answer a useful question; the rest is rebuilt or retired.
Yes, with Terraform, just like detectors.
Yes, so you can jump from a chart to the traces and logs from the same moment.
Tell us about your situation. If this service is not what you need, we will tell you; if it is, we will propose a concrete first step.
Request this service