Understand the Problem
In this chapter, we explore the design of a scalable metrics monitoring and alerting system. A well-designed monitoring and alerting system plays a key role in providing clear visibility into the health of the infrastructure to ensure high availability and reliability.
Figure 1 shows some of the most popular metrics monitoring and alerting services in the marketplace. In this chapter, we design a similar service that can be used internally by a large company.
Step 1 – Understand the problem and establish design scope
A metrics monitoring and alerting system can mean many different things to different companies, so it is essential to nail down the exact requirements first with the interviewer. For example, you do not want to design a system that focuses on logs such as web server error or access logs if the interviewer has only infrastructure metrics in mind.
Let’s first fully understand the problem and establish the scope of the design before diving into the details.
That’s a great question. We are building it for internal use only.
Which metrics do we want to collect?
We want to collect operational system metrics. These can be low-level usage data of the operating system, such as CPU load, memory usage, and disk space consumption. They can also be high-level concepts such as requests per second of a service or the running server count of a web pool. Business metrics are not in the scope of this design.
What is the scale of the infrastructure we are monitoring with this system?
100 million daily active users, 1,000 server pools, and 100 machines per pool.
How long should we keep the data?
Let’s assume we want 1-year retention.
May we reduce the resolution of the metrics data for long-term storage?
That’s a great question. We would like to be able to keep newly received data for 7 days. After 7 days, you may roll them up to a 1-minute resolution for 30 days. After 30 days, you may further roll them up at a 1-hour resolution.
What are the supported alert channels?
Email, phone, PagerDuty, or webhooks (HTTP endpoints).
Do we need to collect logs, such as error log or access log?
No.
Do we need to support distributed system tracing?
No.
High-level requirements and assumptions
Now you have finished gathering requirements from the interviewer and have a clear scope of the design. The requirements are:
-
The infrastructure being monitored is large-scale.
-
100 million daily active users
-
Assume we have 1,000 server pools, 100 machines per pool, 100 metrics per machine => ~10 million metrics
-
1-year data retention
-
Data retention policy: raw form for 7 days, 1-minute resolution for 30 days, 1-hour resolution for 1 year.
-
-
A variety of metrics can be monitored, for example:
-
CPU usage
-
Request count
-
Memory usage
-
Message count in message queues
-
Non-functional requirements
-
Scalability. The system should be scalable to accommodate growing metrics and alert volume.
-
Low latency. The system needs to have low query latency for dashboards and alerts.
-
Reliability. The system should be highly reliable to avoid missing critical alerts.
-
Flexibility. Technology keeps changing, so the pipeline should be flexible enough to easily integrate new technologies in the future.
Which requirements are out of scope?
-
Log monitoring. The Elasticsearch, Logstash, Kibana (ELK) stack is very popular for collecting and monitoring logs 3.
-
Distributed system tracing 4 5. Distributed tracing refers to a tracing solution that tracks service requests as they flow through distributed systems. It collects data as requests go from one service to another.
Finished reading?
Mark it complete to track your progress.