System Design Interview

Understand the Problem

Scope 3 min readLesson 1 of 4

In this chapter, we explore the design of a scalable metrics monitoring and alerting system. A well-designed monitoring and alerting system plays a key role in providing clear visibility into the health of the infrastructure to ensure high availability and reliability.

Figure 1 shows some of the most popular metrics monitoring and alerting services in the marketplace. In this chapter, we design a similar service that can be used internally by a large company.

Figure 1 Popular metrics monitoring and alerting services

Step 1 – Understand the problem and establish design scope

A metrics monitoring and alerting system can mean many different things to different companies, so it is essential to nail down the exact requirements first with the interviewer. For example, you do not want to design a system that focuses on logs such as web server error or access logs if the interviewer has only infrastructure metrics in mind.

Let’s first fully understand the problem and establish the scope of the design before diving into the details.

Clarifying the scope
Candidate

Who are we building the system for? Are we building an in-house system for a large corporation like Facebook or Google, or are we designing a SaaS service like Datadog 1, Splunk 2, etc?

Interviewer

That’s a great question. We are building it for internal use only.

Candidate

Which metrics do we want to collect?

Interviewer

We want to collect operational system metrics. These can be low-level usage data of the operating system, such as CPU load, memory usage, and disk space consumption. They can also be high-level concepts such as requests per second of a service or the running server count of a web pool. Business metrics are not in the scope of this design.

Candidate

What is the scale of the infrastructure we are monitoring with this system?

Interviewer

100 million daily active users, 1,000 server pools, and 100 machines per pool.

Candidate

How long should we keep the data?

Interviewer

Let’s assume we want 1-year retention.

Candidate

May we reduce the resolution of the metrics data for long-term storage?

Interviewer

That’s a great question. We would like to be able to keep newly received data for 7 days. After 7 days, you may roll them up to a 1-minute resolution for 30 days. After 30 days, you may further roll them up at a 1-hour resolution.

Candidate

What are the supported alert channels?

Interviewer

Email, phone, PagerDuty, or webhooks (HTTP endpoints).

Candidate

Do we need to collect logs, such as error log or access log?

Interviewer

No.

Candidate

Do we need to support distributed system tracing?

Interviewer

No.

High-level requirements and assumptions

Now you have finished gathering requirements from the interviewer and have a clear scope of the design. The requirements are:

  • The infrastructure being monitored is large-scale.

    • 100 million daily active users

    • Assume we have 1,000 server pools, 100 machines per pool, 100 metrics per machine => ~10 million metrics

    • 1-year data retention

    • Data retention policy: raw form for 7 days, 1-minute resolution for 30 days, 1-hour resolution for 1 year.

  • A variety of metrics can be monitored, for example:

    • CPU usage

    • Request count

    • Memory usage

    • Message count in message queues

Non-functional requirements

  • Scalability. The system should be scalable to accommodate growing metrics and alert volume.

  • Low latency. The system needs to have low query latency for dashboards and alerts.

  • Reliability. The system should be highly reliable to avoid missing critical alerts.

  • Flexibility. Technology keeps changing, so the pipeline should be flexible enough to easily integrate new technologies in the future.

Which requirements are out of scope?

  • Log monitoring. The Elasticsearch, Logstash, Kibana (ELK) stack is very popular for collecting and monitoring logs 3.

  • Distributed system tracing 4 5. Distributed tracing refers to a tracing solution that tracks service requests as they flow through distributed systems. It collects data as requests go from one service to another.

Finished reading?

Mark it complete to track your progress.