System Design Interview

Wrap Up

Wrap up 1 min readLesson 4 of 4

In this chapter, we presented the design for a metrics monitoring and alerting system. At a high level, we talked about data collection, time-series database, alerts, and visualization. Then we went in-depth into some of the most important techniques/components:

  • Pull vs push model for collecting metrics data.

  • Utilize Kafka to scale the system.

  • Choose the right time-series database.

  • Use downsampling to reduce data size.

  • Build vs buy options for alerting and visualization systems.

We went through a few iterations to refine the design, and our final design looks like this:

Figure 22 Final design

Congratulations on getting this far! Now give yourself a pat on the back. Good job!

Chapter Summary

Reference Materials

  1. Datadog: https://www.datadoghq.com/
  2. Splunk: https://www.splunk.com/
  3. Elastic stack: https://www.elastic.co/elastic-stack
  4. Dapper, a Large-Scale Distributed Systems Tracing Infrastructure: https://research.google/pubs/pub36356/
  5. Distributed Systems Tracing with Zipkin: https://blog.twitter.com/engineering/en_us/a/2012/distributed-systems-tracing-with-zipkin.html
  6. Prometheus: https://prometheus.io/docs/introduction/overview/
  7. OpenTSDB - A Distributed, Scalable Monitoring System: http://opentsdb.net/
  8. Data model: : https://prometheus.io/docs/concepts/data_model/
  9. Schema design for time-series data | Cloud Bigtable Documentation https://cloud.google.com/bigtable/docs/schema-design-time-series
  10. MetricsDB: TimeSeries Database for storing metrics at Twitter: https://blog.twitter.com/engineering/en_us/topics/infrastructure/2019/metricsdb.html
  11. Amazon Timestream: https://aws.amazon.com/timestream/
  12. DB-Engines Ranking of time-series DBMS: https://db-engines.com/en/ranking/time+series+dbms
  13. InfluxDB: https://www.influxdata.com/
  14. etcd: https://etcd.io
  15. Service Discovery with Zookeeper https://cloud.spring.io/spring-cloud-zookeeper/1.2.x/multi/multi_spring-cloud-zookeeper-discovery.html
  16. Amazon CloudWatch: https://aws.amazon.com/cloudwatch/
  17. Graphite: https://graphiteapp.org/
  18. Push vs Pull: http://bit.ly/3aJEPxE
  19. Pull doesn’t scale - or does it?: https://prometheus.io/blog/2016/07/23/pull-does-not-scale-or-does-it/
  20. Monitoring Architecture: https://developer.lightbend.com/guides/monitoring-at-scale/monitoring-architecture/architecture.html
  21. Push vs Pull in Monitoring Systems: https://giedrius.blog/2019/05/11/push-vs-pull-in-monitoring-systems/
  22. Pushgateway: https://github.com/prometheus/pushgateway
  23. Building Applications with Serverless Architectures https://aws.amazon.com/lambda/serverless-architectures-learn-more/
  24. Gorilla: A Fast, Scalable, In-Memory Time Series Database: http://www.vldb.org/pvldb/vol8/p1816-teller.pdf
  25. Why We’re Building Flux, a New Data Scripting and Query Language: https://www.influxdata.com/blog/why-were-building-flux-a-new-data-scripting-and-query-language/
  26. InfluxDB storage engine: https://docs.influxdata.com/influxdb/v2.0/reference/internals/storage-engine/
  27. YAML: https://en.wikipedia.org/wiki/YAML
  28. Grafana Demo: https://play.grafana.org/