Wrap Up
Wrap up 1 min readLesson 4 of 4
In this chapter, we presented the design for a metrics monitoring and alerting system. At a high level, we talked about data collection, time-series database, alerts, and visualization. Then we went in-depth into some of the most important techniques/components:
-
Pull vs push model for collecting metrics data.
-
Utilize Kafka to scale the system.
-
Choose the right time-series database.
-
Use downsampling to reduce data size.
-
Build vs buy options for alerting and visualization systems.
We went through a few iterations to refine the design, and our final design looks like this:
Congratulations on getting this far! Now give yourself a pat on the back. Good job!
Chapter Summary
Reference Materials
- Datadog: https://www.datadoghq.com/
- Splunk: https://www.splunk.com/
- Elastic stack: https://www.elastic.co/elastic-stack
- Dapper, a Large-Scale Distributed Systems Tracing Infrastructure: https://research.google/pubs/pub36356/
- Distributed Systems Tracing with Zipkin: https://blog.twitter.com/engineering/en_us/a/2012/distributed-systems-tracing-with-zipkin.html
- Prometheus: https://prometheus.io/docs/introduction/overview/
- OpenTSDB - A Distributed, Scalable Monitoring System: http://opentsdb.net/
- Data model: : https://prometheus.io/docs/concepts/data_model/
- Schema design for time-series data | Cloud Bigtable Documentation https://cloud.google.com/bigtable/docs/schema-design-time-series
- MetricsDB: TimeSeries Database for storing metrics at Twitter: https://blog.twitter.com/engineering/en_us/topics/infrastructure/2019/metricsdb.html
- Amazon Timestream: https://aws.amazon.com/timestream/
- DB-Engines Ranking of time-series DBMS: https://db-engines.com/en/ranking/time+series+dbms
- InfluxDB: https://www.influxdata.com/
- etcd: https://etcd.io
- Service Discovery with Zookeeper https://cloud.spring.io/spring-cloud-zookeeper/1.2.x/multi/multi_spring-cloud-zookeeper-discovery.html
- Amazon CloudWatch: https://aws.amazon.com/cloudwatch/
- Graphite: https://graphiteapp.org/
- Push vs Pull: http://bit.ly/3aJEPxE
- Pull doesn’t scale - or does it?: https://prometheus.io/blog/2016/07/23/pull-does-not-scale-or-does-it/
- Monitoring Architecture: https://developer.lightbend.com/guides/monitoring-at-scale/monitoring-architecture/architecture.html
- Push vs Pull in Monitoring Systems: https://giedrius.blog/2019/05/11/push-vs-pull-in-monitoring-systems/
- Pushgateway: https://github.com/prometheus/pushgateway
- Building Applications with Serverless Architectures https://aws.amazon.com/lambda/serverless-architectures-learn-more/
- Gorilla: A Fast, Scalable, In-Memory Time Series Database: http://www.vldb.org/pvldb/vol8/p1816-teller.pdf
- Why We’re Building Flux, a New Data Scripting and Query Language: https://www.influxdata.com/blog/why-were-building-flux-a-new-data-scripting-and-query-language/
- InfluxDB storage engine: https://docs.influxdata.com/influxdb/v2.0/reference/internals/storage-engine/
- YAML: https://en.wikipedia.org/wiki/YAML
- Grafana Demo: https://play.grafana.org/
Finished reading?
Mark it complete to track your progress.