System Design Interview

Wrap Up

Wrap up 1 min readLesson 4 of 4

In this chapter, we went through the process of designing an ad click event aggregation system at the scale of Facebook or Google. We covered:

  • Data model and API design.

  • Use MapReduce paradigm to aggregate ad click events.

  • Scale the message queue, aggregation service, and database.

  • Mitigate hotspot issue.

  • Monitor the system continuously.

  • Use reconciliation to ensure correctness.

  • Fault tolerance.

The ad click event aggregation system is a typical big data processing system. It will be easier to understand and design if you have prior knowledge or experience with industry-standard solutions such as Apache Kafka, Apache Flink, or Apache Spark.

Congratulations on getting this far! Now give yourself a pat on the back. Good job!

Chapter Summary

Reference Materials

  1. Clickthrough rate (CTR): Definition: https://support.google.com/google-ads/answer/2615875?hl=en
  2. Conversion rate: Definition: https://support.google.com/google-ads/answer/2684489?hl=en
  3. OLAP functions: https://docs.oracle.com/database/121/OLAXS/olap_functions.htm#OLAXS169
  4. Display Advertising with Real-Time Bidding (RTB) and Behavioural Targeting: https://arxiv.org/pdf/1610.03013.pdf
  5. LanguageManual ORC: https://cwiki.apache.org/confluence/display/hive/languagemanual+orc
  6. Parquet: https://databricks.com/glossary/what-is-parquet
  7. What is avro: https://www.ibm.com/topics/avro
  8. Big Data: https://www.datakwery.com/techniques/big-data/
  9. DAG model https://en.wikipedia.org/wiki/Directed_acyclic_graph
  10. Java stream: https://docs.oracle.com/javase/8/docs/api/java/util/stream/Stream.html
  11. Understand star schema and the importance for Power BI: https://docs.microsoft.com/en-us/power-bi/guidance/star-schema
  12. Martin Kleppmann, “Designing Data-Intensive Applications”, 2017
  13. Apache Flink: https://flink.apache.org/
  14. Lambda architecture: https://databricks.com/glossary/lambda-architecture
  15. Kappa architecture: https://hazelcast.com/glossary/kappa-architecture/
  16. Martin Kleppmann, “Stream Processing, Designing Data-Intensive Applications”, 2017
  17. End-to-end Exactly-once Aggregation Over Ad Streams: https://www.youtube.com/watch?v=hzxytnPcAUM
  18. Ad traffic quality: https://www.google.com/ads/adtrafficquality
  19. An Overview of End-to-End Exactly-Once Processing in Apache Flink: https://flink.apache.org/features/2018/03/01/end-to-end-exactly-once-apache-flink.html
  20. Understanding MapReduce in Hadoop: https://www.section.io/engineering-education/understanding-map-reduce-in-hadoop/
  21. Flink on Apache Yarn https://ci.apache.org/projects/flink/flink-docs-release-1.13/docs/deployment/resource-providers/yarn/
  22. How data is distributed across a cluster (using virtual nodes): https://docs.datastax.com/en/cassandra-oss/3.0/cassandra/architecture/archDataDistributeDistribute.html
  23. Flink performance tuning: https://nightlies.apache.org/flink/flink-docs-master/docs/dev/table/tuning/
  24. ClickHouse: https://clickhouse.com/
  25. Druid: https://druid.apache.org/
  26. Real-Time Exactly-Once Ad Event Processing with Apache Flink, Kafka, and Pinot: https://eng.uber.com/real-time-exactly-once-ad-event-processing/

Finished reading?

Mark it complete to track your progress.