Wrap Up
In this chapter, we went through the process of designing an ad click event aggregation system at the scale of Facebook or Google. We covered:
-
Data model and API design.
-
Use MapReduce paradigm to aggregate ad click events.
-
Scale the message queue, aggregation service, and database.
-
Mitigate hotspot issue.
-
Monitor the system continuously.
-
Use reconciliation to ensure correctness.
-
Fault tolerance.
The ad click event aggregation system is a typical big data processing system. It will be easier to understand and design if you have prior knowledge or experience with industry-standard solutions such as Apache Kafka, Apache Flink, or Apache Spark.
Congratulations on getting this far! Now give yourself a pat on the back. Good job!
Chapter Summary
Reference Materials
- Clickthrough rate (CTR): Definition: https://support.google.com/google-ads/answer/2615875?hl=en
- Conversion rate: Definition: https://support.google.com/google-ads/answer/2684489?hl=en
- OLAP functions: https://docs.oracle.com/database/121/OLAXS/olap_functions.htm#OLAXS169
- Display Advertising with Real-Time Bidding (RTB) and Behavioural Targeting: https://arxiv.org/pdf/1610.03013.pdf
- LanguageManual ORC: https://cwiki.apache.org/confluence/display/hive/languagemanual+orc
- Parquet: https://databricks.com/glossary/what-is-parquet
- What is avro: https://www.ibm.com/topics/avro
- Big Data: https://www.datakwery.com/techniques/big-data/
- DAG model https://en.wikipedia.org/wiki/Directed_acyclic_graph
- Java stream: https://docs.oracle.com/javase/8/docs/api/java/util/stream/Stream.html
- Understand star schema and the importance for Power BI: https://docs.microsoft.com/en-us/power-bi/guidance/star-schema
- Martin Kleppmann, “Designing Data-Intensive Applications”, 2017
- Apache Flink: https://flink.apache.org/
- Lambda architecture: https://databricks.com/glossary/lambda-architecture
- Kappa architecture: https://hazelcast.com/glossary/kappa-architecture/
- Martin Kleppmann, “Stream Processing, Designing Data-Intensive Applications”, 2017
- End-to-end Exactly-once Aggregation Over Ad Streams: https://www.youtube.com/watch?v=hzxytnPcAUM
- Ad traffic quality: https://www.google.com/ads/adtrafficquality
- An Overview of End-to-End Exactly-Once Processing in Apache Flink: https://flink.apache.org/features/2018/03/01/end-to-end-exactly-once-apache-flink.html
- Understanding MapReduce in Hadoop: https://www.section.io/engineering-education/understanding-map-reduce-in-hadoop/
- Flink on Apache Yarn https://ci.apache.org/projects/flink/flink-docs-release-1.13/docs/deployment/resource-providers/yarn/
- How data is distributed across a cluster (using virtual nodes): https://docs.datastax.com/en/cassandra-oss/3.0/cassandra/architecture/archDataDistributeDistribute.html
- Flink performance tuning: https://nightlies.apache.org/flink/flink-docs-master/docs/dev/table/tuning/
- ClickHouse: https://clickhouse.com/
- Druid: https://druid.apache.org/
- Real-Time Exactly-Once Ad Event Processing with Apache Flink, Kafka, and Pinot: https://eng.uber.com/real-time-exactly-once-ad-event-processing/
Finished reading?
Mark it complete to track your progress.