Wrap Up
In this chapter, we have presented the design of a distributed message queue with some advanced features commonly found in data streaming platforms. If there is extra time at the end of the interview, here are some additional talking points:
-
Protocol: it defines rules, syntax, and APIs on how to exchange information and transfer data between different nodes. In a distributed message queue, the protocol should be able to:
-
Cover all the activities such as production, consumption, heartbeat, etc.
-
Effectively transport data with large volumes.
-
Verify the integrity and correctness of the data.
Some popular protocols include Advanced Message Queuing Protocol (AMQP) 18 and Kafka protocol 19.
-
-
Retry consumption. If some messages cannot be consumed successfully, we need to retry the operation. In order not to block incoming messages, how can we retry the operation after a certain time period? One idea is to send failed messages to a dedicated retry topic, so they can be consumed later.
-
Historical data archive. Assume there is a time-based or capacity-based log retention mechanism. If a consumer needs to replay some historical messages that are already truncated, how can we do it? One possible solution is to use storage systems with large capacities, such as HDFS or object storage, to store historical data.
Congratulations on getting this far! Now give yourself a pat on the back. Good job!
Chapter summary
Reference Materials
- Queue Length Limit: https://www.rabbitmq.com/docs/maxlength
- Apache ZooKeeper - Wikipedia: https://en.wikipedia.org/wiki/Apache_ZooKeeper
- etcd: https://etcd.io/
- Comparison of disk and memory performance: https://deliveryimages.acm.org/10.1145/1570000/1563874/jacobs3.jpg
- Cyclic redundancy check: https://en.wikipedia.org/wiki/Cyclic_redundancy_check
- Push vs. pull: https://kafka.apache.org/documentation/#design_pull
- Kafka 2.0 Documentation: https://kafka.apache.org/20/documentation.html#consumerconfigs
- Kafka No Longer Requires ZooKeeper: https://towardsdatascience.com/kafka-no-longer-requires-zookeeper-ebfbf3862104
- Martin Kleppmann. (2017). ‘Replication’ in Designing Data-Intensive Applications. O'Reilly Media. pp. 151-197
- ISR in Apache Kafka: https://www.cloudkarafka.com/docs/dictionary.html
- Apache Kafka allow consumers fetch from closest replica: https://cwiki.apache.org/confluence/display/KAFKA/KIP-392%3A+Allow+consumers+to+fetch+from+closest+replica
- Hands-free Kafka Replication: https://www.confluent.io/blog/hands-free-kafka-replication-a-lesson-in-operational-simplicity/
- Kafka high watermark: https://rongxinblog.wordpress.com/2016/07/29/kafka-high-watermark/
- Kafka mirroring: https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=27846330
- Message filtering in RocketMQ: https://partners-intl.aliyun.com/help/doc-detail/29543.htm
- Scheduled messages and delayed messages in Apache RocketMQ: https://partners-intl.aliyun.com/help/doc-detail/43349.htm
- Hashed and hierarchical timing wheels: http://www.cs.columbia.edu/~nahum/w6998/papers/sosp87-timing-wheels.pdf
- Advanced Message Queuing Protocol: https://en.wikipedia.org/wiki/Advanced_Message_Queuing_Protocol
- Kafka protocol guide: https://kafka.apache.org/protocol
Finished reading?
Mark it complete to track your progress.