Understand the Problem
Payment platforms usually provide a digital wallet service to clients, so they can store money in the wallet and spend it later. For example, you can add money to your digital wallet from your bank card and when you buy products online, you are given the option to pay using the money in your wallet. Figure 1 shows this process.
Spending money is not the only feature that the digital wallet provides. For a payment platform like PayPal, we can directly transfer money to somebody else’s wallet on the same payment platform. Compared with the bank-to-bank transfer, direct transfer between digital wallets is faster, and most importantly, it usually does not charge an extra fee. Figure 2 shows a cross-wallet balance transfer operation.
Suppose we are asked to design the backend of a digital wallet application that supports the cross-wallet balance transfer operation. At the beginning of the interview, we will ask clarification questions to nail down the requirements.
Step 1 – Understand the problem and establish design scope
Should we only focus on balance transfer operations between two digital wallets? Do we need to worry about other features?
Let’s focus on balance transfer operations only.
How many transactions per second (TPS) does the system need to support?
Let’s assume 1,000,000 TPS.
A digital wallet has strict requirements for correctness. Can we assume transactional guarantees 1 are sufficient?
That sounds good.
Do we need to prove correctness?
This is a good question. Correctness is usually only verifiable after a transaction is complete. One way to verify is to compare our internal records with statements from banks. The limitation of reconciliation is that it only shows discrepancies and cannot tell how a difference was generated. Therefore, we would like to design a system with reproducibility, meaning we could always reconstruct historical balance by replaying the data from the very beginning.
Can we assume the availability requirement is 99.99%
Sounds good.
Do we need to take foreign exchange into consideration?
No, it’s out of scope.
In summary, our digital wallet needs to support the following:
-
Support balance transfer operation between two digital wallets.
-
Support 1,000,000 TPS.
-
Reliability is at least 99.99%.
-
Support transactions.
-
Support reproducibility.
Back-of-the-envelope estimation
When we talk about TPS, we imply a transactional database will be used. Today, a relational database running on a typical data center node can support a few thousand transactions per second. For example, reference 2 contains the performance benchmark of some of the popular transactional database servers. Let’s assume a database node can support 1,000 TPS. In order to reach 1 million TPS, we need 1,000 database nodes.
However, this calculation is slightly inaccurate. Each transfer command requires two operations: deducting money from one account and depositing money to the other account. To support 1 million transfers per second, the system actually needs to handle up to 2 million TPS, which means we need 2,000 nodes.
Table 1 shows the total number of nodes required when the “per-node TPS” (the TPS a single node can handle) changes. Assuming hardware remains the same, the more transactions a single node can handle per second, the lower the total number of nodes required, indicating lower hardware cost. So one of our design goals is to increase the number of transactions a single node can handle.
| Per-node TPS | Node Number |
|---|---|
| 100 | 20,000 |
| 1,000 | 2,000 |
| 10,000 | 200 |
Table 1 Mapping between pre-node TPS and node number
Finished reading?
Mark it complete to track your progress.