Considering a bank has 10,000 ATMs, each ATM handling approximately 500 transactions per day, and assuming peak traffic during lunch and evening hours where 10% of the daily transactions occur within a single hour, the estimated peak throughput is:
QPS = (5,000,000 × 10%) / 3600 ≈ 139 transactions per second
The transaction throughput is relatively modest for modern backend systems, so transaction throughput is not expected to be the primary scaling bottleneck. The design should instead prioritize strong consistency, reliability, and fault tolerance over maximizing throughput. yet a latency budget <50 ms per transaction is required, caching relevant authentication information and user data and balances locally for the period of the session is key to make sure no heavy and unecessary server recalls are made especially for read transactions.
Storage Estimation
Considering a transaction record would include:
- Transaction ID (8 Bytes)
- Source Account ID (8 Bytes)
- Destination Account ID (8 Bytes)
- Amount (8 Bytes)
- Timestamp (8 Bytes)
- ATM ID (8 Bytes)
- Status (8 Bytes)
- Bill Scan Metadata (Up to 5 KB for cash deposits)
The storage requirements would be split between the transactional database and object storage, each having different requirements.
Transactional Database
The database would store all transactional records along with any structured bill scan metadata generated during cash deposits.
Assuming a maximum of 5 KB of bill scan metadata per cash deposit, and that 10% of all transactions are cash deposits, the estimated storage requirement is:
- Core transaction data: 5,000,000 × 56 Bytes ≈ 267 MB/day
- Bill scan metadata: 500,000 × 5 KB ≈ 2.5 GB/day
This results in approximately 2.8 GB/day, or roughly 5 TB of raw database storage over a 5-year retention period.
In practice, the actual storage requirements would be significantly higher once database indexes, replication, backups, transaction logs, and additional metadata are taken into account. Depending on the replication strategy and backup retention policy, the total database storage footprint could realistically grow to 15–25 TB.
Object Storage
Assuming 10% of all transactions are check deposits, and that each deposited check stores an image of approximately 30 KB, the estimated storage requirement is:
- 500,000 × 30 KB ≈ 15 GB/day
This results in approximately 27 TB of raw object storage over a 5-year retention period for check images alone.
Similarly, object storage requirements would increase when considering replication, versioning, and backup policies, bringing the total object storage footprint to approximately 50–80 TB, depending on the redundancy strategy.
This separation keeps the relational database focused on structured transactional data, while object storage is used for large binary assets such as check images.
Banking systems usually communicate using financial messaging protocols such as ISO 8583. For the purpose of this interview, we will abstract the communication using RESTful API format.
Authentication:
Getting List of Accounts:
Get specific account information/balance:
Withdraw cash:
Deposit check:
Deposit check:
Funds Transfer Between Accounts:
Update PIN:
Update ATM Availability/Reconcile Cash:
Note that design of all APIs specially the write to the servers or any critical updates are idempotent, a transaction id is added to all requests, to use it to handle request timeouts, disconnections, and other possible issues.
Client/ATM:
The ATM communicates securely with a load balancer over an encrypted connection, The load balancer distributes requests across the available application servers while maintaining high availability.
Each ATM has hardware bound certificate stored in a Trusted Platform Module TPM, and pre registered with the servers. this enables Mutual TLS verification.
ATMs sit on private and secure banking networks, preventing faking ATMs and connecting to the banking servers.
A local cache stores information such as session id, language and other performance enabling and required data for subsequent requests.
Load Balancer:
The load balancer distributes incoming requests across healthy application server instances. It performs health checks, routes traffic away from failed instances, and provides high availability and fault tolerance. Given the estimated throughput of 139 requests per second, the application tier is unlikely to become a bottleneck, yet horizontal scaling would remain available in the case demand increases. at 139 QPS, there is still no need for a distributed cache layer, as it would add unecessary complexity to the architecture, when modern databases can easily handle thousands of requests, therefor session data can be stored on the database, ensuring strong consistency and a simpler architecture.
App Servers:
The application servers process requests arriving from ATMs, execute business logic, persist transactional data to the database cluster, and return the required responses. Although our estimated peak load is approximately 139 transactions per second, the application tier should support both vertical and horizontal scaling if traffic increases.
Fraud detection layer:
A real-time service checking each withdrawal against risk rules (unusual location, rapid successive withdrawals) before authorizing, with a strict latency budget.
Fraud Detection Data:
A NoSQL database such as Cassandra or Dynamo DB would be used, as they have a massive write throughput and low-latency reads.
Database Cluster: A relational SQL database is the appropriate choice for the ATM system's financial transactions because it provides ACID transactions, strong consistency, referential integrity, and mature replication mechanisms, all of which are essential for financial systems. We ensure strong consistency between the required database replicas. The database cluster uses synchronous replication for committed financial transactions. A transaction is acknowledged only after it has been durably written to the primary and replicated to the required number of replicas, ensuring that no committed transaction is lost if the primary fails.
ACID transactions are used to guarantee Atomicity, Consistency, Isolation, and Durability for all financial operations. Each transaction is executed as a single atomic unit, ensuring that account balance updates and transaction records are either fully committed or rolled back, preventing partial updates and ensuring data integrity.
a Postgtes SQL database cluster is best in this scenario, as most data is transactional and strong consistency is required. using ACID transactions guarantees Atomicity making sure transactions are either fully committed or rolled back, Consistency to ensure data is consistent based on preset database rules, Isolation to ensure each transaction is isolated using update locks and Multi-Version Concurrency Control and Durability making sure transactions are never lost and strong consistency is applied.
this database draft schema is created for sake of simplicity, in a real banking system the design would be eventually more complicated taking into consideration many other criteria for fraud detection, currency exchange, more detailed account data, user data etc.
Using a SQL database for financial transactions is the most appropriate in a banking system as it provides ACID transactions and strong consistency and ensures data integrity. The system initially at 139 QPS does not need sharding, but if there was a need for scaling, a sharding mechanism using customer id and location could be put in place to distribute the data and enhance storage ability and performance. in that case a service discovery layer would also be added using a service like zookeper to keep a shard map and react on failovers. as for the fraud and analytics workload, a NoSQL such as DynamoDB or Cassandra would be used to store commited transactions stream to provide the fraud engine the ability to lookup transactions based on key and time range queries ensuring real-time fraud detection.