Cost Optimization

Topics Covered

Understanding Cloud Billing

The Four Billing Dimensions

Multi-Dimensional Billing in Practice

On-Demand, Reserved, and Spot: The Pricing Spectrum

Why Bills Grow Faster Than Traffic

The Shared Responsibility Model for Cost

Hidden Costs That Compound

Reserved and Savings Plans

Savings Plans

Reserved Instances

The Break-Even Decision

Common Commitment Mistakes

Measuring Commitment Effectiveness

Spot and Preemptible Workloads

How Spot Pricing Works

Ideal Spot Workloads

Spot Fleet Strategies

Handling the 2-Minute Warning

Mixing Spot with On-Demand

Right-Sizing and Waste Elimination

Utilization Analysis

Eliminating Idle Resources

Auto-Scaling to Zero

Storage Optimization

The Right-Sizing Process

FinOps Practices

The FinOps Cycle

FinOps Maturity Levels

The FinOps Team Structure

Cost Allocation and Tagging

Data Transfer Cost Architecture

Building a Cost-Aware Culture

Anomaly Detection and Budgets

Governance and Guardrails

Cloud billing catches teams off guard because it works nothing like traditional infrastructure purchasing. With on-premises servers, you pay a large upfront cost and then amortize it over years. With cloud, you pay for what you consume, measured in dimensions most engineers never think about: compute per hour, storage per GB-month, data transfer per GB, and API requests per million.

The same traffic billed before and after optimization, split into compute, storage, transfer and requests.

The Four Billing Dimensions

Every cloud resource maps to one or more of these billing axes:

Compute (per hour or per second): Virtual machines, containers, and serverless functions all charge based on time running. An m5.xlarge instance on AWS costs approximately $0.192/hour. Whether it is serving traffic at 100% CPU or sitting idle at 2% CPU, the charge is the same. This is the single biggest reason cloud bills surprise teams: you pay for provisioned capacity, not utilized capacity.

Storage (per GB-month): EBS volumes, S3 buckets, and database storage all charge monthly per gigabyte stored. The cost seems small ($0.023/GB-month for S3 Standard), but it compounds. A team that creates daily database snapshots and never deletes old ones accumulates terabytes within a year.

Data Transfer (per GB): This is the billing dimension that blindsides most architects. Traffic within a single Availability Zone is free. Cross-AZ traffic costs $0.01/GB. Cross-region costs $0.02/GB. Internet egress costs $0.09/GB. A microservice architecture where services in different AZs exchange 1 TB of data per month pays $10 just for internal communication.

API Requests (per million): Services like S3, SQS, Lambda, and API Gateway charge per request. S3 GET requests cost $0.0004 per 1,000. At 100 million requests per month, that is $40 just for the request overhead, separate from storage and transfer costs.

Multi-Dimensional Billing in Practice

A single user request often touches all four billing dimensions simultaneously. Consider a web application that receives an image upload: the request hits an API Gateway (request charge), triggers a Lambda function (compute charge per 100ms), stores the image in S3 (storage charge), and if the Lambda runs in a different AZ than the S3 bucket, incurs data transfer charges. A seemingly simple feature generates costs across four independent billing axes, each with its own pricing model and scaling behavior.

This is why cloud cost estimation is harder than it appears. Engineers who model only compute costs miss 40-60% of the actual bill. The most expensive architectures are not the ones with the biggest instances. They are the ones with the most data movement between services, regions, and the internet.

A useful exercise before designing any new service is to draw the data flow and annotate each hop with its per-GB cost. Intra-AZ is free, cross-AZ is $0.01/GB, cross-region is $0.02/GB, and internet egress is $0.09/GB. Multiply by expected monthly volume and you have a transfer cost estimate before writing a single line of code. This exercise has killed bad architecture decisions (like putting a cache in a different region than its application) before they became expensive mistakes.

On-Demand, Reserved, and Spot: The Pricing Spectrum

Cloud providers offer the same compute capacity at three price points, each with different flexibility and risk characteristics:

On-demand is the default: pay per hour or per second with no commitment. You start instances when you need them and stop them when you do not. The flexibility is maximum but so is the price. On-demand makes sense for unpredictable workloads, short-lived experiments, and the variable portion of traffic that sits above your steady-state baseline.

Reserved commits you to a specific capacity for 1-3 years in exchange for 30-60% discounts. The savings are substantial but the commitment is real. If your needs change (new instance family, different region, workload decommissioned), the reservation becomes wasted spend.

Spot gives you spare capacity at 60-90% discounts, but the provider can reclaim it with 2 minutes notice. Spot is ideal for fault-tolerant, interruptible workloads: batch processing, CI/CD, training ML models, and stateless web workers that can lose individual instances without user impact.

The cost-optimal architecture layers all three: reserved or savings plans for the predictable baseline, on-demand for variable traffic above baseline, and spot for fault-tolerant workloads that can handle interruption. This layered approach captures 40-50% overall savings compared to running everything on-demand.

Understanding which pricing model fits which workload is one of the highest-value skills in cloud architecture. A common interview question asks candidates to design a cost-efficient infrastructure for a workload with known characteristics. The answer always involves identifying which portion of the workload is predictable (use reserved), which is variable (use on-demand), and which is interruptible (use spot). Getting this decomposition right determines whether the infrastructure costs $10,000/month or $20,000/month for the same performance.

Interview Tip

When designing systems, always check the data transfer pricing before placing services in different availability zones or regions. A decision that improves fault tolerance by spreading across three AZs can quietly add thousands of dollars per month in cross-AZ data transfer charges. Quantify the cost before making the architectural choice.

Why Bills Grow Faster Than Traffic

Cloud bills often grow 2-3x faster than actual traffic because of three compounding effects.

First, engineers provision for peak capacity but traffic is spiky: a server sized for peak handles 10x more than average load, meaning 90% of the time you pay for unused capacity. Without auto-scaling, you are permanently paying for your worst-case scenario.

Second, storage is append-only in practice: logs, snapshots, backups, and old versions accumulate because nobody deletes them. Unlike compute, which resets when you terminate an instance, storage persists indefinitely unless explicitly removed. A team that generates 50 GB of logs per day and never sets a retention policy accumulates 18 TB in a year.

Third, architecture choices multiply costs: adding a CDN, a cache layer, and a monitoring system each add their own billing dimensions on top of the core compute and storage costs. A monitoring agent that samples metrics every 10 seconds and sends them to CloudWatch generates API request charges. A cache cluster in a different AZ generates data transfer charges on every cache hit. Each layer solves a real problem but adds a billing dimension that compounds over time.

The Shared Responsibility Model for Cost

Cloud providers optimize their own infrastructure costs through economies of scale, but cost optimization for your workloads is entirely your responsibility. The provider does not warn you that an instance is over-provisioned, that an EBS volume is unattached, or that your data transfer architecture is expensive. They simply bill you.

This means cost optimization must be an engineering discipline, not a finance afterthought. The engineer who chooses an instance type, designs the data flow between services, or selects a storage class is making cost decisions whether they realize it or not. Every architectural choice has a billing consequence. The difference between a cost-efficient architecture and an expensive one is often not technical complexity, but awareness of the billing dimensions involved.

Hidden Costs That Compound

Several billing dimensions catch even experienced engineers off guard. NAT Gateway data processing charges $0.045/GB on top of standard data transfer costs. A service making frequent calls to AWS APIs (S3, DynamoDB, SQS) through a NAT Gateway pays this processing fee on every byte. Using VPC endpoints eliminates the NAT Gateway data processing charge for supported services.

CloudWatch metrics, logs, and alarms have their own billing dimensions. Custom metrics cost $0.30/month each. Log ingestion costs $0.50/GB. A microservice that publishes 20 custom metrics and generates 100 GB of logs per month adds $210/month in monitoring costs alone, before any compute or storage charges.

Elastic IP addresses are free when attached to a running instance. But a stopped instance with an attached Elastic IP costs $0.005/hour ($3.65/month). Organizations with many stopped dev instances accumulate dozens of Elastic IPs that quietly add up.

Understanding these billing mechanics is the foundation of cost optimization. You cannot reduce costs you do not understand. The first step is always: map every resource to its billing dimension and measure actual utilization against provisioned capacity.