0%
Cloud Architecture Patterns
Cloud Foundations
Storage and Databases
Application Patterns
Infrastructure as Code
Reliability and Operations
Advanced Patterns
Virtual Machines and Auto-Scaling
Choosing a VM instance type is not about picking the biggest machine you can afford. It is about matching the resource profile of your workload to the hardware the cloud provider offers. Every dollar spent on over-provisioned CPU is a dollar wasted. Every minute of memory starvation is a minute of degraded user experience. The right instance type makes your application faster and cheaper simultaneously.
Cloud providers offer hundreds of instance types because workloads are not uniform. A video encoding pipeline has fundamentally different resource needs than an in-memory cache, which has different needs than a machine learning training job. Understanding the major instance families and when to use each one is essential knowledge for designing cost-effective infrastructure. The wrong choice does not just waste money. It can actively degrade performance when the resource bottleneck is not addressed by the instance type you selected.
General Purpose Instances
General purpose instances (AWS m-family, GCP n2-standard, Azure Dv5) offer a balanced ratio of CPU to memory, typically 1 vCPU per 4 GB of RAM. These are your default choice when you do not know your workload's resource profile yet. Web servers, small databases, development environments, and API backends all run well on general purpose instances because they need moderate compute and moderate memory without extreme demands on either.
The trap is staying on general purpose forever. Once your application is in production and you have utilization data, continuing to use balanced instances when your workload is clearly CPU-bound or memory-bound means you are paying for resources you never use.
A common misconception is that general purpose means "slow." These instances deliver strong performance for most applications. The point is not that they are slow. It is that their balanced ratio wastes money when your workload heavily favors one resource dimension. A web server that spends 90% of its time waiting on database responses does not need high CPU. A batch processor that saturates all cores does not need 4 GB per vCPU.
Compute-Optimized Instances
Compute-optimized instances (AWS c-family, GCP c2, Azure Fsv2) provide a high CPU-to-memory ratio, typically 1 vCPU per 2 GB of RAM. They exist for workloads that are bottlenecked on CPU, not memory: video encoding, scientific simulations, batch processing, game servers, and machine learning inference. If your application's memory usage sits at 20% while CPU pegs at 95%, you need compute-optimized, not a bigger general purpose instance.
Compute-optimized instances also tend to use the latest generation processors with higher clock speeds and better per-core performance. This matters for single-threaded workloads where clock speed determines throughput more than core count. A game server running a single-threaded physics loop benefits more from a faster core than from more cores.
Memory-Optimized Instances
Memory-optimized instances (AWS r-family, GCP m2, Azure Ev5) flip the ratio: 1 vCPU per 8 GB of RAM or more. In-memory databases like Redis and Memcached, real-time analytics engines like Apache Spark, and large caches all need enormous amounts of RAM relative to CPU. A Redis instance holding 200 GB of session data needs perhaps 4 vCPUs but 256 GB of RAM. Putting this on a general purpose instance wastes 60 vCPUs worth of cost to get the memory you actually need.
High-memory variants (AWS x-family, u-family) push this even further with up to 24 TB of RAM for specialized workloads like SAP HANA or in-memory OLAP databases. These are rare but illustrate the principle: cloud providers offer extreme configurations because some workloads genuinely need them, and the alternative (sharding across many smaller instances) adds complexity that may exceed the cost of a single large machine.
GPU Instances
GPU instances (AWS p-family, GCP a2, Azure NCv3) attach one or more GPUs for parallel computation. Machine learning training, 3D rendering, and video transcoding all benefit from GPU parallelism. A single GPU can process a batch of 1,000 matrix multiplications in the time a CPU handles 10. GPU instances are expensive, often 5-10x the cost of similarly-sized CPU instances, so they should only run workloads that genuinely benefit from GPU parallelism. Running a REST API on a GPU instance is like using a freight train to deliver a letter.
GPU instance selection also depends on the type of GPU workload. Training large models requires GPUs with high memory (A100 with 80 GB, H100 with 80 GB) to hold model parameters and gradients. Inference serving needs less GPU memory but benefits from high throughput (T4, L4 GPUs). Multi-GPU instances (p4d.24xlarge with 8 A100 GPUs) are necessary for distributed training of models too large to fit on a single GPU. Choosing the wrong GPU type is expensive: paying for 80 GB of GPU memory to serve a model that needs 4 GB wastes 95% of the GPU's most expensive resource.
In system design interviews, specify instance types when discussing infrastructure. Saying 'we need compute-optimized instances for the encoding pipeline and memory-optimized for the Redis cache' demonstrates that you understand workload profiling, not just 'we need more servers.'
How to Choose
The decision process is straightforward. Start with general purpose for new workloads. After one week of production traffic, check your utilization ratios. If CPU is consistently above 70% while memory is below 30%, switch to compute-optimized. If memory is above 70% while CPU is below 30%, switch to memory-optimized. If your workload involves matrix operations, training, or rendering, evaluate GPU instances. The cloud providers make it easy to resize: stop the instance, change the type, start it again. The hard part is remembering to check.
Instance Sizing Within a Family
Within each family, sizes follow a consistent pattern. A "medium" has 1 vCPU, a "large" has 2, an "xlarge" has 4, a "2xlarge" has 8, and so on, doubling each step. Memory scales proportionally. Pricing is linear: a 2xlarge costs exactly twice an xlarge. This means vertical scaling (moving to a larger size within the same family) has no economy of scale. Two xlarge instances cost the same as one 2xlarge but give you redundancy. When possible, prefer horizontal scaling (more smaller instances) over vertical scaling (fewer larger instances) because it improves fault tolerance at the same price point.
Storage-optimized instances (AWS i-family, d-family) deserve a brief mention. These provide high-throughput local NVMe SSDs for workloads like HDFS data nodes, distributed databases, and data warehouses that need fast local disk I/O. They are niche but critical when your bottleneck is neither CPU nor memory but disk throughput.
Networking Considerations
Instance types also differ in network bandwidth. A t3.micro might offer baseline bandwidth of 5 Gbps, while a c5.18xlarge offers 25 Gbps. For data-intensive applications (video streaming, large file transfers, inter-service communication at high throughput), network bandwidth can be the limiting factor even when CPU and memory are underutilized. Enhanced Networking (ENA) on AWS provides higher bandwidth, lower latency, and lower jitter compared to standard networking. If your application does significant network I/O, verify that your chosen instance type supports the bandwidth your workload requires.
Placement groups further influence network performance. Cluster placement groups place instances physically close within a single AZ, minimizing inter-instance latency to sub-millisecond levels. This is essential for HPC workloads and tightly-coupled distributed systems like Elasticsearch clusters where nodes exchange data frequently.
Spread placement groups take the opposite approach: they ensure instances are placed on different physical hardware to minimize correlated failures. Use spread placement groups for small numbers of critical instances (database replicas, ZooKeeper quorum nodes) where the failure of one physical host should never take down more than one instance. Partition placement groups combine both ideas: instances within a partition are spread across hardware, but partitions are isolated from each other, giving you blast radius control for distributed databases like Cassandra and HDFS.