0%
Cloud Architecture Patterns
Cloud Foundations
Storage and Databases
Application Patterns
Infrastructure as Code
Reliability and Operations
Advanced Patterns
Choosing a Compute Model
Every backend system runs on compute. The question is not which compute model is "best" but which one fits your workload. VMs, containers, and serverless represent three points on a spectrum trading flexibility for operational simplicity. Understanding where each model excels (and fails) is the foundation for every infrastructure decision you will make.
Virtual Machines
A VM is a full operating system running on virtualized hardware. You get a complete Linux (or Windows) instance with its own kernel, file system, networking stack, and process isolation. You control everything: the OS version, installed packages, kernel parameters, mounted volumes, and firewall rules. This is the most flexible compute model and also the most operationally expensive one.
VMs boot in minutes, not seconds. Scaling means launching a new instance, waiting for the OS to initialize, then running your application startup sequence. Auto-scaling groups help, but the feedback loop is slow. A traffic spike hits at 2:03 PM and your new VM is ready at 2:07 PM. For steady-state workloads this is fine. For bursty traffic it is a problem.
The operational burden of VMs is significant. You are responsible for:
- Operating system updates and security patches
- Firewall configuration and network security groups
- Disk capacity monitoring and volume management
- SSH key rotation and access control
- Backup and disaster recovery procedures
- Monitoring agent installation and configuration
- Log rotation and retention policies
Each of these tasks is straightforward individually but collectively they consume meaningful engineering time. A team managing 50 VMs can easily dedicate one full-time engineer to infrastructure maintenance alone.
Where VMs still win: GPU workloads (ML training, video transcoding), legacy applications that require specific OS configurations, workloads needing direct hardware access, and anything requiring custom kernel modules. If your application needs NVIDIA CUDA drivers, a specific kernel version, or access to bare-metal performance, VMs are your only option.
The cost model is straightforward but unforgiving. You pay per hour for the instance, regardless of whether it is processing requests or sitting idle. A t3.large instance costs the same running at 5% CPU utilization as it does at 95%.
VM pricing tiers create different cost optimization strategies:
- On-demand: pay by the hour, no commitment, highest per-hour cost
- Reserved (1-year or 3-year): 30-60% discount for guaranteed uptime commitment
- Spot instances: 60-90% discount, but the provider can terminate with 2 minutes notice
Reserved instances work for steady-state workloads. Spot instances work for batch processing and fault-tolerant workloads that can checkpoint and restart. On-demand works for unpredictable or temporary workloads where you cannot commit to a reservation.
Containers
A container packages your application code, dependencies, and runtime into a single image that runs identically everywhere. Unlike a VM, a container shares the host OS kernel. This makes containers dramatically lighter: a container image is typically 50-500 MB versus 2-20 GB for a VM image. Startup time drops from minutes to seconds.
The portability guarantee is the key advantage. A container that passes tests on a developer's laptop will behave identically in staging and production because the image includes everything except the kernel. "Works on my machine" disappears as a category of bug. Docker standardized the image format. Kubernetes standardized orchestration. Together, they made containers the default for microservice architectures.
Container scaling is where the model truly shines. Kubernetes Horizontal Pod Autoscaler (HPA) watches CPU utilization, memory usage, or custom metrics and adds or removes pod replicas in seconds. A traffic spike at 2:03 PM triggers a scale-up, and the new pod is serving traffic by 2:03:15 PM. This is orders of magnitude faster than VM auto-scaling. Combined with rolling deployments (new pods start before old pods stop), containers enable zero-downtime deployments multiple times per day, something that is operationally painful with VMs.
The trade-off is operational complexity. Running containers means running Kubernetes (or ECS, or Nomad), which means managing clusters, configuring networking (service meshes, ingress controllers), handling storage (persistent volumes), setting up monitoring, and debugging pod scheduling. Kubernetes is powerful but not simple. A team that adopts containers without investing in Kubernetes expertise will spend more time fighting infrastructure than building product.
Container networking adds another layer of complexity. Each pod gets its own IP address. Services discover each other through DNS-based service discovery. Ingress controllers route external traffic. Network policies restrict internal communication. Service meshes (Istio, Linkerd) add mTLS, traffic shaping, and circuit breaking. Each layer solves a real problem but adds configuration surface area. A misconfigured network policy can silently drop traffic between services, creating intermittent failures that are notoriously difficult to debug.
Storage in containers follows the same ephemeral principle. By default, container storage disappears when the pod restarts. For workloads that need persistent data (databases, file uploads, ML model checkpoints), Kubernetes provides PersistentVolumes backed by cloud block storage (EBS, GCE Persistent Disk). These volumes survive pod restarts but introduce new failure modes: a volume can only attach to one node at a time, detaching from a crashed node takes minutes, and storage IOPS limits can throttle database performance in ways that are invisible at the application level.
In interviews, when asked to justify your compute choice, lead with the workload characteristics rather than the technology name. Saying 'I would use containers because this is a stateless microservice that needs to scale horizontally and deploy frequently' is stronger than saying 'I would use Kubernetes because it is industry standard.' The reasoning matters more than the conclusion.
Serverless
Serverless (AWS Lambda, Google Cloud Functions, Azure Functions) removes the server entirely from your mental model. You deploy a function. The cloud provider allocates compute when a request arrives, executes your function, and deallocates when it finishes. You pay per invocation and per millisecond of execution, not per hour of uptime. When traffic is zero, your cost is zero.
Startup time is measured in milliseconds for warm invocations. Cold starts (when the provider must initialize a new execution environment) add 100ms to several seconds depending on the runtime and package size. This is the fundamental constraint of serverless: you cannot control when cold starts happen, and for latency-sensitive endpoints, that unpredictability may be unacceptable.
Serverless excels at event-driven, bursty workloads: processing S3 upload events, handling webhook callbacks, running scheduled jobs, transforming data in a pipeline. These workloads are inherently intermittent. Paying for an always-on container or VM to handle a webhook that fires 50 times per day wastes 99.9% of your compute budget.
The concurrency model of serverless is fundamentally different from containers and VMs. A Lambda function handles one request per execution environment. To handle 100 concurrent requests, the provider spins up 100 separate environments. Each has its own memory, its own file system, and its own network connections. There is no shared memory between invocations.
This means traditional patterns do not work in serverless:
- In-memory caches are lost between invocations
- Connection pools are per-environment, not shared
- Background threads are killed when the handler returns
- Local file writes disappear when the environment is recycled
- Singleton objects are per-environment, not per-application
You must rethink your application to be stateless per-invocation: read state from an external store at the beginning, process the request, write state back at the end, and assume the execution environment may not exist for the next invocation.
Where serverless fails: long-running processes (Lambda has a 15-minute timeout), workloads requiring persistent connections (WebSockets, gRPC streams), GPU compute, and applications needing more than 10 GB of memory. The execution environment is a black box. You cannot tune the OS, install system libraries, or mount custom volumes. The provider controls everything below your function code.
Cold Starts: The Hidden Cost
Cold starts deserve special attention because they are the single most misunderstood aspect of serverless. When a Lambda function has not been invoked recently, the provider must allocate an execution environment, download the deployment package, initialize the runtime, and run your initialization code before handling the first request. This process adds latency that ranges from 100ms for a small Python function to 10+ seconds for a large Java application with many dependencies.
The unpredictability is the real problem. You cannot control when cold starts happen. A function that handled 100 requests per second with warm instances will suddenly cold-start if traffic pauses for 15 minutes. Provisioned concurrency (keeping a fixed number of warm instances) mitigates this but adds cost, partially defeating the pay-per-use advantage.
For latency-sensitive endpoints (user-facing APIs with SLA requirements), cold starts may be unacceptable. For background processing (file conversions, data pipeline steps), an extra second of latency is invisible to users and perfectly acceptable.
Resource Limits and Isolation
Understanding how each model isolates resources affects both performance and security decisions.
VMs provide the strongest isolation. Each VM has its own kernel, its own memory space, and its own network stack. A security vulnerability in one VM cannot access another VM's memory or processes. This is why multi-tenant cloud providers use VMs as the trust boundary: your EC2 instance is isolated from other customers at the hypervisor level.
Containers provide process-level isolation through Linux namespaces and cgroups. Containers share the host kernel, which means a kernel vulnerability affects all containers on the host. Within a container, you can set CPU limits (e.g., 0.5 CPU cores) and memory limits (e.g., 512 MB). If a container exceeds its memory limit, Kubernetes kills it (OOMKilled). This prevents one container from consuming all host resources, but the isolation is weaker than VM-level separation.
Serverless provides the strongest resource isolation from the developer's perspective: you specify memory (128 MB to 10 GB), and the provider handles everything else. CPU scales proportionally with memory. You cannot exceed your allocation because the provider enforces it at the infrastructure level. The trade-off is zero visibility into what happens below your function: you cannot see CPU utilization, network bandwidth, or disk I/O. When performance degrades, you have limited debugging tools compared to a container where you can exec into the pod and run diagnostic commands.
The Comparison Matrix
Think of these three models along six dimensions:
Startup time: VMs take minutes. Containers take seconds. Serverless takes milliseconds (warm) to seconds (cold).
Scaling speed: VMs scale in minutes. Containers scale in seconds via Kubernetes HPA. Serverless scales per-request with no pre-provisioning.
Cost model: VMs charge per hour regardless of utilization. Containers charge per cluster node regardless of pod utilization. Serverless charges per invocation and execution duration.
Operational overhead: VMs require OS patching, security updates, capacity planning, and monitoring. Containers require orchestrator management, image building, and networking configuration. Serverless requires only function code and event source configuration.
Flexibility: VMs offer full OS control. Containers offer application-level control. Serverless constrains runtime, memory, timeout, and available libraries.
Vendor lock-in: VMs are portable across any cloud. Containers are portable via the OCI image standard and Kubernetes. Serverless functions are tightly coupled to provider-specific event sources, IAM models, and APIs.
When to Use Each Model: A Summary
Use VMs when you need: full OS control, GPU or custom hardware, kernel-level configuration, long-running batch processing, or legacy applications that cannot be containerized without a rewrite.
Use containers when you need: fast horizontal scaling, portable deployments across environments, frequent releases (multiple per day), microservice architectures with many small services, or workloads that are stateless and short-lived per request.
Use serverless when you need: zero infrastructure management, per-invocation billing for cost efficiency on low-traffic or bursty workloads, event-driven processing triggered by cloud services, or rapid prototyping where time-to-production matters more than long-term cost optimization.
Mid-level engineers are expected to explain the trade-offs between VMs, containers, and serverless for a given workload. Senior engineers should quantify those trade-offs (startup latency, cost per request at a given scale, operational headcount). Staff engineers evaluate organizational readiness: does the team have Kubernetes expertise, is there a platform team to manage the orchestrator, and what is the migration cost from the current model?