Cloud Computing Models

Topics Covered

IaaS, PaaS, and SaaS

Infrastructure as a Service (IaaS)

Platform as a Service (PaaS)

Software as a Service (SaaS)

Function as a Service (FaaS) and Serverless

How to Choose the Right Model

Cost Comparison Across Models

Shared Responsibility Model

The Boundary Shifts at Each Level

The Most Common Mistakes

Configuration Is Your Biggest Attack Surface

Applying the Model in Practice

Compliance Frameworks and the Shared Model

Cloud-Native vs Lift-and-Shift

Lift-and-Shift

Cloud-Native Architecture

The Middle Path: Modernize Incrementally

Multi-Cloud vs Single-Cloud

Containers as the Migration Bridge

Twelve-Factor App Principles

Choosing a Cloud Provider

Pricing Models

Compliance and Certifications

Team Expertise and Ecosystem

Geographic Presence

The Decision Framework

Migration Cost and Lock-in Assessment

Every cloud service falls somewhere on a spectrum of "how much do you manage versus how much does the provider manage." Understanding where each model sits on that spectrum is the single most important decision in cloud architecture because it determines your operational burden, your flexibility, and your cost structure for years to come.

IaaS, PaaS and SaaS as rungs of one stack, with the layer handed over and the operational control given up at every step.

Infrastructure as a Service (IaaS)

IaaS gives you virtual machines, storage volumes, and networking primitives. The provider handles the physical hardware, power, cooling, and hypervisor. You handle everything above that: operating system, runtime, middleware, application code, and data.

Think of IaaS as renting a bare apartment. The landlord maintains the building structure, plumbing, and electrical wiring. You furnish it, paint the walls, and maintain the interior. AWS EC2, Google Compute Engine, and Azure Virtual Machines are the canonical examples.

When you launch an EC2 instance, you choose the OS image, configure security groups, install your runtime, deploy your application, and manage patches. If a critical OpenSSL vulnerability drops, you patch it. If your application needs a specific version of Python, you install it. This control is the point: teams that need custom kernel modules, specific OS configurations, or legacy software dependencies choose IaaS because nothing else gives them that level of control.

The cost of that control is operational burden. You are running an OS in production. That means patch management, capacity planning, monitoring disk space, rotating SSH keys, and debugging network issues at the TCP level. For a team of 3 engineers building a new product, this overhead can consume 30-40% of engineering time.

IaaS also means you own capacity planning. If your application gets featured on a news site and traffic spikes 10x, you need auto-scaling groups configured, launch templates defined, and health checks tuned. None of this comes for free. You build the scaling machinery yourself using the provider's primitives (launch templates, auto-scaling groups, load balancer target groups). The provider gives you the building blocks, but assembly is your job.

Networking in IaaS is another layer of complexity. You design VPCs, configure subnets across availability zones, set up routing tables, manage NAT gateways for private subnet internet access, and define security groups for every service. A misconfigured security group can expose your database to the internet. A missing route table entry can silently break inter-service communication. These networking primitives are powerful but unforgiving. In contrast, PaaS and FaaS abstract networking almost entirely, removing this entire class of operational work and potential misconfiguration.

Monitoring and alerting on IaaS is also your responsibility. You install and configure monitoring agents on each VM, define metric collection for CPU, memory, disk, and network usage, set up dashboards for visibility, and create alerting rules for anomaly detection. If a VM runs out of disk space at 3 AM and nobody is alerted, that is your failure, not the provider's. The provider surfaces basic hypervisor-level metrics (CPU utilization, network bytes), but application-level metrics (request latency, error rates, queue depth) require instrumentation in your application code and a monitoring stack (Prometheus, Grafana, Datadog) that you deploy and maintain.

Platform as a Service (PaaS)

PaaS removes the OS and runtime from your responsibility. You write your application code, push it to the platform, and the platform handles provisioning servers, installing runtimes, scaling instances, applying OS patches, and managing load balancers.

The analogy shifts from a bare apartment to a furnished co-working space. You bring your laptop and work. The building owner handles furniture, internet, cleaning, and security. AWS Elastic Beanstalk, Google Cloud Run, Heroku, and Azure App Service are PaaS offerings.

With Cloud Run, you build a container image with your application code, push it to a registry, and the platform runs it. You never SSH into a server, never configure an OS, never manage a load balancer. Scaling from 1 to 100 instances happens automatically based on request volume. Scaling to zero when there is no traffic saves money.

Interview Tip

PaaS is the sweet spot for most web applications. You retain full control over your application code and dependencies (through your container or buildpack), but you offload OS management, patching, and scaling. The constraint is that your application must be stateless and horizontally scalable. If it can run as a container, PaaS is almost always cheaper in total cost (engineering time plus infrastructure) than IaaS.

The tradeoff is reduced control. If your application needs a custom kernel module, a specific filesystem mount, or GPU access with custom drivers, PaaS cannot accommodate you. The platform dictates the execution environment.

PaaS also introduces platform-specific constraints. Maximum request timeout limits (typically 5-30 minutes), container size limits, ephemeral filesystem storage, and mandatory statelessness are common restrictions. Your application must conform to these constraints, not the other way around. For most web applications and APIs, these constraints align naturally with good architecture practices. For long-running batch processes or applications that write to local disk, PaaS may not fit without significant refactoring.

Debugging on PaaS is different from debugging on IaaS. You cannot SSH into a PaaS container to inspect processes, check memory usage, or run diagnostic tools interactively. Instead, you rely on structured logging, distributed tracing, and the platform's monitoring dashboards. Engineers accustomed to SSH-based debugging find this transition frustrating initially, but structured observability (logs, metrics, traces) is actually more powerful at scale than interactive debugging because it works across hundreds of instances simultaneously. The investment in observability tooling early in a PaaS adoption pays dividends as the system grows.

PaaS platforms also handle TLS certificate management, load balancer configuration, and health checking automatically. On IaaS, each of these is a separate configuration task with its own failure modes. An expired TLS certificate on a self-managed load balancer causes an outage that is entirely preventable. PaaS platforms integrate with certificate authorities (Let's Encrypt, ACM) to provision and rotate certificates automatically. This is a class of operational incident that simply does not exist in a PaaS environment.

Software as a Service (SaaS)

SaaS is the far end of the spectrum. You manage nothing. No servers, no code, no configuration beyond settings in a web dashboard. Gmail, Salesforce, Slack, and Datadog are SaaS products. You consume the service through a browser or API.

The relevance to system design is that SaaS products are building blocks. Instead of building an email delivery system, you use SendGrid. Instead of building a monitoring pipeline, you use Datadog. Instead of building a CRM, you use Salesforce. Each SaaS tool you adopt is a "build versus buy" decision where you trade customization for speed.

The hidden cost of SaaS is vendor dependency. When a SaaS provider has an outage, your dependent features go down with it. When a SaaS provider changes their API, you update your integration on their timeline. When a SaaS provider raises prices, you either pay or undertake a migration project. The mitigation strategy is abstraction: wrap SaaS integrations behind interfaces in your code so you can swap providers without rewriting business logic. A NotificationService interface that can be backed by SendGrid, Mailgun, or SES protects you from single-vendor dependency.

SaaS also introduces data sovereignty concerns. When you store customer data in a SaaS provider's infrastructure, you may not control which region or country that data resides in. For regulated industries, this can be a compliance violation. Before adopting a SaaS tool for any workflow involving sensitive data, verify that the provider offers data residency controls and holds the compliance certifications your industry requires. Many SaaS providers offer enterprise tiers with dedicated infrastructure and data isolation specifically for regulated customers, but these tiers typically cost 3-10x more than standard plans.

Integration complexity is the final SaaS consideration. Each SaaS tool you adopt adds an integration point: API authentication, webhook handling, error management, rate limiting, and data format mapping. Five SaaS integrations means five APIs to monitor, five vendors to manage, and five potential points of failure in your system. Some teams end up building and maintaining a significant "integration layer" that consumes as much engineering effort as building the feature in-house would have. Evaluate not just the SaaS tool's functionality but the total integration cost including ongoing maintenance of the connection.

Function as a Service (FaaS) and Serverless

FaaS extends PaaS to its logical extreme: you deploy individual functions, not applications. AWS Lambda, Google Cloud Functions, and Azure Functions execute your code in response to events (HTTP requests, queue messages, database changes, scheduled triggers). The platform provisions a container, runs your function, returns the result, and deallocates resources. Between invocations, you consume zero resources and pay nothing.

A function invoked after a quiet night, with the container provisioning underneath, the billed milliseconds, and what the caller waited.

The cold start problem is real. When a function has not been invoked recently, the platform must provision a new container, load your code, and initialize your runtime. This adds 100ms to 2 seconds of latency on the first request. Subsequent requests reuse the warm container and complete in milliseconds. For latency-sensitive APIs, this variability can be unacceptable. For event-driven batch processing, it is irrelevant.

Cold start duration depends on three factors: the runtime language, the size of your deployment package, and the initialization work your code performs. A Python function with a 5 MB package that imports TensorFlow on startup might cold start in 3 seconds. A Go function with a 10 MB binary that connects to a database pool might cold start in 200ms. The mitigation strategies include keeping deployment packages small, lazy-loading heavy dependencies, using compiled languages for latency-sensitive functions, and provisioned concurrency (paying to keep a minimum number of containers warm at all times).

FaaS also imposes execution constraints. Maximum execution time (15 minutes on Lambda), maximum memory (10 GB), maximum payload size (6 MB synchronous, 256 KB asynchronous), and no persistent local state. These constraints are intentional: they force you to write small, focused functions that complete quickly and store state externally. Workloads that need longer execution times, larger memory, or persistent local state are better suited to PaaS or IaaS.

Despite these constraints, FaaS excels at event-driven architectures. A file uploaded to object storage triggers a Lambda function that generates thumbnails. A message arriving in a queue triggers a function that sends a notification email. A database change event triggers a function that updates a search index. In each case, the function does one thing in response to one event. The platform handles concurrency: if 1,000 files are uploaded simultaneously, 1,000 function instances run in parallel, each processing one file. You write sequential code that handles one event. The platform provides the parallelism. This combination of simple code and automatic scaling makes FaaS the ideal model for glue logic between cloud services.

FaaS pricing is per-invocation and per-millisecond of execution time. At low volumes (under 1 million invocations per month), FaaS is dramatically cheaper than a reserved EC2 instance running 24/7. At high volumes (10 million+ invocations per month), the per-invocation cost often exceeds the cost of a dedicated server running the same workload continuously. The crossover point depends on your function's execution duration, memory requirements, and traffic patterns.

How to Choose the Right Model

The decision framework is straightforward. Ask three questions in order:

Do you need OS-level control? If you need custom kernel modules, specific filesystem configurations, GPU drivers, or legacy software that requires a particular OS version, IaaS is your only option. PaaS and FaaS abstract the OS away.

Is your workload event-driven and short-lived? If each unit of work completes in under 15 minutes, is triggered by an event (HTTP request, queue message, file upload), and has idle periods, FaaS is the most cost-efficient choice. You pay nothing when idle.

Everything else points to PaaS. Stateless web applications, APIs, background workers that run continuously, and containerized services all fit the PaaS model. You get auto-scaling and managed infrastructure without the operational burden of IaaS or the execution constraints of FaaS.

Many production systems use all three models simultaneously. The web API runs on PaaS, the ML inference pipeline runs on IaaS GPU instances, and image thumbnail generation runs on FaaS. Choosing the right model per component, rather than one model for the entire system, optimizes both cost and operational complexity.

Cost Comparison Across Models

Understanding the cost structure of each model prevents surprises on your monthly bill.

IaaS costs are predictable but inflexible. You pay for VMs by the hour regardless of utilization. A t3.large instance running 24/7 costs approximately the same whether it is handling 1,000 requests per second or sitting idle. You also pay for storage volumes, network transfer, and any additional services (load balancers, elastic IPs). The total cost is the sum of infrastructure components, and you optimize it through right-sizing instances and using reserved pricing for stable workloads.

PaaS costs scale with usage but include a platform premium. Cloud Run charges per request and per vCPU-second of execution. At zero traffic, the cost is zero. At high traffic, the per-request cost accumulates. The platform premium (typically 20-40% over raw compute cost) pays for auto-scaling, load balancing, health checking, and zero-ops deployment. For most teams, this premium is far cheaper than hiring an additional operations engineer.

FaaS costs are the most granular: per invocation (typically $0.20 per million) plus per-millisecond of execution time. The granularity is both an advantage and a trap. At low volumes, FaaS is nearly free. At high volumes, the per-invocation overhead exceeds the cost of a dedicated server. The crossover point varies by workload, but a general rule is that workloads exceeding 10 million invocations per month with average execution times over 500ms should be evaluated against PaaS or IaaS alternatives.

The total cost of ownership extends beyond infrastructure bills. Engineering time spent on operations (IaaS), platform constraints (PaaS), and cold start optimization (FaaS) all have dollar values. A $500/month VM that requires 10 hours of engineering time per month for patching and monitoring costs more than a $700/month PaaS deployment that requires zero operations time, if the engineer's loaded cost exceeds $20/hour.

Managed services add another dimension to cost analysis. A self-managed PostgreSQL instance on IaaS costs less per hour than a managed RDS instance. But the managed instance includes automated backups, failover, patch management, and monitoring. The operations work you avoid with managed services is work that would otherwise consume senior engineering time. For a startup where every engineer's time is the scarcest resource, the managed service premium is almost always worth paying. For a large enterprise with a dedicated database operations team, self-managed databases on IaaS might be cheaper because that team's salary is already paid regardless of which services they manage.

The key principle is to evaluate total cost of ownership (TCO), not sticker price. TCO includes compute costs, storage costs, networking costs, engineering time for operations, training costs for new tools, and opportunity cost of engineering time spent on infrastructure instead of product development. The service model with the lowest sticker price is often not the one with the lowest TCO.

A practical exercise for evaluating TCO: estimate the monthly infrastructure bill for each service model, then add the engineering hours per month for operations (patching, monitoring, scaling, incident response, on-call rotations). Multiply engineering hours by the loaded cost of an engineer (salary plus benefits plus office space, typically $75-150/hour at tech companies). Add that to the infrastructure bill. The result is the true monthly cost. For many teams, the engineering time component is larger than the infrastructure component, which makes PaaS and managed services the cheaper option even when their sticker price is higher.

Teams should also account for the cost of incidents. IaaS workloads tend to have more operational incidents (disk full, OS kernel panic, failed patch, expired certificate) because more components are under your management. Each incident consumes engineering time and potentially causes revenue loss through downtime. PaaS reduces the surface area for operational incidents, which reduces the expected cost of incidents per month. This actuarial view of operations cost often tips the TCO calculation decisively toward managed services.