0%
Cloud Architecture Patterns
Cloud Foundations
Compute Patterns
Storage and Databases
Application Patterns
Reliability and Operations
Advanced Patterns
CI/CD Pipelines
A CI/CD pipeline is a sequence of automated steps that transforms a code commit into a production deployment. The "CI" half (Continuous Integration) validates every change. The "CD" half (Continuous Delivery or Deployment) ships validated changes to users. Separating them matters because CI can run on every pull request while CD only triggers on merges to the main branch.
The reason pipelines exist is not automation for its own sake. It is repeatability. A manual deployment process depends on whoever runs it remembering every step, in the right order, without mistakes. A pipeline encodes those steps once and executes them identically every time, whether it is 2 AM on a Sunday or noon on a Tuesday.
Pipelines also create a feedback loop. When a developer pushes code, the pipeline runs tests in minutes and reports results. Without a pipeline, the feedback loop stretches to hours or days: code gets reviewed, merged, manually deployed to staging, manually tested, and only then does someone discover that the database migration fails. By that point, three more PRs have merged on top of it. Pipelines compress this loop to minutes, catching failures while the developer still has the code fresh in mind.
Continuous Delivery vs. Continuous Deployment
These terms sound interchangeable but describe different levels of automation. Continuous Delivery means every commit that passes the pipeline is ready to deploy. A human still clicks "deploy" to push to production. Continuous Deployment means every commit that passes the pipeline deploys to production automatically, with no human intervention.
Most teams start with Continuous Delivery and graduate to Continuous Deployment as their test suites, monitoring, and rollback capabilities mature. The prerequisite for Continuous Deployment is confidence: if your automated tests catch 99% of bugs and your rollback takes 30 seconds, automated deployment is safe. If your tests catch 70% and rollback takes 20 minutes, you want a human reviewing the staging results before pushing the button.
CI Pipeline Stages
A typical CI pipeline runs these stages in order, where each stage gates the next:
Lint and static analysis runs first because it is the fastest check. It catches formatting violations, unused imports, and common code smells in seconds. Failing here avoids wasting minutes on tests for code that does not meet basic standards.
Unit tests run next. These test individual functions and classes in isolation, typically completing in under a minute. They catch logic errors early. The key metric here is coverage: most teams require 70-80% line coverage as a minimum gate.
Build compiles the code (for compiled languages) or packages it (for interpreted languages). If the build fails, there is no artifact to test further. This stage also produces the build artifact that downstream stages use: a Docker image, a JAR file, a compiled binary.
Integration tests run after the build because they need the actual artifact. These tests verify that components work together: API endpoints return correct responses, database queries execute properly, message queues process events. They take longer (2-10 minutes) but catch issues that unit tests miss, like serialization mismatches between services.
Security scan analyzes the artifact and its dependencies for known vulnerabilities. Tools like Snyk, Trivy, or Dependabot check dependency versions against CVE databases. A critical vulnerability in a dependency should block the pipeline. Medium or low severity findings can be tracked as tech debt without blocking. Static Application Security Testing (SAST) goes further by analyzing your source code for patterns like SQL injection, hardcoded secrets, or insecure cryptographic usage. The combination of dependency scanning and SAST catches both external and internal security issues.
Artifact publish is the final CI stage. The validated, scanned artifact is pushed to a registry (Docker Hub, ECR, Artifactory) with a unique tag, typically the git commit SHA. This tag is immutable: the same SHA always produces the same artifact. Downstream CD stages pull from this registry, never from source code.
Each stage produces metadata that downstream stages and humans can inspect: test results with pass/fail counts and coverage percentages, build logs with compilation output, security scan reports listing discovered vulnerabilities. Storing this metadata alongside the artifact creates a complete provenance record. When a production incident occurs at 3 AM, the on-call engineer can trace the running artifact back to its build, see exactly which tests passed, and determine whether the security scan flagged anything relevant.
CD Pipeline Stages
The CD pipeline takes the published artifact and deploys it through environments:
Deploy to staging pulls the artifact from the registry and deploys it to a staging environment that mirrors production. Same database schema, same configuration structure, same resource limits. The only differences should be scale (fewer instances) and data (synthetic or anonymized).
Smoke tests verify that the deployment actually works. These are not comprehensive tests. They are 5-10 checks that confirm the application starts, serves its health endpoint, connects to its database, and processes a basic request. If smoke tests fail, the deployment is broken at a fundamental level. Keep smoke tests fast (under 60 seconds) and focused. They exist to catch deployment failures (misconfiguration, missing environment variables, connectivity issues), not application logic bugs. That is what the CI tests already verified.
Deploy to production pushes the same artifact to production using one of the deployment strategies covered in the next section. The critical point is that the exact same artifact deployed to staging is deployed to production. No rebuild, no repackage, no "it works on staging but not production" surprises.
Health checks and monitoring confirm the production deployment is healthy. Health checks verify the application is responsive. Monitoring watches error rates, latency percentiles, and business metrics (orders per minute, signups per hour) for the first 15-30 minutes after deployment. The monitoring window matters: some bugs only manifest under sustained load or after cache entries expire. A 5-minute monitoring window catches crashes but misses memory leaks that surface after 20 minutes. Calibrate the window based on your historical deployment failure patterns.
The most common pipeline mistake is rebuilding the artifact for each environment. If you build once for staging and rebuild for production, you are deploying a different artifact than the one you tested. Always build once, tag with the commit SHA, and promote the same artifact through environments.
Quality Gates
A quality gate is a pass/fail checkpoint between stages. It turns a sequence of stages into a pipeline with enforcement. Without gates, a failing security scan produces a warning that everyone ignores. With gates, it blocks the deployment.
Coverage threshold: The pipeline fails if test coverage drops below a configured minimum. This prevents the gradual erosion of test coverage that happens when teams say "we will add tests later." A gate at 75% coverage means every PR must maintain or improve coverage. Some teams use a stricter variant: the coverage of changed lines must be above 90%, even if the overall project coverage is lower. This ensures new code is well-tested without requiring teams to retroactively test legacy code.
Security scan pass: Critical and high severity vulnerabilities block the pipeline. The team must either update the vulnerable dependency, apply a patch, or explicitly approve an exception with a documented reason and timeline.
Manual approval: For production deployments, some organizations require a human to click "approve" after reviewing staging test results. This is the one gate that is not automated. It exists for compliance (SOC 2, HIPAA) or for high-risk changes where human judgment adds value. The tension with manual approval is that it introduces delay: if the approver is in a meeting, the deployment waits. Teams often compromise by requiring manual approval only for specific categories of changes (database migrations, security-sensitive code, changes touching payment flows) while fully automating the rest.
Flaky Tests and Pipeline Reliability
A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are the single biggest threat to pipeline trust. When a pipeline fails and the team's first reaction is "probably a flaky test, re-run it," the pipeline has lost its authority. Developers stop investigating failures and start re-running, which means real bugs slip through alongside the false alarms.
The fix is not tolerating flaky tests. Quarantine them: move flaky tests to a separate non-blocking suite, fix them, and move them back. Track flake rates per test. Any test that fails more than 1% of runs without a code change goes to quarantine. This preserves pipeline trust for the tests that remain: when the pipeline fails, it means something real.
Pipeline as Code
Modern CI/CD pipelines are defined in configuration files committed to the repository: .github/workflows/ci.yml for GitHub Actions, Jenkinsfile for Jenkins, .gitlab-ci.yml for GitLab CI. This means the pipeline definition is versioned, reviewed, and tested alongside the application code.
Pipeline as code solves the "snowflake pipeline" problem. When pipelines are configured through a web UI, changes are untracked, unreviewable, and unreproducible. If the pipeline breaks, nobody knows what changed or when. With pipeline as code, a broken pipeline shows up in git log as a specific commit by a specific person, making debugging straightforward.
It also enables branch-specific pipelines. A feature branch can modify the pipeline to add a new test suite, and that modification only affects builds on that branch. Once the branch merges, the pipeline change merges with it. This is impossible with UI-configured pipelines where the pipeline definition is global.
Environment Parity
The principle of environment parity states that staging should be as close to production as possible. The same Docker image, the same configuration structure, the same database engine version, the same network topology. The only acceptable differences are scale (fewer replicas), data (synthetic or anonymized), and external service endpoints (sandbox payment APIs, test email servers).
Violations of environment parity are the root cause of "works in staging, breaks in production" incidents. Common violations include: staging uses SQLite while production uses PostgreSQL, staging skips the CDN layer, staging has no rate limiting, or staging uses a different version of a shared library. Each violation is a gap where bugs can hide.
Containerization (Docker) helps enforce parity by packaging the application and its dependencies into a single image. If staging runs the same Docker image as production, the application-level parity is guaranteed. Infrastructure parity (load balancers, networking, DNS) requires additional effort through Infrastructure as Code tools like Terraform or Pulumi.
The key principle is that gates should be binary: pass or fail. A gate that says "warning: coverage is low" is not a gate. A gate that blocks the pipeline until coverage reaches 75% is a gate.