Environment Management

Topics Covered

Dev, Staging, and Production

Development Environment

Staging Environment

Production Environment

Deployment Strategies in Production

Preview Environments

Environment Parity

What Must Be the Same

What Can Differ

Infrastructure as Code Enables Parity

Config Drift Detection

Container Images and Build-Once-Deploy-Everywhere

Measuring Parity

Feature Flags

Types of Feature Flags

The Flag Lifecycle

Gradual Rollout Strategy

Implementation Considerations

Flag Targeting and User Segmentation

Kill Switches and Operational Flags

Flag Management Platforms

Configuration Management

Environment Variables

Centralized Configuration Stores

Secret Management

Terraform Workspaces for Environment Configuration

The 12-Factor Approach

Configuration Validation

Configuration Hierarchy and Overrides

Every serious software team runs multiple environments because deploying code directly to production is gambling with user experience. The environment hierarchy exists to catch problems early, when fixing them is cheap, instead of late, when an outage costs revenue and trust.

The standard hierarchy is three tiers: development, staging, and production. Each tier has a distinct purpose, and understanding those purposes prevents the common mistake of treating environments as arbitrary copies of each other.

One artifact promoted through four tiers, with the gate it has to pass and the data it reads at each one.

Development Environment

The development environment is where engineers write and test code locally or on shared dev servers. It prioritizes speed of iteration over stability. Databases are small, seeded with synthetic data. External services are mocked or pointed at sandbox APIs. The goal is fast feedback: write code, run it, see if it works, repeat.

Dev environments tolerate instability. If a dev database gets corrupted, you re-seed it in minutes. If a dev server crashes, nobody's checkout flow breaks. This freedom to break things is the entire point. Engineers experiment without consequences to real users.

Common dev environment patterns include Docker Compose (defining all services in a single file for local orchestration), devcontainers (standardized VS Code development environments), and cloud-based dev environments like GitHub Codespaces or Gitpod that provision remote machines with pre-configured toolchains. The choice depends on the application's complexity. A simple web app works fine with Docker Compose. A system with 15 microservices, a message queue, and three databases may need a cloud-based dev environment because laptops cannot run the full stack.

Staging Environment

Staging is the dress rehearsal. It mirrors production's architecture, configuration structure, and deployment pipeline, but runs at smaller scale. The staging database has the same schema as production. The same container images that will run in production are deployed to staging first. Load balancers, message queues, caches, and monitoring are all present.

The critical rule: code reaches production only after passing through staging. This is where integration tests run against real infrastructure (not mocks), QA teams verify features end-to-end, and performance baselines are validated. If something breaks in staging, you caught it before it reached users.

Staging should also mirror production's data characteristics. While you cannot (and should not) use real user data, staging databases should contain anonymized data at a scale that exercises the same query patterns. An empty staging database will not reveal the slow query that only appears when the users table has 10 million rows. Many teams use production database snapshots with automated anonymization pipelines to keep staging data realistic without compromising privacy.

 
1Promotion pipeline:
2dev -> staging -> production
3
4At each gate:
5  - Automated tests must pass
6  - Security scans must clear
7  - Performance benchmarks must hold
8  - Required approvals must be collected
Interview Tip

In interviews, explain that staging exists to validate the deployment itself, not just the code. The same Terraform, the same Helm charts, the same CI/CD pipeline run in staging first. If the deployment script has a bug, staging catches it before production.

Production Environment

Production serves real users. It has the strongest access controls, the most monitoring, and the most conservative change management. Deployments are automated, gradual (canary or blue-green), and instantly reversible. Every change is audited.

Production differs from staging in scale, not in kind. More replicas, larger databases, higher-tier cloud instances, multi-region distribution. The architecture is identical. If production has a Redis cluster, staging has a Redis cluster (smaller, but present). If production uses a managed Kafka service, staging uses the same Kafka service (fewer partitions, but same configuration structure).

Deployment Strategies in Production

Production deployments use strategies that minimize risk:

Blue-green deployment: Two identical production environments (blue and green) run simultaneously. Traffic points at blue. You deploy the new version to green, run smoke tests against green, then switch the load balancer to route traffic to green. If something breaks, switch back to blue in seconds. The cost is double the infrastructure during the transition window.

Canary deployment: Route 1-5% of traffic to the new version while 95-99% continues hitting the old version. Monitor error rates, latency, and business metrics for the canary group. If metrics look good, gradually shift more traffic. If metrics degrade, route all traffic back to the old version. This is the production equivalent of a feature flag rollout but at the infrastructure level.

Rolling update: Replace instances one at a time. Take instance 1 out of the load balancer, deploy the new version, put it back. Repeat for instances 2, 3, and so on. At any point, a mix of old and new versions handles traffic. This works well when the new version is backward-compatible with the old version.

Each strategy trades off between speed, cost, and risk. Blue-green is the safest but doubles infrastructure cost during deployment. Canary is the most granular but requires robust metric monitoring to detect issues at low traffic percentages. Rolling updates are the cheapest but cannot fully roll back without redeploying the old version. Most teams start with rolling updates and adopt canary deployments as they mature their monitoring infrastructure.

Preview Environments

Modern teams add a fourth tier: preview environments. These are ephemeral, per-pull-request environments that spin up automatically when a developer opens a PR and tear down when the PR is merged or closed.

One environment per pull request, with what it costs to run and what a reviewer can actually do inside it.

Preview environments let reviewers see a running version of the change, not just a code diff. A product manager can click through the new checkout flow in a real browser instead of imagining it from a PR description. When the PR merges, the environment is destroyed, freeing resources. Tools like Vercel, Netlify, and Argo CD make this pattern straightforward.

 
1PR opened    -> spin up preview-env-pr-347
2PR updated   -> redeploy preview-env-pr-347
3PR merged    -> destroy preview-env-pr-347
4PR closed    -> destroy preview-env-pr-347

The cost is infrastructure complexity. Each preview environment needs its own database instance (or schema), its own secrets, and its own DNS entry. Teams that manage this well use templates: a single Terraform module parameterized by PR number creates the full stack in minutes.

Preview environments also improve the code review process. Instead of reading a PR description that says "updated the payment form layout," a reviewer can click a preview URL, fill in a test credit card number, and see the new layout in action. Visual changes, interaction patterns, and edge cases that are invisible in a code diff become obvious in a running preview. Product managers, designers, and QA engineers who cannot read code can still evaluate changes directly.

 
1# CI generates a unique URL per PR
2Preview URL: https://pr-347.preview.myapp.dev
3Status: Ready (deployed 2 minutes ago)
4Branch: feature/new-checkout