Identity and Access Management

Topics Covered

IAM Fundamentals

Users and Groups

Policy Evaluation Logic

Authentication: Proving Identity

The Root Account Problem

Identity Federation

The Principle of Least Surprise

Policies and Permissions

Policy Document Structure

Identity-Based vs Resource-Based Policies

Conditions: Fine-Grained Access Control

Policy Best Practices

Roles and Service Accounts

How Role Assumption Works

Service Accounts for Applications

Cross-Account Access

When Credentials Are Compromised

Confused Deputy Problem

Role Chaining

Session Policies for Dynamic Scoping

Instance Profiles and Metadata Service Security

Least Privilege in Practice

Why Over-Permissioning Happens

Permission Boundaries

Auditing and Continuous Improvement

Access Key Hygiene

Service Control Policies for Organizational Guardrails

Automating Least Privilege with Infrastructure as Code

Tagging for Access Control

Identity and Access Management (IAM) answers two questions for every request that enters your system: who is making this request, and are they allowed to do what they are asking? Authentication answers the first question. Authorization answers the second. Every cloud provider, every database, every API gateway performs this check on every request. If you get IAM wrong, nothing else matters. A perfectly designed database schema and a flawless caching layer are worthless if an attacker can read every record because permissions were misconfigured.

IAM is built around four types of principals: users, groups, roles, and service accounts. Understanding the difference between these is the foundation for everything that follows.

Users and Groups

A user is a single identity. It has credentials (password, access key, MFA device) and is typically associated with a human being. When an engineer joins your team, you create a user for them. When they leave, you deactivate that user.

The problem with managing permissions at the user level is scale. A team of 50 engineers where each person needs access to 10 services means 500 individual permission assignments. When you add a new service, you update 50 users. When someone changes teams, you audit every permission individually.

Groups solve this. A group is a named collection of users that share the same permissions. You create a group called "backend-engineers," attach policies granting access to the application database, deployment pipeline, and log aggregation service. When a new engineer joins, you add them to the group. One action grants all necessary permissions. When someone moves to a different team, you remove them from one group and add them to another.

One request walked through the four gates of policy evaluation in order, with the single step that actually decides the outcome.

Policy Evaluation Logic

When a principal makes a request, the IAM system evaluates all applicable policies. The evaluation follows a strict hierarchy:

Step 1: Gather all policies. Collect identity-based policies (attached to the user, their groups, and any assumed role), resource-based policies (attached to the target resource), permission boundaries (if any), and organization-level service control policies.

Step 2: Check for explicit deny. If any single policy contains an explicit deny that matches the action and resource, the request is denied. Period. It does not matter how many allow statements exist elsewhere. One explicit deny overrides everything.

Step 3: Check for explicit allow. If no deny was found, check whether any policy explicitly allows the action on the resource. If yes, access is granted.

Step 4: Default deny. If no policy explicitly allows the action, access is denied. This is the principle of default deny: everything is forbidden unless explicitly permitted.

This "deny wins" model is critical for security. It means you can grant broad permissions through group policies and then carve out exceptions with deny statements. A developer group might have full access to S3, but a specific bucket containing PII data has an explicit deny for all non-compliance team members.

Interview Tip

In system design interviews, when discussing access control, always mention that IAM follows a default-deny model with explicit-deny-wins evaluation. This signals that you understand the security-first design philosophy. Many candidates describe only allow rules and miss the fact that the absence of an allow is itself a deny.

Authentication: Proving Identity

Proving identity requires credentials, and the type of credential determines the security posture. Password-based authentication is the weakest form. Passwords are reused, phished, and stored in plaintext in configuration files more often than anyone admits.

Multi-factor authentication (MFA) adds a second verification step: something you have (a hardware token, a phone app generating TOTP codes) in addition to something you know (the password). MFA reduces the impact of credential theft dramatically. Even if an attacker obtains a password through phishing, they cannot authenticate without the second factor.

Access keys are long-lived credentials used for programmatic access. They consist of an access key ID and a secret access key. Unlike passwords, access keys do not expire unless you explicitly rotate them. This makes them dangerous: a leaked access key in a Git commit grants an attacker persistent access until someone notices and rotates the key. Automated scanners on GitHub continuously scrape public repositories for AWS access key patterns (keys start with AKIA for permanent keys and ASIA for temporary credentials). Attackers who find valid keys often spin up cryptocurrency mining instances within minutes, racking up thousands of dollars in charges before the key owner is alerted. Best practice is to avoid access keys entirely and use roles with temporary credentials instead.

The Root Account Problem

Every cloud account has a root user with unrestricted access to everything. The root user can delete billing configurations, close the account, and bypass every permission boundary. It is the most powerful and most dangerous identity in your system.

The rule is simple: never use the root account for daily operations. Create individual IAM users for every person and every service, even administrators. Lock the root credentials in a safe (literally, some companies store root credentials in physical safes). Enable MFA on root. Set up alerts for any root login. The root account should only be used for the handful of tasks that require it, like changing the account's billing settings or enabling certain AWS service features for the first time.

The reason this matters at scale is auditability. When three administrators share the root account, CloudTrail logs show "root" performed the action but not which human was responsible. When each administrator has their own IAM user with admin privileges, every action is traceable to a specific person. This traceability is not just for security investigations. It enables change management, compliance reporting, and incident postmortems. "Who changed the VPC configuration at 2 AM?" is answerable with individual users and unanswerable with shared root credentials.

Identity Federation

In organizations with existing identity providers (Active Directory, Okta, Google Workspace), creating separate IAM users for every employee is wasteful and introduces synchronization problems. When someone leaves the company, you must deactivate their corporate account AND their IAM user. If you forget the IAM user, you have an orphaned identity with active permissions.

Identity federation solves this by allowing users to authenticate through the corporate identity provider and receive temporary IAM credentials. The employee signs in through their company's single sign-on portal, the identity provider issues a SAML assertion or OIDC token, and AWS STS exchanges it for temporary credentials mapped to an IAM role. No IAM user is created. When the employee is deactivated in the corporate directory, they can no longer obtain federation tokens, and existing tokens expire naturally.

This approach unifies the identity lifecycle: one place to manage users, one place to enforce password policies, one place to revoke access.

Federation also enables role-based access mapping. Different corporate groups (engineering, finance, security) map to different IAM roles through the identity provider's attribute assertions. An engineer authenticating through the SSO portal receives credentials for the "engineer-role" with access to development resources. A security analyst receives credentials for the "security-audit-role" with read-only access across all accounts. The mapping is configured once in the identity provider and never requires individual IAM user management.

The security benefit of federation extends to incident response. When a security incident requires locking out a compromised employee, deactivating their corporate directory account instantly prevents them from obtaining new federation tokens. With separate IAM users, the security team must remember to deactivate accounts in every AWS account the employee had access to. Organizations with 10+ accounts regularly discover orphaned IAM users during audits, months after the employee departed, because someone forgot to clean up one account.

The Principle of Least Surprise

Well-designed IAM structures follow naming conventions and organizational patterns that any engineer can understand without reading documentation. Group names like backend-engineers-prod-readonly or data-team-staging-admin communicate membership, scope, and access level immediately.

Policy names should describe what they grant, not who they are attached to. A policy named AllowS3ReadAppLogs is reusable across any principal that needs log access. A policy named JohnsPolicy tells you nothing about its contents and cannot be safely reused.

Consistent conventions reduce misconfiguration risk. When every production role follows the pattern {service}-{environment}-role and every permission boundary follows {team}-boundary, an engineer reviewing an IAM configuration can quickly identify anomalies: a role without the expected boundary, a group with an unexpected policy attachment, or a service account that does not follow the naming convention.