Privacy and Compliance

Topics Covered

Regulatory Landscape

GDPR (General Data Protection Regulation)

CCPA (California Consumer Privacy Act)

HIPAA (Health Insurance Portability and Accountability Act)

Data Classification

Cross-Border Data Transfers

Regulatory Overlap and Conflict

Privacy by Design

Anonymization and Pseudonymization

Why True Anonymization Is Hard

K-Anonymity

Differential Privacy

Pseudonymization in Practice

Choosing the Right Approach

Right to Deletion

Finding All Copies of User Data

Cascade Deletion Across Services

The Backup Problem

Derived Data and ML Models

Consent Management

Data Retention Policies

Third-Party Data and Deletion Propagation

Deletion Verification

Audit Trails

What to Log

Immutability

Encryption at Rest and in Transit

Breach Investigation and Response

Real-Time Monitoring and Alerting

Audit Trails in Microservices

Periodic Access Reviews

Privacy regulations are not aspirational guidelines. They are laws with enforcement mechanisms, and violating them carries real financial penalties. GDPR fines can reach 4% of global annual revenue or 20 million euros, whichever is higher. CCPA violations cost $7,500 per intentional violation. HIPAA penalties range from $100 to $50,000 per violation, up to $1.5 million per year per violation category. These numbers make compliance a first-class engineering requirement, not a checkbox exercise.

The reason you need to understand regulations as an engineer (not just as a legal matter) is that regulations impose architectural constraints. They dictate where data can be stored, how long it can be retained, who can access it, and what must happen when a user requests deletion. These constraints flow directly into your database schema, access control layer, backup strategy, and data pipeline design. A system that ignores them from the start will require expensive retrofitting later.

Four processing purposes checked against consent records at the gateway, including one that was withdrawn after the fact.

GDPR (General Data Protection Regulation)

GDPR applies to any company that processes data of EU residents, regardless of where the company is headquartered. A startup in California with a single user in Berlin is subject to GDPR. The regulation rests on several core principles:

Lawful basis for processing: You must have a legal reason to collect and process personal data. The most common bases are consent (the user explicitly agreed), contract (you need the data to provide a service the user requested), and legitimate interest (you have a business reason that does not override the user's rights). You cannot collect data "just in case" or "for future use."

Data minimization: Collect only the data you need for the stated purpose. If your service requires an email address to send order confirmations, you cannot also collect the user's date of birth, phone number, and home address unless each field serves a specific, documented purpose.

Right to access: Users can request a copy of all personal data you hold about them. You must respond within 30 days. This means your system must be able to locate and export every piece of data associated with a user across every service, database, and data warehouse.

Right to deletion (Right to be Forgotten): Users can request that you delete all their personal data. This is not a soft delete or a status flag. The data must be removed from operational databases, data warehouses, analytics pipelines, backups (within reason), and any derived datasets. We cover the engineering challenges of this in a later section.

Right to portability: Users can request their data in a machine-readable format (typically JSON or CSV) so they can transfer it to another service.

Data breach notification: You must notify the supervisory authority within 72 hours of discovering a breach that affects personal data. You must also notify affected users if the breach poses a high risk to their rights.

CCPA (California Consumer Privacy Act)

CCPA applies to businesses that collect personal information of California residents and meet certain thresholds (annual revenue over $25 million, data on 100,000+ consumers, or 50%+ revenue from selling personal data). The key rights mirror GDPR but with differences:

Right to know: Consumers can request what personal information you collect, where it comes from, and who you share it with.

Right to delete: Similar to GDPR, but with broader exemptions (you can retain data needed for completing a transaction, security, legal obligations, or internal use compatible with the original collection purpose).

Right to opt-out of sale: Consumers can direct you to stop selling their personal information. This requires a "Do Not Sell My Personal Information" link on your website.

HIPAA (Health Insurance Portability and Accountability Act)

HIPAA governs Protected Health Information (PHI) in the United States. It applies to healthcare providers, health plans, healthcare clearinghouses, and their business associates (any vendor that handles PHI on their behalf). HIPAA mandates encryption, access controls, audit trails, and breach notification for any system that stores or transmits PHI. The standard is stricter than GDPR in some respects: PHI must be encrypted at rest and in transit, access must be role-based with minimum necessary privileges, and every access must be logged.

Interview Tip

In interviews, when discussing data storage or API design, mention regulatory constraints early. Saying 'we need to track consent per field because of GDPR' or 'PHI requires encryption at rest per HIPAA' signals architectural maturity. Interviewers want to see that you consider compliance as a design input, not an afterthought.

Data Classification

Before you can comply with any regulation, you need to know what data you have and how sensitive it is. Data classification is the foundation:

PII (Personally Identifiable Information): Any data that can identify a specific individual. Names, email addresses, phone numbers, IP addresses, device IDs, location data. GDPR and CCPA regulate PII.

Sensitive personal data: A subset of PII with stricter protections. Racial or ethnic origin, political opinions, religious beliefs, health data, biometric data, sexual orientation. GDPR requires explicit consent (not just legitimate interest) for processing sensitive data.

Internal data: Business data that is not personal. Revenue figures, server logs without user identifiers, aggregated analytics. Not subject to privacy regulations but may have internal access controls.

Public data: Data that is intentionally made public. Published blog posts, public social media profiles, open-source code. No privacy restrictions, but you still need to respect terms of service.

Classify every data field in your system. This classification determines encryption requirements, access controls, retention policies, and deletion scope. A system that treats all data equally either over-protects non-sensitive data (wasting engineering effort) or under-protects sensitive data (risking violations).

Cross-Border Data Transfers

GDPR restricts transferring personal data outside the EU to countries that the European Commission has not deemed "adequate" in data protection. The United States, for example, does not have an adequacy decision for general data transfers (the EU-US Data Privacy Framework covers only certified companies). This means you cannot simply store EU user data in a US-based AWS region without a legal mechanism.

The most common mechanisms for lawful cross-border transfers are:

Standard Contractual Clauses (SCCs): Pre-approved contract templates from the European Commission that your company signs with the data recipient. They impose GDPR-equivalent obligations on the recipient regardless of local law.

Binding Corporate Rules (BCRs): Internal policies approved by EU supervisory authorities for multinational companies transferring data between their own entities. BCRs are expensive and time-consuming to establish but cover all intra-group transfers once approved.

Data residency: The simplest technical solution is to keep EU data in EU regions. Deploy your database replicas and processing infrastructure within the EU. This avoids the legal complexity entirely but may increase infrastructure costs and operational complexity if your engineering team is based elsewhere.

For system design interviews, mentioning data residency requirements signals that you understand the real-world constraints on where data can be physically stored and processed. A globally distributed system cannot simply replicate data everywhere without considering regulatory boundaries.

Four data flows mapped to where the bytes physically land, with the transfer mechanism each one depends on to stay lawful.

Regulatory Overlap and Conflict

In practice, most companies face multiple regulations simultaneously. A healthcare company serving patients in California and Germany must comply with HIPAA, CCPA, and GDPR. The requirements sometimes conflict: HIPAA may require retaining records for 6 years while a GDPR deletion request demands removal within 30 days. The resolution depends on which legal obligation takes priority, and this is where legal counsel works with engineering to define data retention hierarchies.

The engineering implication is that your compliance system must be flexible enough to apply different rules to different data subjects based on their jurisdiction, the type of data, and the applicable regulations. A one-size-fits-all approach fails when regulations contradict each other.

Privacy by Design

GDPR explicitly mandates "data protection by design and by default." This means privacy cannot be an afterthought bolted onto a finished system. It must be embedded in the architecture from the start.

In practice, privacy by design means:

Default to minimum access: New features should collect the least data possible and restrict access to the smallest group needed. If a feature can work with aggregated data instead of individual records, use aggregated data. If a service does not need a user's email to function, do not pass the email to that service.

Separation of concerns for PII: Route PII through a dedicated data gateway that handles encryption, tokenization, and access logging. Downstream services receive only the tokens or derived values they need. This creates a single point of control for privacy policies rather than scattering PII handling across dozens of microservices.

Privacy impact assessments: Before launching any feature that collects or processes personal data, evaluate what data is collected, why, how long it will be retained, who will access it, and what safeguards protect it. Document these decisions. They become evidence of compliance during audits and help catch privacy issues before they reach production.

The cost of retrofitting privacy into an existing system is typically 5-10x higher than designing it in from the start. A system built without privacy controls stores PII everywhere, shares it freely between services, and has no deletion capability. Adding these capabilities after the fact requires touching every service, migrating every database, and auditing every data pipeline. Starting with privacy by design avoids this technical debt entirely.