Schema Evolution and Compatibility

Topics Covered

Forward and Backward Compatibility

Backward Compatibility

Forward Compatibility

Full Compatibility

Database Schema Migration

The Problem with ALTER TABLE

Expand and Contract Pattern

Online Schema Change Tools

Migration Ordering and Dependencies

API Versioning Strategies

Why APIs Break

URL Path Versioning

Header and Content Negotiation Versioning

Query Parameter Versioning

Choosing a Strategy

Sunset Periods and Deprecation

Evolving Data in Practice

Event Schema Evolution in Message Queues

The Cost of Schema Coupling

Practical Rules for Safe Evolution

Systems do not upgrade all at once. During a rolling deploy, old servers and new servers run side by side for minutes or hours. A mobile app released today coexists with versions from six months ago. A Kafka consumer deployed last week reads messages produced by code deployed today. Every one of these situations requires one version of code to correctly interpret data written by a different version. That is what compatibility is about.

How much freedom you have here is decided by the encoding, covered in data serialization formats. It matters most where the data outlives the code that wrote it: a topic replayed from its beginning in event streaming fundamentals, a table rewritten by a new engine in lakehouse and open table formats, and a change stream consumed by systems the database owner has never met, in change data capture.

The size of the problem is combinatorial, which is the argument for a rule rather than a test matrix. With nn versions live at once, the number of reader and writer pairings that must work is

P=n(n−1)P = n(n-1)

Three live versions is 6 pairings, five is 20, and ten is 90. Nobody tests 90 pairings. What makes this tractable is transitive compatibility: if every single change is independently backward and forward compatible, then every pairing is compatible by induction, and you verify one change against the rule rather than every version against every other. That is why the useful discipline is a per-change rule enforced in CI, not an integration suite.

The reason this matters is that data outlives code. A database record created two years ago is still read by code deployed today. A message sitting in a Kafka topic for a week must be parseable by consumers from last month and consumers from next week. If your schema changes break the ability to read old data or write data that old code can understand, you create a coupling between deployment schedules that destroys the independence your architecture was designed to provide.

Think of compatibility as a contract between the past and the future. Backward compatibility is a promise from the future: "I will understand your old data." Forward compatibility is a promise from the past: "I will tolerate your new data, even if I do not understand all of it."

Six instances moving from v1 to v2 one at a time, both versions writing the same data while the deploy is still in flight.

Backward Compatibility

Backward compatibility means new code can read old data. When you deploy a new version of your service, it will encounter records, messages, and API responses created by the previous version. If a field was added in the new schema, old data will not contain it. The new code must handle that absence gracefully, typically by filling in a default value.

New code reading records written before a field existed, filling in the default instead of failing.

This is the easier direction. You control the new code. You know which fields are new. You write the defaults. Most serialization frameworks (Protobuf, Avro, Thrift) handle this automatically: every field has a default value, and missing fields resolve to that default during deserialization.

Consider a concrete example. Your Customer service adds a loyalty tier feature. The new schema includes a loyalty_tier field. But the database still has millions of customer records created before this field existed. When the new code reads those old records, loyalty_tier is missing. Rather than crashing, the code fills in the default value (empty string or "NONE") and continues processing.

 
1// New schema adds "loyalty_tier" field
2message Customer {
3  string name = 1;
4  string email = 2;
5  string loyalty_tier = 3;  // default: ""
6}
7// Old data without loyalty_tier deserializes fine:
8// loyalty_tier = "" (empty string default)

The danger with backward compatibility is assuming new fields have real data. During the transition period, most records will have the default value. If your business logic treats an empty loyalty tier as "not a loyalty member" rather than "unknown, needs backfill," you might accidentally downgrade existing loyal customers. Defaults should be semantically safe, not just syntactically valid.

Forward Compatibility

Forward compatibility means old code can read new data. This is harder because you are asking code that was written before a field existed to handle data that contains that field. The old code must ignore fields it does not recognize rather than crashing on unexpected input.

Why does this matter? During a rolling deploy, the new servers start producing data with the new field. Old servers that have not been upgraded yet must still read that data. If old code throws an error on an unknown field, your rolling deploy becomes a cascading failure.

Protobuf handles this well: unknown field tags are preserved in the binary but skipped during deserialization. JSON-based systems need explicit coding discipline. If your REST API consumer uses strict deserialization that rejects unknown fields, adding a field to the response is a breaking change even though it looks harmless.

The asymmetry between forward and backward compatibility explains a common deployment ordering rule: deploy consumers before producers when adding fields. If you deploy the new producer first, it starts emitting data with new fields that old consumers have not seen. If old consumers reject unknown fields, they fail. Deploying consumers first ensures they already know how to ignore the new field by the time it appears. Of course, if your consumers are third-party applications you do not control, you cannot dictate their deployment order. That is why forward compatibility must be a property of the consumer's parser, not something you coordinate operationally.

Interview Tip

Configure JSON deserializers to ignore unknown fields by default. In Jackson, use @JsonIgnoreProperties(ignoreUnknown = true). In Go, encoding/json already ignores unknown fields. In Python, use model_config = ConfigDict(extra='ignore') with Pydantic. This single setting is the difference between forward-compatible and forward-breaking.

Full Compatibility

Full compatibility means a schema change is both forward and backward compatible. Old code reads new data without breaking. New code reads old data without breaking. This is the gold standard for any system where producers and consumers deploy independently, which is nearly every distributed system.

Achieving full compatibility constrains what changes you can make:

Safe changes (fully compatible): Adding an optional field with a default. Renaming a field in Protobuf (only the tag number matters, not the name).

Backward compatible only: Adding a required field. New code handles the default, but old code that does not know about the field cannot produce valid new data.

Breaking (neither direction): Removing a field that old code depends on. Changing a field's type. Renaming a JSON field.

The practical implication is that fully compatible changes are almost always additive and optional. You can grow a schema by adding new fields forever without breaking anyone. Shrinking a schema (removing fields) or mutating it (changing types or meanings) requires coordination. This is why mature schemas tend to accumulate fields over time. A Protobuf message that started with 5 fields might have 30 fields after three years of evolution. Each field tells the story of a feature that was added, and removing any of them risks breaking a consumer somewhere that still depends on it.

Level Expectations

Mid-level engineers understand that adding fields is safe. Senior engineers design schemas with evolution in mind from day one: every field is optional, every message has a version indicator, and removal follows a deprecation timeline. Staff engineers establish organization-wide compatibility policies and enforce them through schema registries and CI checks that reject breaking changes before they merge.