Data Serialization Formats

Topics Covered

Text Formats JSON and XML

JSON: The Universal Lingua Franca

XML: Powerful But Verbose

JSON Variants and Extensions

Limitations Shared by All Text Formats

Binary Formats Protobuf and Avro

Protocol Buffers (Protobuf)

Apache Thrift

Apache Avro

Comparing the Three Formats

Schema Evolution Guarantees

Field Tags and Compatibility Rules

Avro's Schema Resolution

Schema Registries

Schemas as Documentation and Contracts

When to Use Which Format

External APIs and Human-Readable Interfaces

Internal Service Communication

Data Pipelines and Storage

Decision Framework

Every time two processes exchange data, they face the same problem: in-memory data structures (objects, hash maps, trees) use pointers and memory layouts that are meaningless to another process. A Java object on Server A contains references to heap addresses that mean nothing to a Python process on Server B. Serialization converts those structures into a sequence of bytes that can be written to a file, sent over a network, or stored in a database. The receiving side deserializes those bytes back into usable data structures in its own language and memory model.

The format is a long-lived decision because it is the contract between systems that deploy independently, which is why it is inseparable from schema evolution and compatibility. It also sets the cost of everything downstream: the bytes on the wire in event streaming fundamentals, the bytes at rest in data warehouse architecture, and the shape of the change events in change data capture.

This conversion happens constantly in distributed systems. Every API call serializes a request and deserializes a response. Every database write serializes a row and every read deserializes it. Every message on a queue is serialized by the producer and deserialized by the consumer. The format you choose for that byte sequence determines how fast it encodes, how large it is on the wire, whether humans can read it, and how safely you can change the schema over time.

The two broad categories are text formats (human-readable) and binary formats (machine-optimized). This section covers the text side: JSON and XML. Understanding their strengths and limitations is essential because you need to know exactly when human readability is worth the performance cost, and when it is not.

JSON: The Universal Lingua Franca

JSON won the web because it is dead simple. Objects are key-value pairs. Arrays are ordered lists. Values are strings, numbers, booleans, or null. No schema needed. No code generator required. Any language can parse it. Any developer can read it.

json
1{
2  "userId": 12345,
3  "name": "Alice Chen",
4  "roles": ["admin", "editor"],
5  "active": true
6}

This readability is why JSON dominates external APIs. A developer debugging a failed request can copy the response body into a text editor and immediately understand the structure. Logs, error messages, and monitoring dashboards show JSON payloads without special tooling. When an on-call engineer is investigating a 3 AM incident, they can curl an endpoint, pipe the response through jq, and understand the problem in seconds. No code generator, no schema file, no special tooling needed.

JSON also maps directly to native data structures in most languages. A JSON object becomes a Python dictionary, a JavaScript object, a Java HashMap, or a Go map. This zero-friction mapping is why every HTTP client library, every test framework, and every logging system speaks JSON natively.

But JSON pays a tax for that readability. Every field name is repeated in every single record. In a batch of 10,000 user objects, the string "userId" appears 10,000 times. Numbers are stored as decimal text, so the integer 12345 takes 5 bytes instead of 2. There is no distinction between integers and floating-point numbers: the number 42 and the number 42.0 may parse to different types depending on the language. There is no native date, timestamp, or binary type: dates are transmitted as strings like "2026-01-15T10:30:00Z" with no guarantee that producer and consumer parse them identically.

JSON has no concept of a schema: nothing prevents a producer from sending "user_id" in one response and "userId" in the next. A consumer that expects "email" will silently get null if the producer renames it to "emailAddress". There is no compile-time check, no validation step, and no error until the consumer's code encounters a missing field at runtime.

These limitations are acceptable for request-response APIs at moderate volume. They become painful at data pipeline scale.

One record encoded as JSON and as Protobuf byte by byte, with the field names one repeats and the other replaces with tags.

XML: Powerful But Verbose

XML predates JSON and offers features JSON lacks: namespaces, attributes versus elements, mixed content (text interleaved with tags), and schemas (XSD) that validate structure at parse time. Industries like healthcare (HL7/FHIR), finance (FIX/FIXML), and government use XML because their standards were defined before JSON existed and the validation guarantees of XSD matter for compliance.

xml
1<user id="12345">
2  <name>Alice Chen</name>
3  <roles>
4    <role>admin</role>
5    <role>editor</role>
6  </roles>
7  <active>true</active>
8</user>

XML is even more verbose than JSON. Every element has an opening and closing tag. The same user object that is 95 bytes in JSON is 180 bytes in XML. Parsing XML is also slower because the parser must handle namespaces, entity references, CDATA sections, and DTD processing. An XML parser must also handle two entirely different parsing models: DOM (load the entire document into a tree structure in memory) and SAX (stream through the document firing events). JSON parsers are simpler because the format itself is simpler.

For new systems, JSON has almost entirely replaced XML for data interchange. The exceptions are legacy integrations and domains where XSD validation is a regulatory requirement. If you encounter XML in a modern system, it is almost always because the system integrates with a standard that predates JSON (SOAP web services, SAML authentication, SVG graphics) or operates in a regulated industry where XSD validation is mandated by compliance requirements.

Interview Tip

In interviews, if asked why JSON replaced XML for web APIs, focus on developer ergonomics rather than performance. JSON maps directly to JavaScript objects and Python dictionaries. XML requires a parser that produces a DOM tree or SAX events, adding a translation layer. The lower cognitive load of JSON is what drove adoption, not the bandwidth savings.

JSON Variants and Extensions

Several formats attempt to fix JSON's limitations while keeping its readability. JSONL (JSON Lines) puts one JSON object per line, making it streamable and appendable, which standard JSON arrays are not. A consumer can read one line at a time without loading the entire file into memory, which matters for files with millions of records.

MessagePack and CBOR are binary encodings of JSON's data model. They reduce payload size by 20-30% compared to JSON while preserving the same structure (no schema required). However, they sacrifice human readability without gaining the full performance benefits of schema-based formats like Protobuf. They occupy an awkward middle ground: not readable enough for debugging, not efficient enough for high-performance pipelines. They are most useful when you want a drop-in JSON replacement with modest size savings and no schema overhead.

JSON Schema provides optional validation for JSON documents: you can define required fields, types, enums, and value constraints. Unlike XSD for XML, JSON Schema is not integrated into the parser. It is a separate validation step that developers must explicitly invoke. This means it is only as reliable as the team's discipline in applying it. Some teams enforce JSON Schema validation in CI/CD pipelines or API gateways, which provides strong guarantees. Others treat it as optional documentation that gradually drifts from the actual payloads.

Limitations Shared by All Text Formats

Both JSON and XML share fundamental limitations that matter at scale. Field names are transmitted with every record, wasting bandwidth. There is no built-in schema enforcement: a producer can add, remove, or rename fields without warning, and the consumer discovers the change when parsing fails at runtime. There is no compression of repeated values. There is no way to encode binary data (images, encrypted payloads) without base64 encoding, which inflates size by 33%.

These limitations are not bugs. They are the cost of human readability and self-describing structure. When those properties matter more than efficiency (external APIs, configuration files, debugging), text formats are the right choice. When they do not (internal service communication, data pipelines, storage), binary formats exist for a reason.

One common workaround is compressing JSON with gzip before transmission. This reduces the bandwidth cost significantly (gzip compresses JSON well because of the repeated field names), but it does not solve the parsing speed problem. The CPU must still decompress the payload and then parse the text. Compression buys you bandwidth at the cost of CPU cycles. For moderate-traffic APIs, this tradeoff is fine. For high-throughput data pipelines processing millions of messages per second, compression is a band-aid on a fundamental inefficiency.