Last updated: September 29, 2026
By Kris Drouet, Engineering Executive, in partnership with KORE1
Schema evolution is how an event’s structure changes over time without breaking the services that read it, and doing it without downtime means pairing registry compatibility checks with a known consumer list and an expand, migrate, contract rollout. The registry catches broken shapes. It can’t catch broken meaning, and it has never met your consumers.
The registry said compatible. It was telling the truth.
About eight months after we moved our loan pipelines onto Kafka, a producer team added one value to an enum on the rate-lock event. ACTIVE, EXPIRED, CANCELLED, and now EXTENDED. Product had asked for lock extensions, the change was small, the pull request was clean, and the schema registry accepted the new version without complaint, the same way it had accepted a couple hundred other changes that year. They deployed on a Thursday afternoon. Routine.
By Thursday evening the dead-letter topic had a few hundred messages in it. Every one was an extended lock. Two downstream consumers, a pricing-audit job and a reporting feed that finance leaned on at month-end, were still running classes generated from the old schema. They met a symbol they had never heard of, had no default to fall back on, and refused to deserialize the record. Nothing crashed. The messages went exactly where we had designed failed messages to go. That dead-letter topic was one of the guardrails from the Kafka pub/sub migration that cut our downstream latency 45 percent, and it did its job. Quietly. Which is the only reason nobody outside engineering heard about it.
I still chew on this one. No tool failed. The registry checked the change in the direction it was configured to check, and the change passed. What failed was an assumption that “compatible” meant “safe to ship in any order,” plus the fact that nobody on the producer team could have named those two consumers if you had asked them that morning.
Go-live is where most writing about event-driven architecture stops. In the point-to-point integration decoupling playbook, move five is to put a contract on every new boundary and version it. This is the long tail of that move. What versioning actually looks like in month eight, when the events have real consumers, some of them owned by teams you’ve never met.

What Schema Evolution Means for an Event Stream
Schema evolution for events is the practice of changing an event’s structure (adding, removing, or retyping fields) while producers and consumers upgrade on different days. Old records stay in the log, new records arrive beside them, and every reader has to cope with both until the old ones age out.
If you search the term, most of what comes back is about tables. Databricks, Delta Lake, a flag that lets a new column merge into an existing schema on write. Real problem. Easier one, though. A table usually has one writer and a small set of readers, and if something goes wrong you can rewrite the data.
A stream is different in two ways that matter. You can’t rewrite a log that other teams have already read, and you don’t control when they upgrade. The producer ships on Tuesday, the billing consumer ships three weeks later because its team is in the middle of its own release and has no appetite for an unplanned deploy, and nobody is wrong. The data science team replays the topic from the beginning next quarter to train something, reading records that were written under a schema two versions back. All three are normal. Your versioning strategy has to survive all of them at once, with no maintenance window, because a pub/sub platform doesn’t really have one. We built it that way on purpose.
Compatible With What, Exactly?
Every schema registry asks the same question when a new version arrives. Compatible with which versions, in which direction? The answer is a setting, and most teams never look at it after day one. I didn’t.
Confluent’s documentation on schema evolution and compatibility lays out the modes, and the default is BACKWARD. That default checks a new schema only against the latest one. Not the version before that. Not the one your oldest retained records were written with. Just the latest.
| Mode | What it guarantees | Upgrade first | Where it bites |
|---|---|---|---|
| BACKWARD (default) | A consumer on the new schema can read data written with the previous one | Consumers | Producer ships first and old consumers meet data they can’t read |
| BACKWARD_TRANSITIVE | Same, checked against every earlier version | Consumers | Stricter reviews, which is the point |
| FORWARD | A consumer on the previous schema can read data written with the new one | Producers | Rewinding a new consumer to old data isn’t guaranteed |
| FULL / FULL_TRANSITIVE | Both directions | Either, in any order | Fewer changes pass, so teams are tempted to switch it off |
Read the third column again. That column bit us. Under BACKWARD, the registry is promising that upgraded consumers can read old data. It promises nothing about old consumers reading new data, and the documentation says so plainly: upgrade every consumer before you start producing the new events. Adding an enum symbol passes a backward check. It is also exactly the kind of change an old reader can choke on, and the Apache Avro specification is blunt about why. If the writer’s symbol isn’t in the reader’s enum and the reader has no default, “an error is signalled.” Our producer shipped first. That’s the whole incident, in one table row.
Two more details from the same Confluent page are worth pinning to a wall. For Kafka Streams, only backward compatibility is supported. And for Protobuf, the stated best practice is BACKWARD_TRANSITIVE, because adding new message types isn’t forward compatible. If your topics carry Protobuf and nobody has changed the default, you’re running a looser mode than the vendor’s own documentation recommends.
My own rule is simpler. Transitive, on every topic that anyone might replay, which in practice is every topic. A non-transitive check protects the next deploy. A transitive one protects the replay nobody has scheduled yet.
The Consumer List Nobody Keeps
Ask a producer team who reads their topic. You’ll get a confident answer and it’ll be wrong by at least one. Every time.
Pub/sub decouples producers from consumers on purpose, so the producer genuinely doesn’t need to know who is listening. That’s great for latency and terrible for change management. Both at once. Six months in, the topic you built for underwriting has a reporting feed on it, a fraud model, a compliance archive, and a script somebody in finance ops wired up with a service account and never told anyone about. A topic with no documented consumers doesn’t have zero consumers. It has unknown ones.
So before we changed another schema, we built the list, and the broker was the source of truth, not a wiki. The steps were dull.
- Pull every consumer group that has committed offsets on the topic in the last 30 days. That’s your real population, whatever the architecture diagram says.
- Match each group ID to a named team and a human who will answer a message. Groups nobody claims go on a list with a date by which they get claimed or their access gets reviewed.
- Record what each consumer actually reads. Most read four fields out of twenty, and that’s useful, because a change to field seventeen can’t hurt them.
- Put the list next to the schema, in the same repository, so a pull request that changes one shows the other.
That last step carries more weight than it looks like it should. Once the consumer list sits beside the schema file, a reviewer can see who a change touches before it merges instead of after. It turned schema review from a type-checking exercise into a conversation with the people downstream, which is what Ian Robinson was arguing for years ago in consumer-driven contracts. Consumers state what they depend on. Producers can’t break it without knowing. You don’t need a contract-testing framework on day one. A list with names on it gets you most of the way there. Start there.
It took about a week for our twelve busiest topics. We found three consumers nobody could name. One turned out to be a load test someone forgot to delete in March. Still running.

Expand, Migrate, Contract
This is the rollout that lets a schema change ship without a maintenance window. It isn’t clever. It’s just slow on purpose, and each step is safe on its own.
- Expand. Add the new field alongside the old one, optional, with a default. In Avro, a reader field with no default, missing from the writer, is an error, so the default is not a nicety. In Protobuf, adding a field is wire-safe.
- Upgrade the consumers on the list to read the new field if it’s present and fall back to the old one if it isn’t. Order matters here, and your compatibility mode tells you which side goes first.
- Switch producers to write both fields. Now every new record carries the old shape and the new one.
- Wait. Longer than feels necessary.
- Contract. Stop writing the old field, then remove it from the schema. In Protobuf, reserve the number and the name so nobody reuses them.
Step four is where teams lose patience, usually right around the moment a product manager asks why the old field is still in the payload three weeks later, so it’s worth being precise about how long “wait” actually is. The old field can’t come out until three things are true. Every consumer on the list is reading the new field in production, not in a branch. The topic’s retention has rolled past the last record that carried only the old shape, which on a Kafka topic with the default retention is seven days and on a compacted topic may be never, because compaction keeps the latest record for every key indefinitely. And nobody has a replay planned against the older history. Ask around. Compacted topics catch people constantly. The customer record written in 2024 is still sitting there, in the old shape, and it will be read again the next time someone rebuilds a state store.
The Protobuf reserve step deserves one more sentence, because the failure is nasty. The Protocol Buffers language guide lists what happens when a deleted field number gets reused: parse errors if you’re lucky, and at the bad end, data corruption and leaked personal data. Someone renumbers fields to make the file look tidy. Old records now decode into the wrong field. On a lending platform that’s a borrower’s income landing somewhere it was never meant to be. Reserve the numbers. It’s one line.
When the Change Really Is Breaking
Some changes can’t be expanded and contracted. The key changes. The event splits into two. The meaning of the thing itself moves. For those, stop trying to version inside the same stream. Start a new one.
Publish a new event type on a new topic, and run both. The CloudEvents specification puts the principle in one line for its dataschema attribute: incompatible changes to the schema “SHOULD be reflected by a different URI.” A breaking change gets a new identity. It doesn’t get to pretend to be version 7 of the old one.
Mechanically, you have two ways to keep the old consumers fed while they move. The producer can dual-publish to both topics for a migration window, which is simple and doubles your chances of the two drifting. They will drift. Or a small translator service consumes the new topic and republishes the old shape, which keeps the producer clean and gives you one obvious thing to delete at the end. I lean toward the translator. It puts the migration cost in a box with an owner and an end date, and when the last consumer moves you turn it off and the old topic goes quiet.
Retry-safety matters twice here, since a translator that restarts mid-batch will republish some events. Everything I wrote about idempotent integration design applies to the bridge as much as the endpoint.
The Break No Registry Can See
A field called amount used to exclude the appraisal fee. Then, one release, it included it. Same name. Same type. Same schema version. Every compatibility check on earth passes that change, and a reconciliation job downstream was off by a few hundred dollars per loan for eleven days before anyone in finance noticed the pattern and walked it back to us.
Registries check shape. They don’t check meaning. Units change from dollars to cents. Timestamps move from local time to UTC. A null that used to mean “unknown” starts meaning “not applicable.” All of these are breaking changes that look compatible, and the only defense is a rule plus a habit. The rule is that a field never changes meaning under the same name. New meaning, new field, and the old one goes through expand, migrate, contract like anything else. The habit is a line in the pull request template asking, in plain words, whether any existing field now means something different. It sounds too simple to work. It catches more than the registry does. By a lot.
Readers have a part to play too. Martin Fowler’s Tolerant Reader pattern says to read only what you need and ignore the rest, which protects a consumer from most additive changes, since a field you never asked for can’t hurt a reader that never looks for it. Watch for one trap under the hood of a lot of Java consumers. A plain Jackson ObjectMapper fails on unknown JSON properties by default, so a consumer deserializing JSON events into strict classes will reject a perfectly harmless new field. Turn that off deliberately, or find out on a Thursday.

Who Owns the Contract After Go-Live
Full disclosure before this section. I write these in partnership with KORE1, and KORE1 earns a fee when a company hires through it. Weigh what follows with that in mind.
The migration gets a team. The year after it usually gets nobody. Schemas pile up, the compatibility setting stays at whatever it was on day one, and topic ownership quietly defaults to “the platform team,” which means nobody in particular. Most of the incidents I’ve seen on mature Kafka platforms, the ones that ate a weekend or showed up in a finance review, weren’t broker failures or disk failures or anything a vendor could have fixed for us. They were ownership failures that happened to show up in a deserializer.
Somebody has to own the event model as a product. Set the compatibility mode per topic. Keep the consumer lists honest. Say no to a field that quietly changed meaning. Retire the v1 topic when it’s time instead of letting it run forever because turning it off feels risky. That person is usually a senior integration or streaming architect, and they’re hard to find, because the skill is part distributed systems and part diplomacy with five teams who all want their change shipped this sprint. If you’re hiring for it, this is what the event-driven Kafka architect search looks like in practice. KORE1 has placed engineers in more than 30 U.S. metros since 2005, averages 17 days to fill IT roles, and keeps 92 percent of placements in the seat at twelve months. Retention is the number I’d weigh. The person who knows why topic 14 has a translator in front of it is the person you can least afford to lose. Plenty of teams bring that person in on contract staffing for the first year to set the rules, then hire permanently once the platform has settled.
Questions Teams Ask the Week After Go-Live
Can Adding a Field Ever Break a Consumer?
It can, when a consumer upgrades to an Avro schema whose new field has no default and then reads older records, or when it deserializes JSON into strict classes that reject unknown properties.
Both are fixable in an afternoon. Neither shows up in a unit test that only uses the new schema. Both hide well.
Version Number in the Payload or in the Topic Name?
Neither, for compatible changes, because the serializer already stamps every record with its registry schema ID and a version field in the payload just duplicates that.
Put a version in the topic or event type name only when the change is breaking and you’re standing up a new stream. Something like loan.rate-lock.v2 is fine. A field called schemaVersion that consumers branch on is how you end up with six code paths in every reader and no plan to delete any of them.
How Long Do We Keep Supporting an Old Schema Version?
Keep it until the slowest consumer on your list has moved and topic retention has rolled past the last old-only record, and on compacted topics that second condition may never arrive by itself.
On those topics, plan a one-time rewrite of the old records or keep the reader tolerant indefinitely. Pick one on purpose. Most teams pick neither.
Is JSON Without a Registry Good Enough?
Plain JSON is fine for two services owned by one team, and not for anything with consumers you don’t control, because nothing stops a producer from shipping a shape nobody agreed to.
JSON Schema in a registry closes most of that gap. Plain JSON on a topic with eight consumers is a contract enforced by memory and good intentions, and memory is the first thing to go when the engineer who wrote the producer moves to another team. I’ve run that setup. I wouldn’t again. Show me the data on how many unannounced payload changes that topic absorbed last quarter, and I’ll show you a team that’s been getting lucky.
Can We Rename a Field Without a Breaking Change?
Avro lets a reader map an old field name through aliases, but a rename is still safest handled as a new field plus expand, migrate, contract.
Protobuf’s binary format doesn’t carry field names at all, so a rename is invisible on the wire and very visible to anyone reading the JSON encoding. Reserve the old name when you retire it.
Who Should Approve a Schema Change?
The producing team owns the decision, and every consumer on the list gets a veto window before it merges, usually two or three working days.
A central architecture board reviewing every field is too slow and learns to rubber-stamp. I’ve watched it happen. No review at all is how the enum story at the top of this post happens. The consumer list sits between the two. It tells you exactly whose sign-off matters for this change and whose doesn’t.
Version for the Consumer You Haven’t Met
The registry will tell you whether a schema is compatible. It won’t tell you who is reading, what they think a field means, or whether anyone is about to replay two years of history into a new model on Monday morning. That’s still the job, and it doesn’t end at go-live. It starts there.
If your streaming platform has outgrown the people who built it, KORE1 can help you find the architect who owns what comes next. Talk to a KORE1 recruiter about the role. And if you want to compare compatibility settings or consumer-list horror stories, connect with me on LinkedIn. I’m always up for that conversation.

