Practice a Code Migration Interview With a Rollback Plan
Practice a schema migration interview with staged compatibility, concurrent backfill checks and a rollback boundary that respects data state.
TL;DR
- Practice this fictional interview prompt: a service stores a customer's display name in fullname.
- Test it with concurrent updates rather than claiming that batching alone guarantees correctness.
- Choose explicit release gates: compatible writers deployed, reconciliation complete, representative reads verified and an agreed observation period before contraction.
Use a migration where reverting code is not enough
Practice this fictional interview prompt: a service stores a customer's display name in full_name. A new release introduces display_name, but old application instances may remain active during deployment. Explain how you would migrate without losing updates, and what rollback means at each stage.
The exercise is deliberately narrower than redesigning the whole customer system. The interviewer should hear how schema changes, application versions and data movement interact. Begin by clarifying whether the new field has exactly the same meaning. If the new representation changes semantics, a simple copy may be inadequate and reversal may not preserve all information.
State the compatibility requirement first
Assume for this exercise that both fields represent the same string and that the service controls every writer. Old code reads and writes full_name. New code will eventually use display_name. During a rolling deployment, both versions can run at once.
A direct rename makes the old code incompatible. An additive first step leaves the old field available while the application learns the new representation. PostgreSQL's table-modification documentation explains the underlying schema operations; the rollout sequence below is an original application-level exercise, not a claim that an ALTER TABLE command alone makes migration safe.
Write the invariant aloud: every accepted name update must remain available to whichever application version is serving the customer. That gives you a standard against which to judge each phase.
Build a staged plan with one authoritative value
One possible plan is to keep full_name authoritative initially, add the new nullable field, and deploy a bridge version that maintains both fields atomically for new writes. All writers must reach that bridge version before relying on dual-written values. Otherwise, old-only writes can make the new field stale.
| Phase | Read behavior | Write behavior | Main check |
|---|---|---|---|
| Expand | Old field | Old field | Old code still works |
| Bridge | Old field | Both fields atomically | Every writer upgraded |
| Backfill | Old field | Both fields | Historical rows reconciled |
| Switch reads | New field with defined fallback | Both fields | Values agree and errors remain acceptable |
| Contract later | New field | New field | Old code no longer eligible for rollback |
The table is a candidate design, not the only valid one. A database trigger or another synchronization mechanism may fit different constraints. Explain the operational cost and ownership of whichever mechanism you choose.
Make the backfill safe against concurrent updates
A naive job can read an old value, pause, then overwrite a newer value in the new column. Do not describe the backfill as a blind export-and-import. Explain how the database operation preserves the current row state or uses a version check, and how you will handle conflicts.
For this simple same-value exercise, a batched database update from the authoritative column can be designed to avoid stale application snapshots. The exact locking and concurrency behavior depends on the database and statement. Test it with concurrent updates rather than claiming that batching alone guarantees correctness.
Track progress with stable row identifiers and make retries safe. A failed batch should not require guessing which rows were processed. Count remaining missing values and inspect mismatches, but do not treat a zero-null count as proof that every value is correct.
Explain rollback by phase
Before switching reads, reverting application code to the old version may be straightforward because the old column remains authoritative and available. After switching reads but while dual writes continue, reverting reads can still be feasible if you have verified agreement and retained compatible code.
After old writes stop or the old column is removed, the rollback story changes. You may need a forward repair or data reconciliation rather than simply redeploying an old binary. State that boundary explicitly. “We have a rollback button” is not a data-recovery plan.
Our system design frameworks guide can help structure the answer around requirements and failure modes. In this exercise, the most important failure is a version transition that leaves two inconsistent sources of truth.
Rehearse a failure injection and a stopping rule
Ask a practice partner to introduce one event: an old worker remains running, a backfill batch crashes or mismatch counts rise after the read switch. Explain what you would pause, what evidence you would inspect and which operations remain safe. Do not respond to every event by dropping the new column.
Choose explicit release gates: compatible writers deployed, reconciliation complete, representative reads verified and an agreed observation period before contraction. The exercise does not supply universal thresholds; explain how traffic, risk and service expectations would determine them.
Use the mock interview strategy guide to retry the same reasoning with a harder variant, such as splitting one field into two. Your answer is ready when you can explain not just the happy path, but why a particular rollback remains valid at one phase and becomes unsafe at another.