Migrate live data while preserving writes, deletions and recovery
A reading-list service stores tags as strings and is moving to stable tag IDs. Users keep editing and deleting bookmarks while a background process copies old rows. The new store can receive a live edit before the older copied row arrives.
You will define which store owns writes during each phase, how copied and live changes are ordered, and what permits the final switch. A backfill is the bulk copy of existing data. It does not include all changes that happen after the copy begins.
Project connection · feeds Reading-list stage 5: evolve the running application
Follow one bookmark through an out-of-order copy
| Arrival at the target | Version | Correct target state |
|---|---|---|
| Live edit changes tags | 2 | Store v2 |
| Live deletion arrives | 3 | Retain a deletion marker at v3 |
| Delayed snapshot row arrives | 1 | Ignore v1 and remain deleted at v3 |
The deletion marker is a tombstone. It records that the absence is newer than the copied value. Deleting the target row without retaining that fact would let the delayed copy recreate it. Your apply operation must compare versions atomically. A final row-count match cannot detect two different identities with the same total count.
The migration lab supplies the local fixture. Use it to observe copy, live update, deletion and stale replay before designing a cloud cutover.
Reason through the changed situation
“We are replacing an event store while reads and writes continue. The old write succeeds and the new write times out. What do we return, which copy is authoritative, and how will we find and repair the difference?”
A migration must preserve a defined contract while versions and stores coexist. Two writes are not one atomic operation merely because they share a function.
| Event | Expected policy to specify |
|---|---|
Old store accepts event e7; new store is unavailable |
Durable record of replication work |
Change v8 arrives before replayed v7 |
Older replay cannot overwrite newer state |
| A record is deleted during backfill | Tombstone or equivalent deletion semantics survive |
| Read cutover fails | Reversal has a defined data and routing boundary |
Establish authority before cutover
- Choose an authoritative write path and a durable change-capture boundary. Sampling detects some divergence; it does not itself repair missing writes.
- Backfill from a defined snapshot or watermark while replaying changes with version and deletion rules. Make both stages restartable.
- Compare invariants and representative workloads. Define repair, catch-up, and the evidence required before switching reads.
- Stop old writes and retire compatibility only after rollback requirements expire. Track intermediate benefits separately from retirement savings.
“The new store already accepts writes. Can a DNS change instantly restore the old system?” No: account for new data, resolver TTLs, and established connections before claiming a rollback bound.
Senior depth makes partial failure and replay safe. Lead depth coordinates consumer versions, owners, staged value, stop conditions, and actual retirement.
The mechanism: de-risk, enable, finish
De-risk the uncertain requirement early, not the riskiest production cohort. A replay, shadow read, disposable test environment or limited pilot can expose a hard compatibility requirement without putting the largest customer at risk. Choose initial live exposure for representativeness, observability, reversibility and bounded impact. An easy pilot tests deployment mechanics; it does not prove that every difficult consumer can migrate.
| Question | Evidence to collect before expanding |
|---|---|
| Can the new store represent the hardest record? | Replayed fixture with keys, versions and tombstones preserved |
| Does the rollout machinery work? | Small live cohort, monitored errors and an exercised stop control |
| What happens during a partial write? | Durable replication intent plus restart/reconciliation trace |
| Can old clients read new writes? | Compatibility tests across supported versions |
Enable adoption. Supply a supported adapter, migration command and progress report. Make retries safe. A checkpoint means “these inputs are durably applied,” not merely “the loop visited them.” Keep a bridge for consumers that cannot yet move, with an owner and an expiry/review condition.
Finish deliberately. Prevent new unapproved use of the old path, count remaining writers and readers, observe actual traffic, and remove the old system only after the rollback window and retention requirements permit it. A static search can miss dynamic callers; zero matches is not proof of zero traffic.
Separate intermediate gains from retirement savings
Coexistence can cost more while already helping users. These are constructed units per month, not provider prices:
| Stage | Old-system cost | New-system cost | Total | Migrated cohort latency |
|---|---|---|---|---|
| Before | 100 | 0 | 100 | 600 ms |
| Partial rollout | 100 | 30 | 130 | 180 ms |
| Retired old path | 0 | 80 | 80 | 180 ms |
The partial rollout has a real latency benefit and a real cost penalty. Retirement removes the old-system cost; it is not the first moment any benefit exists. Track migration labor, ongoing support, incidents and capacity separately. Do not turn the drawing into a universal cost curve or an “80% abandoned” statistic.
A rollback is four separate questions
| Boundary | What reversal can and cannot do |
|---|---|
| Code | Revert a change only while the old version understands current state |
| Traffic | Change routing; account for cached routes and existing connections |
| Data | Restore retained values or replay a trustworthy log; recreating a deleted column does not restore its contents |
| Business effect | Reconcile or compensate an already executed payment; disabling a flag does not undo it |
Keep a defined authority while old and new versions coexist. If new-only writes have started, redirecting reads to an unreconciled old store loses visible state. A safe reversal needs reverse replication, a write pause with reconciliation, or an explicitly narrowed forward-recovery policy.
Practice on P5
Choose one load-bearing change in the continuing project. Record the invariant, hardest compatibility case, first low-impact live cohort, replication/repair mechanism, stop condition and retirement owner. Include a late old writer and a delete during backfill in the test data.
Ask an assistant to find a counterexample to that plan, then check it against the actual contracts. Do not require an objection when the plan is sound. A useful review may approve it unchanged and record why.
Acceptance: restart a partial backfill without corrupting newer writes; reconcile values, not just row counts; exercise the supported reversal; and show that retiring the old path leaves no supported caller dependent on it. An intentionally irreversible step needs a tested restore or forward-fix policy.
Words to keep: de-risk reduces uncertainty; coexistence runs both paths; cutover changes authority or routing; retirement removes the old obligation.
Database foundations · Recovery lab · Technical decisions and engineering effectiveness
Draw it from memory · Draw coexistence before drawing cutover
Circle every writer that still targets the old representation. What evidence permits retirement?