Migrate tenant data with a resumable backfillLESSON 16.05 · 5 OF 9 IN CHAPTER
PART E / Live migrations
Step 219 of 252
LESSON 16.05 · 5 OF 9 IN CHAPTERTry it, then open the solution

Migrate tenant data with a resumable backfill

Application background

A SaaS product stores several companies' records in an old database layout. The team needs a new layout, but customers must keep using the application during the move. Some still use older mobile apps that cannot change immediately.

You will move one company's data at a time. Copying existing records is called a backfill. While that copy runs, new edits still arrive, so the system must know which database is allowed to accept authoritative writes.

Example walkthrough

01 · Try this input

Input / starting state
Start moving tenant Acme
Expected result
Record where its copy has reached.

02 · Try this input

Input / starting state
The copy process stops and restarts
Expected result
Continue from saved progress instead of starting blindly.

03 · Try this input

Input / starting state
Acme is ready to move to the new store
Expected result
Compare the records, then explicitly switch which store accepts writes.

A cutover is that controlled switch. Copying data and changing write authority are different operations, and the old client contract still needs to work through both.

Your assignment

Deliver: Build a copy process that can resume for each company, compare the old and new records, and control when the new store becomes responsible for writes.

Required behavior: Each tenant has a recorded migration state and one write authority. Backfill, updates and deletions carry monotonically comparable source versions. Reads never resurrect an older row after a newer deletion.

The required first milestone is a working local implementation of the behavior above. The numbered implementation steps define the scope. The cloud architecture is a later extension, not something the starter has already provisioned.

Get the code and run the supplied example

The code is in the public junior-to-staff repository. Install Git and Python 3.12+. No AWS account or Python packages are required for this first run. If you already have a checkout, use it and skip cloning.

git clone https://github.com/Soulful-Iris/junior-to-staff.git
cd junior-to-staff
python3 examples/architecture-starts/multi_tenant_migration.py

Supplied file: examples/architecture-starts/multi_tenant_migration.py. You can also read or download the source here (download file, source below).

Read the supplied code · multi_tenant_migration.py
read or download the source here · multi_tenant_migration.py
"""Local mechanism demonstration for multi-tenant-migration. No AWS resources are created."""
target={}
def apply(key,version,value):
    old=target.get(key)
    if old and old['version']>=version: return 'ignored stale/duplicate'
    target[key]={'version':version,'value':value}; return 'applied'
for v,value in [(5,'live edit'),(6,None),(4,'old backfill')]:
    print(v,apply('acme:7',v,value))
print('Target:',target,'visible:',target['acme:7']['value'])

This program is a mechanism demonstration: it runs the small scenario in one process and prints the result. It is not an HTTP service, a complete application, or an AWS deployment. A successful run demonstrates this mechanism only. It does not establish the workload or failure guarantees of the application you will build.

Example output from the supplied run:

Generated IDs and timestamps may differ. Compare the state transitions and outcomes.

5 applied
6 applied
4 ignored stale/duplicate
Target: {'acme:7': {'version': 6, 'value': None}} visible: None

Set up your implementation workspace

Create work/multi-tenant-migration/ in your checkout (or use a separate repository). Copy the supplied mechanism into that directory as mechanism.py, then extract its state transitions into functions you can call from your implementation. The record and module names below describe what you must implement. They are not a promise that files with those names already exist. Keep a README.md beside your implementation with its exact run commands and observed results.

Local components and state to implement

This table names the records, interfaces or decision inputs for your deliverable. Unless a name is explicitly linked to supplied source above, it is something you create. Implement the local state transitions first, then connect the HTTP, storage or worker boundaries required by the steps.

Record / module Key or interface Responsibility
tenant_migration tenant,phase,source_watermark,writer_epoch Controls routing and cutover eligibility.
target_records tenant,id,source_version,deleted Conditional apply prevents stale backfill and update replay.
reconciliation tenant,range,count,canonical_hash Evidence that copied data agrees at a defined watermark.

Implement the assignment

1. Make old and new contracts coexist

Write down the old client request/response shape and the target schema. Add an adapter that preserves old behavior while storing the new representation. Record the last client version requiring the adapter. Do not remove fields merely because the new UI stopped using them.

2. Backfill through the same versioned apply path

Take a consistent snapshot boundary and start change capture without a gap. Persist range checkpoints only after durable target writes. Use a conditional source-version comparison for snapshot rows, live changes and tombstones. Retry a range safely after process death.

3. Reconcile before changing reads

Compare canonical records or range hashes at an aligned source watermark. Account for deletes and null/default conversions. Shadow reads are useful only when you distinguish replication lag from transformation errors. Sample by tenant size and unusual schema values.

4. Transfer write authority once

Drain or fence old writers, apply the final change watermark, then advance the tenant’s writer epoch and routing state. A DNS change alone is not a write fence. Before new-format writes, document whether rollback is still possible or requires reverse transformation.

Demonstrate the completed local result

01 · Try this input

Input / starting state
Run the starting program
Expected result
Version 6 deletion remains after late version 4 backfill.

02 · Try this input

Input / starting state
Restart halfway through a range
Expected result
The same range replays without duplicate logical rows.

03 · Try this input

Input / starting state
Send an old-client write after cutover
Expected result
The compatibility adapter accepts the supported contract. The fenced old database rejects direct writes.

Handoff: In your implementation README, include the start command, one successful operation, the failure case above and the resulting stored state or decision. State which dependencies are simulated. Someone with a fresh checkout should be able to reproduce this without your chat history.

Workload assumptions and capacity decisions

These are constructed exercise assumptions. The stated workload is a design target. The local demonstration does not establish that throughput. Use the estimation constants to check units before choosing capacity.

Input or objective Calculation / consequence
10 TB at 100 MB/s effective copy 10,000,000 MB / 100 = 100,000 s ≈ 27.8 hours ideal. Retries, indexes and live changes extend this.
Security deadline: four weeks A full client replacement cannot fit a 12-week compatibility window. Add an adapter or narrow the deadline scope.
Live writes: 1,000/s assumption A one-hour CDC pause adds 3.6 million changes. Copy throughput alone does not prove catch-up.

Map the local implementation to AWS

Deployment status: local only. Running the supplied command creates no AWS resources and configures no cloud connections. The diagram is a proposed deployment of the completed application. Each box needs either a deployed runtime, a provisioned service or an explicitly external dependency.

Read the diagram by following the arrows from the entry point: application code accepts the request or event, the state owner commits it, and any worker produces the later result. The table ties those roles to code and adapter work. Multiple boxes do not imply multiple Python files already exist.

Migrate tenant data with a resumable backfill: AWS services, their general roles, and the primary data flow

DMS can move rows and changes, but application compatibility, semantic transformations and write fencing remain your responsibility. The design uses per-tenant cutover so one problematic customer does not force a global switch.

Local responsibility Cloud destination and role Implementation still required
Local records and transaction boundary Amazon Aurora PostgreSQL: existing write authority Write PostgreSQL schema/migrations and a database adapter. Configure credentials, connection limits and recovery.
Local snapshot/change-copy input AWS DMS: change transport Configure source/target replication and observe copy positions. Implement application cutover and reconciliation separately.
Application or worker process Amazon ECS: versioned apply workers Build a container and task definition. Supply configuration, task roles and graceful shutdown behavior.
Local records and transaction boundary Amazon Aurora PostgreSQL: target authority Write PostgreSQL schema/migrations and a database adapter. Configure credentials, connection limits and recovery.
Local versioned configuration AWS AppConfig: tenant routing configuration Publish validated configuration versions and consume them with bounded caching and rollback behavior.
Local counters, timestamps and diagnostic output Amazon CloudWatch: migration operations Emit bounded metrics and logs, build the named operational view and configure retention and access.

Provision resources, then connect the application

Resource or boundary Initial configuration and reason
AWS DMS Use only if the source/target pair and required CDC semantics are supported. Record snapshot/CDC start boundary and task lag.
Aurora source and target Separate credentials and writer roles. Preserve tombstones and source versions through transformation.
ECS migration workers Bound copy concurrency to protect live traffic. Checkpoint ranges and expose per-tenant progress/lag.
AppConfig routing Application consumes tenant placement changes. Epoch enforcement occurs at the write authority, not only in cached routing.

Use one disposable AWS environment for the cloud exercise. Put the named resources in infra/template.yaml or your existing IaC tool, pass resource IDs through configuration, and scope each runtime role to its own tables, buckets and queues. The diagram is a design to implement. It is not a claim that these resources have been deployed. Record the commands you used to deploy and remove the exercise resources.

For concrete provisioning commands, configuration wiring and cleanup, use the AWS foundation guide. It includes a deployable table/queue/object-storage foundation and explains which application and service adapters you still implement.

A provisioned queue or table does not make the local program use it. Configure resource IDs in the deployed runtime, replace the local adapter, and replay the same successful and failing operation against that runtime. Record the deployed commit and observable result, then remove the disposable resources using your infrastructure tool.

Extend the design after the baseline works

Worked follow-up: Cross the point where an old schema cannot represent new writes

The new tag model permits color and organization ownership, while the old model stores only strings. Sending traffic back to old code can silently discard meaning even when both databases are healthy.

Starting design Changed requirement
Old and new readers can interpret every accepted write. The target permits a value that the old schema cannot encode.

Revised architecture. Follow the changed responsibility and failure path below. This is a design to implement. The supplied local example does not provision these components.

Diagram: Worked follow-up: Cross the point where an old schema cannot represent new writes

What to implement. Use explicit migration phases: copy, catch up, compatible serving, then target-only writes. Before that last phase, either implement a lossless reverse representation or declare forward repair as the recovery strategy. Fence old writers and retain versioned change records. Keep a per-tenant cutover record so one tenant's target-only write does not accidentally authorize all tenants to cross the boundary. AWS routing changes select servers, while the database migration controller owns data readiness.

Walk through the result. Create tag {name: urgent, color: red} after cutover. Attempt the documented rollback. It must either preserve color through the compatibility adapter or refuse and invoke forward repair. Supply a phase table with permitted writers and recovery action for each phase.

The target accepts data the old schema cannot represent. Mark that first write as an explicit rollback boundary. Design a forward repair path and explain why flipping traffic back would lose meaning even if every server is healthy.

Additional design cases, alternatives and original source notes

All prompts here are constructed practice, without company attribution.

THE PROBLEM

Candidate opening: “Security needs tenant isolation in four weeks. Mobile clients will write the old schema for twelve weeks. Copy live data without losing edits or resurrecting deletes, and choose a rollback boundary the teams can use.”

01 · Try this input

Input / starting state
Old v2 commits. Target write fails
Expected result
Durable replay repairs target

Scope: Independent dual writes are not atomic

02 · Try this input

Input / starting state
Live delete v2 followed by backfill v1
Expected result
Tombstone v2 survives

Scope: Versioned apply, explicit delete retention

03 · Try this input

Input / starting state
New-only write after cutover
Expected result
Rollback waits for reverse repair

Scope: Routing alone cannot restore data compatibility

Prerequisites: migration concepts and transaction lab. Start with the runnable migration fixture. Deliver compatibility/authority first, then snapshot plus replay, then full reconciliation and a rehearsed writer/reader rollback.

Name source authority and the durable change record before designing traffic percentages. Enforce versions and tombstones at apply, and advance checkpoints only after durable writes. Then separate reader routing from writer authority.

Worked approach and AWS mapping — open after your attempt

Prompt: move from a shared relational schema to tenant-isolated storage while teams release independently. Define why: contractual isolation, noisy neighbors, or operating limits. “New technology” alone is not a benefit. Assume 10 TB to backfill and a 100 MB/s safe copy budget: about 28 hours of ideal transfer, before verification, ongoing writes, retries, and throttling.

Choose a routing registry with tenant migration state. Expand client contracts first. Capture changes with an outbox or CDC. Backfill a consistent baseline and apply ordered updates. Shadow reads compare meaningful values. Move a small tenant only when reconciliation and latency pass. Keep rollback routing and change capture until a stated point of no return.

AWS mapping: RDS source, a target selected by access pattern, DMS/CDC where the supported source/target behavior fits, S3 for checkpoints/export artifacts, CloudWatch for lag and mismatch rate. Validate CDC ordering, schema-change handling, and transaction boundaries before committing to the mechanism.

Decision memo: source and target owners. Acceptance metric. Customer cohorts. Capacity/cost budget. Rollback trigger. Data repair procedure. Old-path retirement owner. Failure drill: change a record during backfill, crash the worker, then resume. Row-count equality alone does not prove correctness.

Junior: explain old/new compatibility. Senior: implement resumable copy and reconciliation. Staff: negotiate sequencing, avoid two years of dual operation, and identify what would make you cancel the migration.

Follow-up: region failover during migration

Senior: run partial-write, stale-backfill, delete, replay and rollback tests. Lead follow-up: negotiate the mobile/retention/capacity conflict in the three-team memo, then explain why DNS TTL and connection drain prevent an instantaneous universal rollback. Assessor checks include zero value/version/tombstone mismatches, old-writer fencing, a point of no return, and partial benefits versus retirement-only benefits.

Design route · Practice rubric