Migrate tags across data, clients and workersLESSON 16.08 · 8 OF 9 IN CHAPTER
PART E / Live migrations
Step 222 of 252
LESSON 16.08 · 8 OF 9 IN CHAPTERBuild brief

Migrate tags across data, clients and workers

Application background

A reading list currently stores tags as words such as databases. The team wants tags to have permanent IDs so their display names can change. A tag might become {id: "tag-7", name: "databases"} while keeping the same identity after a rename.

The change affects more than database rows. Browsers display tags, workers read them, exports include them and optional AI suggestions use them. Some of those consumers will still use the old format while the move is underway.

The data shape being changed

These are proposed old and new representations for the same bookmark:

{"id":41,"tags":["databases"]}
{"id":41,"tagObjects":[{"id":"tag-7","name":"databases"}]}

During migration, the server may need to derive both representations from one authoritative model. An old browser still expects strings, while a new browser can keep tag-7 when its display name changes.

A backfill converts existing records. Compatibility keeps old and new consumers working during that conversion. The migration is unfinished while an active consumer still depends on a retired format.

Your assignment

Deliver: Move tags to stable IDs while old clients and jobs continue working. Build the conversion, compatibility handling and record comparison, then retire the old format only when its remaining users are gone.

Required behavior: One writer authority exists at each phase. Backfill and live changes carry source versions, including deletions. Old clients remain supported through an adapter until a documented retirement point. A progress percentage alone does not prove completion.

The required first milestone is a working local implementation of the behavior above. The numbered implementation steps define the scope. The cloud architecture is a later extension, not something the starter has already provisioned.

Get the code and run the supplied example

The code is in the public junior-to-staff repository. Install Git and Python 3.12+. No AWS account or Python packages are required for this first run. If you already have a checkout, use it and skip cloning.

git clone https://github.com/Soulful-Iris/junior-to-staff.git
cd junior-to-staff
python3 examples/architecture-starts/the_migration_you_actually_finish.py

Supplied file: examples/architecture-starts/the_migration_you_actually_finish.py. You can also read or download the source here (download file, source below).

Read the supplied code · the_migration_you_actually_finish.py
read or download the source here · the_migration_you_actually_finish.py
"""Local mechanism demonstration for the-migration-you-actually-finish. No AWS resources are created."""
target={}
def apply(version,tags,deleted=False):
    if target and version<=target['version']: return 'ignored stale'
    target.update(version=version,tags=tags,deleted=deleted); return 'applied'
for v,tags,deleted in [(5,['tag-db'],False),(6,[],True),(4,['old-text'],False)]:
    print(v,apply(v,tags,deleted))
print('Final target:',target)

This program is a mechanism demonstration: it runs the small scenario in one process and prints the result. It is not an HTTP service, a complete application, or an AWS deployment. A successful run demonstrates this mechanism only. It does not establish the workload or failure guarantees of the application you will build.

Example output from the supplied run:

Generated IDs and timestamps may differ. Compare the state transitions and outcomes.

5 applied
6 applied
4 ignored stale
Final target: {'version': 6, 'tags': [], 'deleted': True}

Run the application you will extend

The reading-list API setup guide gives you a real local HTTP server, SQLite database, save/list/edit requests and controlled title success/timeout behavior. Start it in one terminal and send the documented curl requests from another. Read that setup before following the implementation steps below. The demo above isolates this lesson's mechanism. The server is where you integrate it.

For a first run, start this in terminal 1 from the repository root:

python3 examples/reading-list-starter/app.py --db /tmp/reading-list.sqlite3

In terminal 2, save one bookmark with a controlled title timeout:

curl -i http://127.0.0.1:8080/bookmarks \
  -H 'X-Demo-User: alice' -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com/docs","title_mode":"timeout"}'

Expect 201 Created, a bookmark id and title_status: "timeout". The URL is persisted despite the title failure. This is the supplied baseline. The assignment adds the behavior described above. The lookup is a fixture, so no external website is contacted. For members Bob or Ben in a scenario, use the starter's second demo identity bob. Alice or Ana corresponds to alice.

Work in your own branch or copy examples/reading-list-starter/ to work/the-migration-you-actually-finish/. app.py exists in that directory. Add the modules named below there as you separate HTTP, storage and background work. The server has demo membership, not production authentication.

Local components and state to implement

This table names the records, interfaces or decision inputs for your deliverable. Unless a name is explicitly linked to supplied source above, it is something you create. Implement the local state transitions first, then connect the HTTP, storage or worker boundaries required by the steps.

Record / module Key or interface Responsibility
migration_state phase,checkpoint,writer_epoch Resumable progress and authority.
link_tags_v2 link_id,tag_ids,source_version,deleted Target representation with monotonic apply.
consumer_inventory owner,version,reads,writes,retirement Evidence that old contracts can actually be removed.

Implement the assignment

1. Inventory before changing storage

List old/new API shapes, browser versions, worker payloads, exports and AI inputs. Name the owner of each reader/writer. Add a canonical tag mapping and compatibility serializer while the old path remains the write authority.

2. Implement resumable versioned backfill

Read a consistent source boundary and capture subsequent changes without a gap. Apply source versions conditionally, including tombstones. Persist checkpoints after durable target application so restarting a range is safe.

3. Reconcile and transfer authority

Compare canonical records at an aligned watermark, accounting for deletes and tag normalization. Fence old direct writers, apply the final change prefix, then move routing/epoch to the target. Shadow reads alone do not stop an old worker from writing stale data.

4. Finish retirement and repair

Observe supported-client usage, drain or adapt old queue payloads, update exports and remove old fields only after every required consumer is covered. Name the first target write that old code cannot interpret. After that point use a prepared reverse adapter or forward repair rather than a misleading rollback button.

Demonstrate the completed local result

Action Expected visible result
Run the starting program Version-4 backfill cannot resurrect the version-6 deletion.
Restart mid-range The range replays safely and progress resumes.
Send an old worker payload after cutover The adapter handles it or the old writer is explicitly rejected.

Handoff: In your implementation README, include the start command, one successful operation, the failure case above and the resulting stored state or decision. State which dependencies are simulated. Someone with a fresh checkout should be able to reproduce this without your chat history.

Workload assumptions and capacity decisions

These are constructed exercise assumptions. The stated workload is a design target. The local demonstration does not establish that throughput. Use the estimation constants to check units before choosing capacity.

Input or objective Calculation / consequence
One million links. 500 rows/s backfill assumption About 33 minutes ideal copy time, excluding live changes, indexes and retries.
Live update version 5. Deletion version 6. Late backfill version 4 Target must retain the version-6 tombstone.
Four consumer types Browser, API, worker and AI/tagging input all belong in the compatibility inventory.

Map the local implementation to AWS

Deployment status: local only. Running the supplied command creates no AWS resources and configures no cloud connections. The diagram is a proposed deployment of the completed application. Each box needs either a deployed runtime, a provisioned service or an explicitly external dependency.

Read the diagram by following the arrows from the entry point: application code accepts the request or event, the state owner commits it, and any worker produces the later result. The table ties those roles to code and adapter work. Multiple boxes do not imply multiple Python files already exist.

Migrate tags across data, clients and workers: AWS services, their general roles, and the primary data flow

The backfill worker transports and transforms data. Source versions and writer fencing preserve correctness. Configuration routing is coordination metadata, not a substitute for rejecting stale writers.

Local responsibility Cloud destination and role Implementation still required
Local records and transaction boundary Amazon Aurora PostgreSQL: existing data authority Write PostgreSQL schema/migrations and a database adapter. Configure credentials, connection limits and recovery.
Application or worker process Amazon ECS: backfill and change workers Build a container and task definition. Supply configuration, task roles and graceful shutdown behavior.
Local dictionary, SQLite records or state model Amazon DynamoDB: target tag representation Design partition/sort keys and write a storage adapter with conditional updates or transactions. Python state and SQL are not uploaded as a database.
Local versioned configuration AWS AppConfig: migration phase routing Publish validated configuration versions and consume them with bounded caching and rollback behavior.
Local counters, timestamps and diagnostic output Amazon CloudWatch: migration evidence Emit bounded metrics and logs, build the named operational view and configure retention and access.
Local file, object fixture or exported payload Amazon S3: reconciliation artifacts Implement upload/download and metadata adapters, scoped access, object naming, retention and incomplete-upload cleanup.

Provision resources, then connect the application

Resource or boundary Initial configuration and reason
Workers Bound copy load so live requests retain capacity. Checkpoints are durable and replay-safe.
Routing Enforce writer authority at the write boundary, not only in cached client routing.
Completion Record old consumer usage and schema compatibility. Remove old resources only after the declared retirement conditions.

Use one disposable AWS environment for the cloud exercise. Put the named resources in infra/template.yaml or your existing IaC tool, pass resource IDs through configuration, and scope each runtime role to its own tables, buckets and queues. The diagram is a design to implement. It is not a claim that these resources have been deployed. Record the commands you used to deploy and remove the exercise resources.

For concrete provisioning commands, configuration wiring and cleanup, use the AWS foundation guide. It includes a deployable table/queue/object-storage foundation and explains which application and service adapters you still implement.

A provisioned queue or table does not make the local program use it. Configure resource IDs in the deployed runtime, replace the local adapter, and replay the same successful and failing operation against that runtime. Record the deployed commit and observable result, then remove the disposable resources using your infrastructure tool.

Extend the design after the baseline works

A previously unknown export tool still reads the old table. Add it to the inventory and decide whether to adapt or retire it. Migration completion is about actual consumers, not only the services you remembered initially.

Follow-up scenarios and worked designs

Follow-up 1 · The backfill meets live writes

How do you combine snapshot v1 with live v2 and delete v3 without losing either? Predict the failure before opening the design.

Worked design and implementation

Apply only increasing versions and retain tombstones for the replay horizon. Track checkpoint coverage, source counts/checksums and semantic mismatches. A counter at zero needs a blind-spot analysis, including dynamic consumers and delayed/offline writers.

Work through one record. The snapshot contains item 7 at v1, a live edit produces v2, and deletion produces tombstone v3. Deliver them to the target in the order v2, v3, v1. The target must retain deleted v3 because every apply checks that the incoming version is newer. A tombstone is a versioned deletion record, not an ordinary missing row.

Persist copy checkpoints and stream coverage separately. If change history expires before catch-up, stop and resnapshot the affected range. Deliver the record timeline, a version/value/deletion reconciliation result and the point where you permit cutover.

Revised flow. These are proposed components to implement, not extra services started by the supplied demo.

Diagram: Follow-up 1 · The backfill meets live writes

Follow-up 2 · A team cannot cut over

One team must keep old clients for a quarter. DNS caches last 300 seconds. What can rollback promise?

Worked design and implementation

Preserve old/new data compatibility and choose explicit cohorts. Load-balancer admission, DNS propagation and existing connection drain have different timing. Keep partial gains measurable, but retire duplicated maintenance only after consumers and replay obligations are gone.

Keep compatibility alive for the actual consumer horizon. Route migrated cohorts to the new path while the delayed team keeps an owned adapter. Old writers must still preserve the new system's invariants, or be fenced before target-only semantics begin. A quarter of dual support has an operational cost that belongs in the plan.

Show an old mobile write arriving during the overlap and a rollback after a target-only field is created. DNS with a 300-second TTL does not instantly move cached clients or existing connections. Deliver a compatibility table and distinguish traffic rollback from data repair.

Supplied mechanism practice

These verify specific boundaries. Passing their reference tests does not implement or assess the full project.