Collect news feeds with freshness and deduplication
Application background
A news reader regularly downloads lists of new articles from publishers. It stores article links, groups repeated stories and builds the feed a reader sees. Visiting each publisher on a schedule is called polling.
Some publishers update quickly, some stop responding and some ask the reader to slow down. An empty download does not always mean there is no news. Readers need to know when each source was last checked successfully.
Example walkthrough
01 · Try this input
- Input / starting state
- Publisher A returns a new article
- Expected result
- Store it and make it eligible for the reader's feed.
02 · Try this input
- Input / starting state
- Another feed includes the same article
- Expected result
- Recognize the duplicate rather than displaying another copy.
03 · Try this input
- Input / starting state
- Publisher A has failed for an hour
- Expected result
- Show the last successful check time instead of claiming the feed is current.
Freshness describes how recent the information is. Track it separately from whether your own feed endpoint can return a response.
Your assignment
Deliver: Build publisher polling, duplicate grouping and a reader feed. Show when each source was last checked successfully and preserve that evidence during publisher failures.
Required behavior: Ingest RSS/Atom updates, retain source attribution and publish normalized stories with stable IDs. Freshness is measured per reachable source. A publisher outage must not be presented as an empty successful feed.
The required first milestone is a working local implementation of the behavior above. The numbered implementation steps define the scope. The cloud architecture is a later extension, not something the starter has already provisioned.
Get the code and run the supplied example
The code is in the public junior-to-staff repository. Install Git and Python 3.12+. No AWS account or Python packages are required for this first run. If you already have a checkout, use it and skip cloning.
git clone https://github.com/Soulful-Iris/junior-to-staff.git
cd junior-to-staff
python3 examples/architecture-starts/news_aggregator.py
Supplied file: examples/architecture-starts/news_aggregator.py. You can also read or download the source here (download file, source below).
Read the supplied code · news_aggregator.py
"""Local mechanism demonstration for news-aggregator. No AWS resources are created."""
articles={}; etag='v1'
def ingest(source,item,url,title):
key=(source,item); articles[key]={'url':url,'title':title}
return 'upserted '+str(key)
print(ingest('publisher-a','42','https://example.invalid/story','First title'))
print(ingest('publisher-a','42','https://example.invalid/story','Corrected title'))
print('304 response: retain existing articles',articles)
This program is a mechanism demonstration: it runs the small scenario in one process and prints the result. It is not an HTTP service, a complete application, or an AWS deployment. A successful run demonstrates this mechanism only. It does not establish the workload or failure guarantees of the application you will build.
Example output from the supplied run:
Generated IDs and timestamps may differ. Compare the state transitions and outcomes.
upserted ('publisher-a', '42')
upserted ('publisher-a', '42')
304 response: retain existing articles {('publisher-a', '42'): {'url': 'https://example.invalid/story', 'title': 'Corrected title'}}
Set up your implementation workspace
Create work/news-aggregator/ in your checkout (or use a separate repository). Copy the supplied mechanism into that directory as mechanism.py, then extract its state transitions into functions you can call from your implementation. The record and module names below describe what you must implement. They are not a promise that files with those names already exist. Keep a README.md beside your implementation with its exact run commands and observed results.
Local components and state to implement
This table names the records, interfaces or decision inputs for your deliverable. Unless a name is explicitly linked to supplied source above, it is something you create. Implement the local state transitions first, then connect the HTTP, storage or worker boundaries required by the steps.
| Record / module | Key or interface | Responsibility |
|---|---|---|
| sources | feed_id,url,etag,last_modified,next_poll | Poll state and publisher-specific backoff. |
| articles | source_id,source_item_id,canonical_url,revision | Provenance-preserving normalized records. |
| clusters | cluster_id,article_ids,merge_version | Reversible grouping. Not destructive deduplication. |
Implement the assignment
1. Poll one source correctly
Send If-None-Match/If-Modified-Since when supported. Treat 304 as successful unchanged content, not deletion. Bound response size and parsing time, disable unsafe XML entity expansion, and preserve the original feed/item identity.
2. Coordinate host load
Schedule next_poll by source freshness and host allowance. Respect bounded Retry-After and back off failing publishers. Apply outbound destination validation to feed URLs and redirects. A submitted feed must not become an internal-network fetch endpoint.
3. Normalize without losing attribution
Keep publisher item IDs, observed canonical URLs, published/updated times and retrieved time separately. Group likely duplicates while retaining each source record. A corrected headline updates a revision. A mistaken cluster merge must be reversible.
4. Serve an explicit read model
Index published stories and serve bounded pages with cache headers. Show source freshness and partial ingestion status in the editor view. A feed disappearing should trigger a source incident, not mass deletion of historical articles.
Demonstrate the completed local result
01 · Try this input
- Input / starting state
- Run the starting program
- Expected result
- A corrected title replaces one source item. 304 retains the existing record.
02 · Try this input
- Input / starting state
- Return 429 from one publisher
- Expected result
- That host backs off while unrelated hosts continue.
03 · Try this input
- Input / starting state
- Merge two stories incorrectly
- Expected result
- Editors can split the cluster without losing either source article.
Handoff: In your implementation README, include the start command, one successful operation, the failure case above and the resulting stored state or decision. State which dependencies are simulated. Someone with a fresh checkout should be able to reproduce this without your chat history.
Workload assumptions and capacity decisions
These are constructed exercise assumptions. The stated workload is a design target. The local demonstration does not establish that throughput. Use the estimation constants to check units before choosing capacity.
| Input or objective | Calculation / consequence |
|---|---|
| 50,000 feeds. One poll/minute baseline | About 833 requests/s before retries. Coordinate by host as well as feed. |
| 20 million daily readers | Reader traffic belongs behind a cacheable read model, separate from outbound polling capacity. |
| Average feed response 50 KiB assumption | About 41 MiB/s if every poll returns full content. Conditional requests can materially reduce transfer. |
Map the local implementation to AWS
Deployment status: local only. Running the supplied command creates no AWS resources and configures no cloud connections. The diagram is a proposed deployment of the completed application. Each box needs either a deployed runtime, a provisioned service or an explicitly external dependency.
Read the diagram by following the arrows from the entry point: application code accepts the request or event, the state owner commits it, and any worker produces the later result. The table ties those roles to code and adapter work. Multiple boxes do not imply multiple Python files already exist.
Polling and reader delivery have different scaling patterns. Conditional requests save transfer, while source/item identity preserves corrections without manufacturing duplicate stories.
| Local responsibility | Cloud destination and role | Implementation still required |
|---|---|---|
| Local event dispatch | Amazon EventBridge: polling wake-up schedule | Define event rules/targets and delivery failure handling. Persist logical event/run identity in the application. |
| Application or worker process | Amazon ECS: feed fetch workers | Build a container and task definition. Supply configuration, task roles and graceful shutdown behavior. |
| Local dictionary, SQLite records or state model | Amazon DynamoDB: source and article state | Design partition/sort keys and write a storage adapter with conditional updates or transactions. Python state and SQL are not uploaded as a database. |
| Local file, object fixture or exported payload | Amazon S3: original feed snapshots | Implement upload/download and metadata adapters, scoped access, object naming, retention and incomplete-upload cleanup. |
| Local derived search records | Amazon OpenSearch Service: story read index | Implement indexing, updates/deletions and queries. Recheck current authorization before returning sensitive results. |
| Local static/media delivery path | Amazon CloudFront: public read delivery | Configure an origin, cache policy and private-content access. Distinguish cached bytes from current authorization. |
Provision resources, then connect the application
| Resource or boundary | Initial configuration and reason |
|---|---|
| Fetch workers | Fixed outbound concurrency, host limits, response byte cap and a total request deadline. |
| Read projection | Version updates and expose oldest indexing lag. Keep source originals outside the index. |
| Retention | Store only required feed snapshots and attribution. Separate diagnostic retention from reader history. |
Use one disposable AWS environment for the cloud exercise. Put the named resources in infra/template.yaml or your existing IaC tool, pass resource IDs through configuration, and scope each runtime role to its own tables, buckets and queues. The diagram is a design to implement. It is not a claim that these resources have been deployed. Record the commands you used to deploy and remove the exercise resources.
For concrete provisioning commands, configuration wiring and cleanup, use the AWS foundation guide. It includes a deployable table/queue/object-storage foundation and explains which application and service adapters you still implement.
A provisioned queue or table does not make the local program use it. Configure resource IDs in the deployed runtime, replace the local adapter, and replay the same successful and failing operation against that runtime. Record the deployed commit and observable result, then remove the disposable resources using your infrastructure tool.
Extend the design after the baseline works
Worked follow-up: Rebuild multilingual story groups without losing attribution
Two headlines about an election may describe different events. A similarity score can nominate a pair but cannot safely replace stable article identity. A mistaken merge must be reversible.
| Starting design | Changed requirement |
|---|---|
| Stories are grouped using one canonicalization policy. | A new policy groups related stories across languages and can split old groups. |
Revised architecture. Follow the changed responsibility and failure path below. This is a design to implement. The supplied local example does not provision these components.
What to implement. Keep immutable article IDs, source URLs and language metadata. Build a separately versioned article-to-story mapping, leaving the old mapping available while the new one is evaluated. Compare known false merges and missed duplicates by language. Switch readers to one mapping version and preserve cursor behavior across the switch. Use S3 for retained source material and a separate index generation for derived groups, with explicit attribution in each result.
Walk through the result. Construct three articles: two translations of the same announcement and one later correction. Show the desired grouping and the evidence that distinguishes the correction. Introduce an incorrect merge, roll back the mapping pointer and demonstrate that source URLs and article identities never changed. Deliver a before/after membership table.
Add multilingual story grouping. Preserve original language and attribution, and evaluate false merges separately from missed duplicates. Similarity is a candidate signal, not identity.
Additional design cases, alternatives and original source notes
This is a commonly listed system-design interview prompt with a concrete practice contract. Assume 50,000 publisher feeds, 20 million daily readers, and new stories visible within a 60-second target for responsive, successfully polled sources. Publisher outages cannot meet that target. Surface stale-source status. Clarify service guarantees and a first version before filling the board with services.
01 · Try this input
- Input / starting state
- Same story
- Expected result
- Cluster canonical story while retaining source attribution.
Input / condition: Three publishers syndicate one article
02 · Try this input
- Input / starting state
- Feed refresh
- Expected result
- Refresh the item without creating an unrelated duplicate.
Input / condition: One publisher updates title
03 · Try this input
- Input / starting state
- Breaking news
- Expected result
- Hot-topic cache and read fanout do not block ingestion.
Input / condition: Topic query receives 5× normal traffic
04 · Try this input
- Input / starting state
- Unfollow
- Expected result
- Feed response stops showing it within the declared privacy/freshness bound.
Input / condition: Reader unfollows a source
Think from the contract to the boxes
Treat collection, canonicalization, ranking and feed reads as separate stages. Keep source article IDs and a canonical cluster ID. Dedupe is probabilistic candidate grouping followed by explainable rules. Precompute ordinary feeds where it helps, but do not copy every breaking story to every user synchronously. Recheck follows and muted topics when composing the page.
Publication and freshness budget
| Arrow | Identity and accepted state | Failure policy |
|---|---|---|
| Scheduler → fetcher | (publisher_id, poll_slot). A scheduled tick means work is due, not fetched. |
Per-domain concurrency cap and 5 s HTTP deadline. Record stale status when publisher limits prevent freshness. |
| Fetcher → article authority | (publisher_id, source_article_id, source_version) plus content hash. Commit article revision and outbox together. |
Replay revision without duplicating it. Credentials are scoped to that publisher/domain. |
| Outbox → indexer | (article_id, revision, cluster_version). Index is a projection. |
Reject older versions, retry failed writes. Outbox survives a crash before send. |
| Index/cache → feed composer | Candidate IDs, then current follow/mute and content policy checks. | Never treat a personalized cached response as current permission. |
A worked 60 s target is 20 s maximum poll delay + 5 s fetch + 10 s queue + 15 s normalize/index + 10 s index/cache visibility. Polling 50,000 feeds every 20 s needs 2,500 fetch starts/s before retries. Provider limits can make this target infeasible, requiring push feeds or an explicitly relaxed promise.
Keep publisher articles immutable by revision. Cluster merges store aliases from old cluster IDs to the chosen cluster. Splits increment cluster version and emit membership corrections. Preserve source attribution. Rebuild projected feeds from those corrections rather than changing article identity.
Check: commit revision 8, crash before indexing, replay twice, then deliver revision 7. One current revision 8 is visible. Pause invalidation, unfollow a source, and confirm feed composition removes it on the next authorized read. Record detection-to-visibility latency, not merely worker execution time.
First diagram: Trace publisher fetch → canonical story → topic index → feed composition. Mark which copy is source and which is projection.
| AWS service / general role | Why it fits this design | Alternative and when it fits better |
|---|---|---|
| Amazon EventBridge Scheduler / poll schedule | Trigger publisher fetches at per-source cadence. | SQS delayed work when retry timing needs tighter control. |
| AWS Lambda / fetch/normalize worker | Parse small source feeds and enforce per-domain limits. | ECS for heavy parsers and persistent connections. |
| Amazon OpenSearch Service / article search index | Find text/topic candidates and support filtered discovery. | Aurora full-text search at lower volume. |
| Amazon DynamoDB / article + follow state | Store canonical IDs and reader/source relationships. | Aurora when joins and consistency across follows dominate. |
| Amazon ElastiCache / feed/result cache | Protect hot stories and repeated feed reads. | CloudFront for public non-personalized pages. |
Service choice follows the contract: the box label gives the generic job, while the table explains the AWS product and a reasonable substitute. Name which component owns durable truth, where retries happen, and the guarantee each managed service does not provide by itself.
Pressure-test the design
Follow-up: Fanout-on-write can multiply one breaking story into millions of queue tasks. Compare a normal source with one followed by ten million readers.
Publisher fetches become rate-limited and malformed. Isolate domains, add backoff and quarantine, and tell readers when content is stale.
A new editorial policy changes canonicalization across regions. Define data ownership, safe reindexing, relevance quality measures, and rollback without duplicate feed entries.
Practice artifact: Trace publisher fetch → canonical story → topic index → feed composition. Mark which copy is source and which is projection. Then trace every row in the table, draw one failure, and state what the customer observes. Suggested rehearsal: 35 minutes design, 10 minutes to challenge the guarantees.
Evidence and origin: The current community interview-question catalog lists personalized news-aggregation reports at Amazon, Microsoft and Rippling. Interview dates are not shown. The entry does not show the interview date and is not a verified company rubric. The prompt contract, workload, outcomes, diagrams and solution here are original practice material. Treat company tags as reported sightings, not a prediction of your interview loop.
Interview report listing: Open the community question entry.