Independent implementation review — F1/F3/F4/F6SUPPORTING MATERIAL
REFERENCE SHELF
Your guided curriculum
SUPPORTING MATERIALGUIDED READING

Independent implementation review — F1/F3/F4/F6

Integration note: this report preserves the independent reviewer’s observations and the test counts at each review stage. Subsequent integrated validation passed 209 methods across 55 suites; see the final validation record. The editor evidence and diagram review mentioned as finishing below are now complete. The small F1 calibration concern is also resolved: the first follow-up now requires deduplication across calls and stable returned snapshots, rather than another permutation of an already order-independent sum.

Reviewed the changing working tree on 2026-09-22 against the original findings in audit issue #1 and the user-priority standard in the teaching standard. This is a read-only review of the repository; test databases and browser processes were disposable. No issue/comment/push was performed.

Current verdict

The new work addresses the original F1/F3/F4/F6 acceptance criteria in substance with actual implementations, meaningful tests, and usable assessor packs. All four scoped findings can be considered implemented, subject to the final validation-record/visual-review limits below. Three uncovered reference defects were reproduced, sent to the owner, fixed, and independently rechecked successfully. A review-patch packaging gap was also fixed and tested. No remaining reference-code blocker has been established in the reviewed paths.

This assesses supplied preparation material. It does not demonstrate any learner's independent mastery or claim a hiring outcome. The incomplete importer starter/ is deliberately broken candidate material; it is not judged as a failed reference implementation.

Acceptance evidence

Finding Verified implementation evidence Status/remaining check
F1: independent assessment practice/README.md:12 defines candidate/assessor separation, choosing unseen variants, and retiring known variants; :40 defines two independent occasions, per-dimension gates, no averaging away invariant failures, and explicit limits. Five candidate and five assessor pages supply contracts, timed follow-ups, observable weak/adequate/strong anchors, failing approaches, and debrief links. attempt-record.md:6 records allowed files/tools, timing, artifacts, evidence and a second occasion. Met. The finishing files were subsequently inspected: assessor/scored-examples.md:9 and :41 score fictional senior and lead performances across all eight dimensions and two occasions, with explicit evidence and limitations; four executable held-back importer tests passed independently.
F3: practical/debug/review Importer has separate starter/reference packages across coordinator, domain validation, transport, and real SQLite persistence; visible baseline syntax failure; page and retry bounds; fake-clock/server seams; crash after page persistence; replay/conflict handling; whole-page rollback; three proposed PRs and assessor decisions with regression names. Met in scoped teaching/reference support. Reference money defect fixed/rechecked. PR files were converted to applicable unified patches; three new review tests passed independently, applying each in a disposable copy and proving the named failure or logging assertion.
F4: runtime/thread concurrency bounded-executor/executor.py:66 admission predicate loop; :112 worker predicate and execution outside lock; :137 exception-safe accounting; :100 queued/running cancellation; :143 shutdown wake/drain/cancel/join. Runtime lesson distinguishes promises, threads, processes, database atomicity and cancellation. Tests use barriers/events and logical clock; watchdog detects an intentionally deadlocking nested-future implementation. Met in declared scope. Forced cancellation, strict producer fairness, fleet-wide limits and process durability are explicitly excluded or lead follow-ups, not falsely supplied. No defect established.
F6: runnable full-stack slice Real HTTP router, SQLite schema/store, TypeScript view state, named controls/live status/focus; owner/version predicates and atomic idempotency response; tied (timestamp,id) cursor; browser fixtures for save A/type B, 409, stale GET, lost acknowledgment, ignored abort, search races, keyboard retry. measure.py compares equal query results/plans before and after a composite index. Met in explicit edit/search/paginate scope; create/delete and production authentication are honestly excluded. Stale-query cursor and padded-title defects fixed/rechecked. Full browser validation ledger was still being finalized; this review independently ran the targeted failed-query regression, not the entire seven-scenario browser suite.

The teaching format is visible in the actual pages, not only the standard: importer, executor, runtime and editor open with concrete user/system problems, contract tables, expected values, excluded scope, prerequisite links and runnable commands. They explain a reusable sequence of tracing inputs, identifying state authority/invariants, changing a boundary and testing failures. Each has at least two useful diagrams; the larger labs add changed-requirement/failure diagrams. Candidate pages keep reference implementations separate. I inspected diagram source and explanatory context, but did not independently render the full new Mermaid corpus.

Reproduced reference defects, now fixed

  1. F3: exact money lost through ambient Decimal rounding. Initial coding/labs/importer/reference/model.py:10–17 multiplied Decimal(amount) * 100 before checking integral cents. Input 0.290000000000000000000000000001 returned ('x',29,'USD'); 9999999999.999999999999999999999 returned 1_000_000_000_000 cents. Both should reject fractional-cent input under the stated contract. All 10 existing importer tests initially passed. The owner replaced multiplication with exact coefficient/exponent conversion and added precision regressions. Independently rechecked: both inputs now raise ValueError. Current implementation is model.py:11–33; range checking and a 128-character input bound prevent unbounded conversion work.

  2. F6: failed new query reused old query's cursor/results. Initial full-stack/bookmark-editor/web/app.ts:156–193 changed currentQuery immediately but retained rows/nextCursor, then enabled Load more in finally. Real Chromium/HTTP/SQLite reproduction: initial blank search returns Beta/Alpha with cursor [100,"a"]; intercept new query Gamma with 503; Load more remains enabled; restore server and click it; request uses q=Gamma with the old cursor and screen displays Beta, Alpha, Gamma. Expected: only Gamma results, or no continuation until the new query succeeds. The owner now clears rows/cursor and hides continuation when starting a new query (app.ts:160–168) and added a regression. Independently rechecked current TypeScript via in-memory Node syntax stripping: after failure More is hidden and zero stale rows remain; retry shows only Gamma.

  3. F6: title validation disagreed with the schema and dropped the connection. Initial server.py:137 checked trimmed length but stored the original string against schema.sql:5's 200-character maximum. PATCH {"title":"A" + 200 spaces,"expectedVersion":1} with valid token/key passed validation, then raised sqlite3.IntegrityError; HTTP client received RemoteDisconnected. Expected either documented normalization or HTTP 400. The owner now checks the stored title's actual length and rejects empty/whitespace-only values. Independently rechecked: the same input returns HTTP 400. Integer version bounds were also tightened by the owner.

Final caveats and resolved packaging check

  • F1 finishing files verified: the scored examples and held-back fixtures now exist and have been reviewed. Fictional scored examples are appropriately labeled; they do not masquerade as actual observed learner or employer results.
  • F3 review patch usability, resolved: initial review snippets used bare @@; git apply --check returned error: No valid patches in input. The owner replaced them with valid unified patches and added review/test_review.py:13–51. Independently ran all three tests: PR101 demonstrably suppresses the expected conflict error, PR102 retries at 0.05 instead of 0.75, and PR103 logs one bounded metadata record without transaction ID/amount. The disposable-copy harness meets the original review-exercise acceptance better than static diff excerpts.
  • F1 small calibration caution: the coding assessor's minute-15 “any order” fixture adds little to a baseline that already contains nonadjacent duplicate e1,e2,e1 and commutative deltas. The conflict and atomic-batch follow-ups are genuinely stronger. Improve the first change if desired; it does not negate the useful later stages and is not a new blocker.
  • Validation honesty: editor VALIDATION.md currently lists five API methods and says browser results will follow. Keep final browser/type-check evidence synchronized with actual commands; Node syntax stripping is correctly distinguished from strict type checking.

Checks independently run

  • Bounded executor: PYTHONDONTWRITEBYTECODE=1 python -m unittest -v in its directory — 6 tests passed, including subprocess deadlock detection.
  • Importer reference: same command in importer directory — 10 tests passed before the money correction; direct precision counterexamples then failed as described; corrected direct repros now reject both inputs.
  • Editor HTTP/SQLite: python -m unittest discover -s tests -p 'test_*.py' -v with bytecode disabled — 5 tests passed before the title correction; corrected padded-title direct repro returns 400.
  • Real Chromium targeted query-failure scenario — defect reproduced, then fix verified against current TypeScript and temporary real HTTP/SQLite server. Runtime: supplied Chromium plus Playwright, no AWS or production resources.
  • Assessor held-back importer pack: python practice/assessor/heldback_importer.py --package reference -v — 4 tests passed, covering multi-hop cycle, total transport/retry budget, crash/reopen normalized replay, and late malformed-page atomicity.
  • Review patch harness: python -m unittest discover -s review -p 'test_*.py' -v in importer directory — 3 tests passed, including actual patch application and targeted broken/preserved behavior.

Limits: this is the requested focused review, not a repeat whole-repository audit or new interview research. Code and assessments were still changing during review. I did not demand a production auth system, create/delete editor features, a frontend framework, distributed executor, strict fairness, or other scope excluded by the new contracts.