Independent implementation review — F1/F3/F4/F6
Integration note: this report preserves the independent reviewer’s observations and the test counts at each review stage. Subsequent integrated validation passed 209 methods across 55 suites; see the final validation record. The editor evidence and diagram review mentioned as finishing below are now complete. The small F1 calibration concern is also resolved: the first follow-up now requires deduplication across calls and stable returned snapshots, rather than another permutation of an already order-independent sum.
Reviewed the changing working tree on 2026-09-22 against the original findings in audit issue #1 and the user-priority standard in the teaching standard. This is a read-only review of the repository; test databases and browser processes were disposable. No issue/comment/push was performed.
Current verdict
The new work addresses the original F1/F3/F4/F6 acceptance criteria in substance with actual implementations, meaningful tests, and usable assessor packs. All four scoped findings can be considered implemented, subject to the final validation-record/visual-review limits below. Three uncovered reference defects were reproduced, sent to the owner, fixed, and independently rechecked successfully. A review-patch packaging gap was also fixed and tested. No remaining reference-code blocker has been established in the reviewed paths.
This assesses supplied preparation material. It does not demonstrate any learner's independent mastery or claim a hiring outcome. The incomplete importer starter/ is deliberately broken candidate material; it is not judged as a failed reference implementation.
Acceptance evidence
| Finding | Verified implementation evidence | Status/remaining check |
|---|---|---|
| F1: independent assessment | practice/README.md:12 defines candidate/assessor separation, choosing unseen variants, and retiring known variants; :40 defines two independent occasions, per-dimension gates, no averaging away invariant failures, and explicit limits. Five candidate and five assessor pages supply contracts, timed follow-ups, observable weak/adequate/strong anchors, failing approaches, and debrief links. attempt-record.md:6 records allowed files/tools, timing, artifacts, evidence and a second occasion. |
Met. The finishing files were subsequently inspected: assessor/scored-examples.md:9 and :41 score fictional senior and lead performances across all eight dimensions and two occasions, with explicit evidence and limitations; four executable held-back importer tests passed independently. |
| F3: practical/debug/review | Importer has separate starter/reference packages across coordinator, domain validation, transport, and real SQLite persistence; visible baseline syntax failure; page and retry bounds; fake-clock/server seams; crash after page persistence; replay/conflict handling; whole-page rollback; three proposed PRs and assessor decisions with regression names. | Met in scoped teaching/reference support. Reference money defect fixed/rechecked. PR files were converted to applicable unified patches; three new review tests passed independently, applying each in a disposable copy and proving the named failure or logging assertion. |
| F4: runtime/thread concurrency | bounded-executor/executor.py:66 admission predicate loop; :112 worker predicate and execution outside lock; :137 exception-safe accounting; :100 queued/running cancellation; :143 shutdown wake/drain/cancel/join. Runtime lesson distinguishes promises, threads, processes, database atomicity and cancellation. Tests use barriers/events and logical clock; watchdog detects an intentionally deadlocking nested-future implementation. |
Met in declared scope. Forced cancellation, strict producer fairness, fleet-wide limits and process durability are explicitly excluded or lead follow-ups, not falsely supplied. No defect established. |
| F6: runnable full-stack slice | Real HTTP router, SQLite schema/store, TypeScript view state, named controls/live status/focus; owner/version predicates and atomic idempotency response; tied (timestamp,id) cursor; browser fixtures for save A/type B, 409, stale GET, lost acknowledgment, ignored abort, search races, keyboard retry. measure.py compares equal query results/plans before and after a composite index. |
Met in explicit edit/search/paginate scope; create/delete and production authentication are honestly excluded. Stale-query cursor and padded-title defects fixed/rechecked. Full browser validation ledger was still being finalized; this review independently ran the targeted failed-query regression, not the entire seven-scenario browser suite. |
The teaching format is visible in the actual pages, not only the standard: importer, executor, runtime and editor open with concrete user/system problems, contract tables, expected values, excluded scope, prerequisite links and runnable commands. They explain a reusable sequence of tracing inputs, identifying state authority/invariants, changing a boundary and testing failures. Each has at least two useful diagrams; the larger labs add changed-requirement/failure diagrams. Candidate pages keep reference implementations separate. I inspected diagram source and explanatory context, but did not independently render the full new Mermaid corpus.
Reproduced reference defects, now fixed
-
F3: exact money lost through ambient Decimal rounding. Initial
coding/labs/importer/reference/model.py:10–17multipliedDecimal(amount) * 100before checking integral cents. Input0.290000000000000000000000000001returned('x',29,'USD');9999999999.999999999999999999999returned1_000_000_000_000cents. Both should reject fractional-cent input under the stated contract. All 10 existing importer tests initially passed. The owner replaced multiplication with exact coefficient/exponent conversion and added precision regressions. Independently rechecked: both inputs now raiseValueError. Current implementation ismodel.py:11–33; range checking and a 128-character input bound prevent unbounded conversion work. -
F6: failed new query reused old query's cursor/results. Initial
full-stack/bookmark-editor/web/app.ts:156–193changedcurrentQueryimmediately but retainedrows/nextCursor, then enabled Load more infinally. Real Chromium/HTTP/SQLite reproduction: initial blank search returns Beta/Alpha with cursor[100,"a"]; intercept new query Gamma with 503; Load more remains enabled; restore server and click it; request usesq=Gammawith the old cursor and screen displays Beta, Alpha, Gamma. Expected: only Gamma results, or no continuation until the new query succeeds. The owner now clears rows/cursor and hides continuation when starting a new query (app.ts:160–168) and added a regression. Independently rechecked current TypeScript via in-memory Node syntax stripping: after failure More is hidden and zero stale rows remain; retry shows only Gamma. -
F6: title validation disagreed with the schema and dropped the connection. Initial
server.py:137checked trimmed length but stored the original string againstschema.sql:5's 200-character maximum. PATCH{"title":"A" + 200 spaces,"expectedVersion":1}with valid token/key passed validation, then raisedsqlite3.IntegrityError; HTTP client receivedRemoteDisconnected. Expected either documented normalization or HTTP 400. The owner now checks the stored title's actual length and rejects empty/whitespace-only values. Independently rechecked: the same input returns HTTP 400. Integer version bounds were also tightened by the owner.
Final caveats and resolved packaging check
- F1 finishing files verified: the scored examples and held-back fixtures now exist and have been reviewed. Fictional scored examples are appropriately labeled; they do not masquerade as actual observed learner or employer results.
- F3 review patch usability, resolved: initial review snippets used bare
@@;git apply --checkreturnederror: No valid patches in input. The owner replaced them with valid unified patches and addedreview/test_review.py:13–51. Independently ran all three tests: PR101 demonstrably suppresses the expected conflict error, PR102 retries at 0.05 instead of 0.75, and PR103 logs one bounded metadata record without transaction ID/amount. The disposable-copy harness meets the original review-exercise acceptance better than static diff excerpts. - F1 small calibration caution: the coding assessor's minute-15 “any order” fixture adds little to a baseline that already contains nonadjacent duplicate
e1,e2,e1and commutative deltas. The conflict and atomic-batch follow-ups are genuinely stronger. Improve the first change if desired; it does not negate the useful later stages and is not a new blocker. - Validation honesty: editor
VALIDATION.mdcurrently lists five API methods and says browser results will follow. Keep final browser/type-check evidence synchronized with actual commands; Node syntax stripping is correctly distinguished from strict type checking.
Checks independently run
- Bounded executor:
PYTHONDONTWRITEBYTECODE=1 python -m unittest -vin its directory — 6 tests passed, including subprocess deadlock detection. - Importer reference: same command in importer directory — 10 tests passed before the money correction; direct precision counterexamples then failed as described; corrected direct repros now reject both inputs.
- Editor HTTP/SQLite:
python -m unittest discover -s tests -p 'test_*.py' -vwith bytecode disabled — 5 tests passed before the title correction; corrected padded-title direct repro returns 400. - Real Chromium targeted query-failure scenario — defect reproduced, then fix verified against current TypeScript and temporary real HTTP/SQLite server. Runtime: supplied Chromium plus Playwright, no AWS or production resources.
- Assessor held-back importer pack:
python practice/assessor/heldback_importer.py --package reference -v— 4 tests passed, covering multi-hop cycle, total transport/retry budget, crash/reopen normalized replay, and late malformed-page atomicity. - Review patch harness:
python -m unittest discover -s review -p 'test_*.py' -vin importer directory — 3 tests passed, including actual patch application and targeted broken/preserved behavior.
Limits: this is the requested focused review, not a repeat whole-repository audit or new interview research. Code and assessments were still changing during review. I did not demand a production auth system, create/delete editor features, a frontend framework, distributed executor, strict fairness, or other scope excluded by the new contracts.