소스 검색

docs: resolve adversarial review of Windows completion scope

Drew Ritter 2 주 전
부모
커밋
4ce8c884b6

+ 18 - 11
docs/superpowers/specs/2026-09-09-proof-movie-windows-completion-design.md

@@ -1,6 +1,6 @@
 # Finish Windows support for the movie skill
 # Finish Windows support for the movie skill
 
 
-**Status:** Draft for human review; specification only. Implementation remains stopped.
+**Status:** Adversarial review complete; revised design ready for human review. Implementation remains stopped. See the [findings and scoped recheck](2026-09-09-proof-movie-windows-completion-review.md).
 
 
 **Baseline:** `feat/movie-os-compatibility` at `442a48d9`, based on the movie import at `f6617db1`.
 **Baseline:** `feat/movie-os-compatibility` at `442a48d9`, based on the movie import at `f6617db1`.
 
 
@@ -37,7 +37,7 @@ The probe's TUI success used a bounded checkpoint. Its event-driven screencast a
 | Browser and cards | Discover installed Chrome or Edge in standard Windows user/machine locations and on PATH; keep existing Mac/Linux discovery. An explicit `--browser` is authoritative: an unusable value reports an error. Render local HTML using `Path.resolve().as_uri()`, an owned temporary browser profile, and a bounded timeout. Verify the screenshot exists. Release the browser processes launched by this render on success/failure without touching the user's browser. |
 | Browser and cards | Discover installed Chrome or Edge in standard Windows user/machine locations and on PATH; keep existing Mac/Linux discovery. An explicit `--browser` is authoritative: an unusable value reports an error. Render local HTML using `Path.resolve().as_uri()`, an owned temporary browser profile, and a bounded timeout. Verify the screenshot exists. Release the browser processes launched by this render on success/failure without touching the user's browser. |
 | Subtitle paths | For hard subtitles, copy the SRT to a safe fixed basename in a temporary directory and run FFmpeg there, using absolute movie/output paths. This avoids interpreting drive letters, apostrophes, and backslashes as filter syntax. Preserve explicit soft-subtitle mode and the existing no-libass fallback, with accurate diagnostics. A burn failure must not be reported as missing libass. |
 | Subtitle paths | For hard subtitles, copy the SRT to a safe fixed basename in a temporary directory and run FFmpeg there, using absolute movie/output paths. This avoids interpreting drive letters, apostrophes, and backslashes as filter syntax. Preserve explicit soft-subtitle mode and the existing no-libass fallback, with accurate diagnostics. A burn failure must not be reported as missing libass. |
 | Text | Read/write YAML, JSON, HTML, SRT, and transcription text with explicit UTF-8, accepting UTF-8 BOM where shell-generated input requires it. CRLF input is valid. Unicode paths/content must survive native Windows defaults, including PowerShell 5.1. Machine-readable helper results must not depend on the console code page. |
 | Text | Read/write YAML, JSON, HTML, SRT, and transcription text with explicit UTF-8, accepting UTF-8 BOM where shell-generated input requires it. CRLF input is valid. Unicode paths/content must survive native Windows defaults, including PowerShell 5.1. Machine-readable helper results must not depend on the console code page. |
-| Local transcription | Replace the failing nested `python3` launch with a Windows-compatible uv-managed Python invocation isolated from the project being filmed. Return transcript data separately from library stdout diagnostics. Missing, malformed, or failed transcription remains a verification failure when `--verify on` is requested; `auto` may report verification unavailable, and `off` remains explicit. Preserve existing drift thresholds. |
+| Local transcription | Replace the failing nested `python3` launch with a Windows-compatible uv-managed Python invocation isolated from the project being filmed. Return transcript data separately from library stdout diagnostics. Under `--verify on`, transcribe both newly generated and reused cached WAVs: an unchanged manifest or a prior `--verify off` run does not establish verification. Missing, malformed, or failed transcription is nonzero under `on`; `auto` may report verification unavailable, and `off` remains explicit. Preserve existing drift thresholds. |
 
 
 The narration change addresses observed failures: native Windows could synthesize audio but silently skip requested transcription, and a native-library stdout warning could be mistaken for transcript text. Do not change the checker's general acceptance policy as part of this fix.
 The narration change addresses observed failures: native Windows could synthesize audio but silently skip requested transcription, and a native-library stdout warning could be mistaken for transcript text. Do not change the checker's general acceptance policy as part of this fix.
 
 
@@ -49,13 +49,16 @@ Add `skills/proving-it-works-with-a-movie/examples/film-terminal.py` as the Wind
 
 
 - `serve --shell powershell51|powershell7|gitbash --directory <new-session-dir>` owns one recorded shell and one browser page. Allow explicit shell/browser/ttyd executable paths and `--cwd` (default: the invoking working directory). Launch ttyd with explicit `-w` for that cwd, writable mode, one client, and loopback binding.
 - `serve --shell powershell51|powershell7|gitbash --directory <new-session-dir>` owns one recorded shell and one browser page. Allow explicit shell/browser/ttyd executable paths and `--cwd` (default: the invoking working directory). Launch ttyd with explicit `-w` for that cwd, writable mode, one client, and loopback binding.
 - Keep `serve` in a foreground/background task owned by the invoking harness, as the visual companion does. Later shell tool calls control the same process. It is not an installed service or self-daemonizing launcher.
 - Keep `serve` in a foreground/background task owned by the invoking harness, as the visual companion does. Later shell tool calls control the same process. It is not an installed service or self-daemonizing launcher.
-- `request --directory <session-dir> --file <UTF-8-JSON>` submits a request. Files avoid inline command length and cross-shell JSON quoting problems. Each request has a positive unique `id` and an `operation`; duplicate IDs are rejected. Writes and replies are atomic. Default output is acknowledgment JSON, which does not claim command completion. `--wait-result` waits up to `--timeout` (default: 30 seconds) and exits nonzero for failure, unknown/interrupted outcome, or missing result. A client wait timeout does not cancel the pending command.
-- Supported operations are `run` (`command`, `timeout_seconds`), `begin-take` (`name`), `end-take`, `key` (`key`), `inspect`, `close`, and `cancel`. Support printable keys, Enter, Escape, arrows, Tab, and Ctrl-C. Diagnostic probe phases, crash fixtures, and aggregation are test-only.
-- `run` acknowledges acceptance without waiting for the command to finish. A separate result records its outcome; its `timeout_seconds` defaults to 90 seconds and is independent of the client's wait timeout. Reuse the probe's numbered request/ack/result files under `control/`, with session ID and request ID in replies. `inspect` exposes pending request and active take. A second `run` while one is pending fails before typing; take controls and intentional `key` input remain available. A known failing command can be followed by another command; it does not itself terminate recording.
+- `request --directory <session-dir> --file <UTF-8-JSON>` submits a request. Files avoid inline command length and cross-shell JSON quoting problems. A session has one requesting controller; IDs are consecutive integers starting at one, matching the probe's ordered reader. Expose `next_request_id` in the ready/status files and `inspect` result; publish the updated status before acknowledging a consumed ID. Reject duplicate or non-next IDs before publication, with an immediate error rather than waiting for a missing file. Writes and replies are atomic. Default output is acknowledgment JSON, which does not claim command completion; a rejection is nonzero. `--wait-result` waits up to `--timeout` (default: 30 seconds) and exits nonzero for command/control failure, unknown/interrupted outcome, or missing result. A client wait timeout does not cancel the pending command.
+- `result --directory <session-dir> --id <accepted-id> --timeout <seconds>` waits for or reads an existing result without submitting another request or consuming an ID. Default timeout is 30 seconds; use the same completion/exit-status rules as `request --wait-result`. This is the supported follow-up after acknowledgment or client timeout.
+- Supported request operations are `run` (`command`, optional `timeout_seconds`, optional `native_producer`), `begin-take` (`name`), `end-take`, `key` (`key`), `inspect`, `close`, and `cancel`. Support printable keys, Enter, Escape, arrows, Tab, and Ctrl-C. Diagnostic probe phases, crash fixtures, and aggregation are test-only.
+- `run` acknowledges acceptance without waiting for the command to finish. Its result deadline defaults to 90 seconds and is independent of the client's wait timeout. Reuse the probe's numbered request/ack/result files under `control/`, with session ID and request ID in replies. `inspect` exposes pending request and active take. A second `run` while one is pending fails before typing, leaving the session usable; take controls and intentional `key` input remain available. A known failing command can be followed by another command; it does not itself terminate recording.
 - `end-take` stops capture only. It preserves the shell, working directory, variables, pending command, and browser connection for a later take, including across separate harness tool calls.
 - `end-take` stops capture only. It preserves the shell, working directory, variables, pending command, and browser connection for a later take, including across separate harness tool calls.
 - Readiness requires a nonce, shell identity, and cwd returned through the filmed terminal, plus a readable preflight image. Opening another ttyd connection is not an observation mechanism. A reconnect fails the session instead of silently replacing its identity.
 - Readiness requires a nonce, shell identity, and cwd returned through the filmed terminal, plus a readable preflight image. Opening another ttyd connection is not an observation mechanism. A reconnect fails the session instead of silently replacing its identity.
 
 
-Command results distinguish `completed`, `unknown`, and `interrupted`; shell success is nullable. PowerShell must preserve the submitted statement's semantics and distinguish cmdlet errors from an attributed native exit code. Capture status before logging. Bash must preserve an identified producer's failure through a logging pipeline. An opaque script's final status does not establish success of every internal command. Retain the existing probe's relevant outcome cases as regressions for the adapted code.
+Command results distinguish `completed`, `unknown`, and `interrupted`; shell success is nullable. Preserve raw PowerShell `$?` separately from observed request-attributed errors, including PS5.1 expression-wrapper behavior; instrumentation must not change the submitted statement's semantics. Capture status before logging.
+
+`native_producer` identifies the native executable for a direct invocation or the first stage of a pipeline followed by logging commands. This is the bounded attribution contract already exercised by the probe: Bash uses the first saved `PIPESTATUS` entry; PowerShell uses that request's native exit status. Opaque/mixed scripts omit this field and report their observed shell outcome without claiming individual internal commands succeeded. Do not infer producer identity or reuse stale native status. A completed result is successful only when shell success is true, no request-attributed shell error is present, and any explicitly identified producer has a known zero exit code. Thus successful `tee`/`Tee-Object` cannot hide a producer failure. Retain the existing probe's relevant outcome cases as regressions for the adapted code.
 
 
 ### Capture choice and deliberately limited geometry
 ### Capture choice and deliberately limited geometry
 
 
@@ -63,9 +66,9 @@ Use **bounded `Page.captureScreenshot` PNG capture at 5 fps**, with a fixed **16
 
 
 - Observe and record actual terminal rows/columns; require at least 80 columns for the existing 60-character completion framing. Do not offer resize controls. A detected terminal geometry change fails the session with an actionable message.
 - Observe and record actual terminal rows/columns; require at least 80 columns for the existing 60-character completion framing. Do not offer resize controls. A detected terminal geometry change fails the session with an actionable message.
 - The CDP receiver must continue handling terminal output, command replies, and cancellation during capture. Use a two-second screenshot response deadline; a timeout or lost browser fails the active take.
 - The CDP receiver must continue handling terminal output, command replies, and cancellation during capture. Use a two-second screenshot response deadline; a timeout or lost browser fails the active take.
-- Save ordered PNG frames and capture timestamps. Normalize them to the declared 5 fps without shortening elapsed take time; duplicates can fill the interval between successful samples, but cannot hide a capture gap exceeding two seconds. The exported duration must match the recorded take within one frame (0.2 seconds).
+- Treat 5 fps as the target sampling cadence, not a guarantee that every 0.2-second event is observed. Record monotonic take boundaries and screenshot request/completion times. A take starts at its first completed screenshot; `begin-take` completes only then. Place frames on the 0.2-second output grid using the latest screenshot completed at or before that grid time, never a future screenshot. Record duplicated intervals; a gap between completed captures, or from the last capture to take end, exceeding two seconds fails the take. Match total take duration within one frame (0.2 seconds). Total duration alone does not establish event timing.
 - `end-take` returns a take directory and the corresponding existing `kind: frames` scene fields (`src`, `rate: 5`). These artifacts feed the existing assembler directly.
 - `end-take` returns a take directory and the corresponding existing `kind: frames` scene fields (`src`, `rate: 5`). These artifacts feed the existing assembler directly.
-- Check preflight pixels and final exported frames. A static application can legitimately produce identical frames; image similarity alone is not a stalled-capture test. Test visible command output and a TUI while it still owns input, before its exit key.
+- Check preflight pixels and final exported frames. A static application can legitimately produce identical frames; image similarity alone is not a stalled-capture test. Acceptance must show several successive visible TUI states, each held for at least 1.3 seconds, in the automatically captured sequence while the command owns input, before its exit key. Check their order and placement against capture times. A manually requested snapshot cannot rescue this acceptance check; missing states fail it.
 - At the fixed geometry, long wrapping output followed by a command completion must work. Invalid/truncated completion records time out as unknown; they never produce success. A general VT emulator and arbitrary resizing are outside this delivery.
 - At the fixed geometry, long wrapping output followed by a command completion must work. Invalid/truncated completion records time out as unknown; they never produce success. A general VT emulator and arbitrary resizing are outside this delivery.
 
 
 The screenshot cadence and fixed-geometry adapter require a focused actual Windows test before committing to further recorder extraction. If either fails, retain the failure and propose a bounded alternative; do not silently resume the old recorder architecture.
 The screenshot cadence and fixed-geometry adapter require a focused actual Windows test before committing to further recorder extraction. If either fails, retain the failure and propose a bounded alternative; do not silently resume the old recorder architecture.
@@ -74,6 +77,10 @@ The screenshot cadence and fixed-geometry adapter require a focused actual Windo
 
 
 Reuse the probe's Windows Job Object mechanism: establish containment before recorder-owned roots execute, and retain kill-on-close semantics. Normal close, cancel, browser loss, and forced recorder exit must remove its browser/ttyd/shell descendants within ten seconds. An unrelated process must survive. Cleanup still runs when capture, control-file, or report writes fail; such failures remain nonzero results. Pending command success becomes unknown/interrupted when observation is lost.
 Reuse the probe's Windows Job Object mechanism: establish containment before recorder-owned roots execute, and retain kill-on-close semantics. Normal close, cancel, browser loss, and forced recorder exit must remove its browser/ttyd/shell descendants within ten seconds. An unrelated process must survive. Cleanup still runs when capture, control-file, or report writes fail; such failures remain nonzero results. Pending command success becomes unknown/interrupted when observation is lost.
 
 
+A command-result deadline expiring records an unknown outcome, fails the active take, and terminates the session with owned-resource cleanup. It must not clear pending state and type another command into the still-running program. Browser loss, capture timeout, or a forbidden geometry change likewise ends the session. Client wait deadlines have no such effect.
+
+`close` finalizes a healthy active take; `cancel` marks an active take incomplete. Both mark any pending command interrupted with null success and release owned processes. Earlier completed takes remain available. A shutdown acknowledgment means only accepted; publish the successful close/cancel result only after cleanup finishes. That success describes resource release, not the interrupted command. A finalization, cleanup, or result-write failure makes shutdown nonzero or leaves no successful result; the client must not infer success from a missing result.
+
 Keep ownership code specific to this recorder and its launched browser. Extract a small Windows helper only where it avoids duplicating the proven mechanism; do not introduce the previous generic `OwnedProcesses` framework.
 Keep ownership code specific to this recorder and its launched browser. Extract a small Windows helper only where it avoids duplicating the proven mechanism; do not introduce the previous generic `OwnedProcesses` framework.
 
 
 ## 3. Update only the Windows-facing guidance
 ## 3. Update only the Windows-facing guidance
@@ -94,16 +101,16 @@ All results below are required unless explicitly conditional. Reuse prepared too
 | --- | --- |
 | --- | --- |
 | Three native Windows workflow runs | One run invoked from each of PS5.1, PS7, and Git Bash on Ballmer, recording the corresponding shell. Record actual tool paths, native Windows Python identity, shell versions, and ordinary-user token. |
 | Three native Windows workflow runs | One run invoked from each of PS5.1, PS7, and Git Bash on Ballmer, recording the corresponding shell. Record actual tool paths, native Windows Python identity, shell versions, and ordinary-user token. |
 | One repeatable fixture per run | Combine a rendered title card, still image, changing real browser content, terminal takes, and an existing movie segment into a narrated, hard-subtitled movie. Use a path such as `movie O'Brien λ & [take]`, CRLF scene input, and a separate work directory. Preserve source movie audio and measured offsets; cover both narration-longer and visuals-longer scene timing. |
 | One repeatable fixture per run | Combine a rendered title card, still image, changing real browser content, terminal takes, and an existing movie segment into a narrated, hard-subtitled movie. Use a path such as `movie O'Brien λ & [take]`, CRLF scene input, and a separate work directory. Preserve source movie audio and measured offsets; cover both narration-longer and visuals-longer scene timing. |
-| Terminal behavior per recorded shell | Known success/failure; PowerShell cmdlet/native distinctions; failure through logging; one command running across two takes and separate tool calls; cwd/variable persistence; long wrapped output; active TUI pixels before the exit key. Validate the fixed viewport/capture cadence and reject a geometry change. |
+| Terminal behavior per recorded shell | Known success/failure; PowerShell cmdlet/native distinctions and raw expression-wrapper status; explicitly attributed producer failure through logging must make `result` nonzero; one command running across two takes and separate tool calls; cwd/variable persistence; long wrapped output; several ordered TUI states in automatic frames before the exit key. Validate timestamp placement and fixed geometry. |
 | Local voice per workflow | Fresh local Piper synthesis and actual transcription with no cloud key, including a repeat using cached models. Invoke the finished helper through `narrate --verify on`. Synthetic audio/text comparisons and skipped ASR do not satisfy this check. Additional OS network-isolation experiments are unnecessary. |
 | Local voice per workflow | Fresh local Piper synthesis and actual transcription with no cloud key, including a repeat using cached models. Invoke the finished helper through `narrate --verify on`. Synthetic audio/text comparisons and skipped ASR do not satisfy this check. Additional OS network-isolation experiments are unnecessary. |
 | Final movie per workflow | All five tools complete; hard subtitles are visible, narration audible, action and timing visible in the finished movie; checker passes; inspect its contact sheet and transcribe/check the rendered audio. Record artifacts and commands. |
 | Final movie per workflow | All five tools complete; hard subtitles are visible, narration audible, action and timing visible in the finished movie; checker passes; inspect its contact sheet and transcribe/check the rendered audio. Record artifacts and commands. |
 | Browser fallback and desktop | Render an actual title card with Chrome and Edge once each. Run the desktop preflight in the existing interactive session and inspect real app pixels; separately verify that unavailable capture is reported as unavailable. No RDP/compositor matrix. |
 | Browser fallback and desktop | Render an actual title card with Chrome and Edge once each. Run the desktop preflight in the existing interactive session and inspect real app pixels; separately verify that unavailable capture is reported as unavailable. No RDP/compositor matrix. |
-| Focused negative tests | Bad browser override, browser timeout, missing libass/soft fallback, malformed or unavailable ASR under `--verify on`, no-Bash PowerShell launch, screenshot timeout, geometry change, and cleanup after control/report failures. Use controlled tests for these failure cases. |
+| Focused negative tests | Bad browser override, browser timeout, missing libass/soft fallback, malformed or unavailable ASR under `--verify on` for both fresh and off→on cached clips, no-Bash PowerShell launch, screenshot timeout, geometry change, and cleanup after control/report failures. Control cases cover immediate gap/duplicate-ID rejection, wait-only result retrieval after client timeout, second-run rejection, command-timeout teardown, and no successful shutdown result before cleanup. Use controlled tests for these failure cases. |
 | Ownership regressions | Normal close, cancel, browser loss, and forced exit on the final recorder for each recorded shell. Reuse the existing child/grandchild and unrelated-sentinel fixtures. |
 | Ownership regressions | Normal close, cancel, browser loss, and forced exit on the final recorder for each recorded shell. Reuse the existing child/grandchild and unrelated-sentinel fixtures. |
 | Existing platforms | Run the existing portable regressions on the available native Mac and Linux runner, including the modified browser/subtitle/narration code paths. Preserve Unix instructions. No additional WSL, Rosetta, Intel-hardware, or architecture certification gate. |
 | Existing platforms | Run the existing portable regressions on the available native Mac and Linux runner, including the modified browser/subtitle/narration code paths. Preserve Unix instructions. No additional WSL, Rosetta, Intel-hardware, or architecture certification gate. |
 | Skill instructions | Four focused agent sessions: before/after using actual PowerShell tools and before/after using actual Git Bash tools. Use the same Windows scenario to exercise paths, no-key narration, terminal commands, and honest capture failure. Follow `superpowers:writing-skills`; retain transcripts and before/after results. No new eval platform or 36-session matrix. |
 | Skill instructions | Four focused agent sessions: before/after using actual PowerShell tools and before/after using actual Git Bash tools. Use the same Windows scenario to exercise paths, no-key narration, terminal commands, and honest capture failure. Follow `superpowers:writing-skills`; retain transcripts and before/after results. No new eval platform or 36-session matrix. |
 
 
-Fix demonstrated failures and rerun the affected checks. A final integration run uses the final code; unchanged prerequisite investigations and already-reviewed unrelated fixtures do not need repeating.
+Share the fixture and downloaded model caches across shell runs; each run must still produce its own final-code artifacts. Fix demonstrated failures and rerun the affected checks. A final integration run uses the final code; unchanged prerequisite investigations and already-reviewed unrelated fixtures do not need repeating.
 
 
 ## Delivery and stopping point
 ## Delivery and stopping point
 
 

+ 36 - 0
docs/superpowers/specs/2026-09-09-proof-movie-windows-completion-review.md

@@ -0,0 +1,36 @@
+# Adversarial review of the Windows completion spec
+
+**Reviewed draft:** `cd5e0bd7`.
+**Spec:** [Windows completion design](2026-09-09-proof-movie-windows-completion-design.md).
+**Reviewer:** Independent Codex subagent `/root/review_windows_completion`, given the user goal, current spec, and relevant source paths without the controller's full conversation.
+**Scope:** Design correctness, scope discipline, existing-code compatibility, and sufficient Windows acceptance. Read-only review; no implementation, runtime tests, provisioning, remote-host work, or additional agents.
+
+## Initial verdict
+
+Revise before implementation. The reviewer judged the narrowed architecture proportionate and found five mandatory contract corrections. None required restoring the old framework, platform matrix, or eval infrastructure.
+
+| Finding | Evidence in the reviewed source | Spec correction |
+| --- | --- | --- |
+| R1 — Producer attribution and success (P1) | Probe `dispatch` depends on `native_producer`; Bash `emit_bash` selects the first saved pipeline status. The draft required attribution but omitted the input field and aggregate success rule. | Include the field; limit its promise to a direct native invocation or first pipeline stage followed by logging. Keep raw PowerShell status and observed errors distinct. A failed/unknown identified producer cannot yield a successful command result. |
+| R2 — Cached ASR verification (P1) | `narrate`'s unchanged-WAV/manifest branch bypasses all transcription, including under `--verify on`. The controller independently identified and sent this case to the reviewer, who confirmed it. | Explicit `on` verifies reused WAVs as well as new audio. Add the off→on cached-clip case to the focused checks; unchanged text does not establish verified audio. |
+| R3 — Complete asynchronous control (P2) | Probe reads exactly the next consecutive request filename; arbitrary positive IDs can block it. Duplicate rejection prevents resubmitting solely to wait for an existing result. | Specify consecutive IDs, one controller, visible next ID, and prompt gap/duplicate rejection. Add a wait-only `result` command; distinguish acknowledgment from completion and client timeout from cancellation. |
+| R4 — Timeout/shutdown state (P2) | Probe command timeout terminates the session, but the draft left this ambiguous. Probe dispatch publishes close/cancel replies before cleanup. | Command timeout records unknown and tears down the session. Define active-take/pending-command behavior for close/cancel. Successful shutdown completion is published only after resource cleanup, separately from acknowledgment. |
+| R5 — Event timing and automatic capture (P2) | Draft allowed nearly two-second duplicate intervals while requiring only total-duration accuracy; a transient could be omitted without failing that duration check. | Specify monotonic boundaries, capture request/completion timestamps, conservative placement on the output grid, and visible gaps. Require several ordered TUI states in automatic frames before exit; a manual checkpoint cannot substitute. |
+
+The controller checked R1–R4 against the cited implementation paths and R5 against the actual capture/export contract before editing. The revisions add focused checks inside the existing Windows validation work, not new infrastructure.
+
+## Boundaries retained
+
+- Fixed geometry removes runtime resize support. It does not prove wrapping/redraw correctness.
+- Periodic screenshots are a proposed implementation choice. Existing successful single screenshots do not establish sustained capture correctness.
+- A focused actual Windows cadence/wrapping test remains required before further recorder extraction. Its failure requires a bounded design decision, not automatic expansion into the old plan.
+- The three Windows shell runs may share downloaded model caches and fixture setup, but each must produce its own final-code evidence.
+- Existing agents remain stopped. This newly authorized reviewer performs review only.
+
+## Scoped recheck
+
+The same reviewer rechecked R1–R5 and contradictions introduced by their revisions. **All five are resolved sufficiently for implementation; no new concrete blocker or mandatory scope reduction was found.** The revised design is ready to guide implementation when execution is authorized.
+
+The reviewer found no evidence that fixed-geometry periodic screenshots fail as a design. Sustained capture and wrapping correctness remain unproven runtime questions covered by the focused Windows decision gate. Design approval does not establish Windows support or replace that test.
+
+Both review passes were read-only. The controller changed only this review record and the completion spec; implementation did not resume.