Explorar o código

chore(movie): drop the abandoned OS-rollout probe and superseded plans

The first, broader OS-compatibility rollout was stopped and replaced by the
narrower Windows completion. Its feasibility probe, probe cleanup test,
design, review, 12-task plan, results report, and the completion plan and
review record were internal execution artifacts with machine-specific paths.
The one probe-derived test list is inlined into the terminal suite.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Drew Ritter hai 2 semanas
pai
achega
b2f28a41c4

+ 0 - 709
docs/superpowers/plans/2026-09-09-proof-movie-os-compatibility.md

@@ -1,709 +0,0 @@
-# Proof Movie OS Compatibility Implementation Plan
-
-> **SUPERSEDED — DO NOT EXECUTE.** The user stopped this rollout. The [focused Windows completion spec](../specs/2026-09-09-proof-movie-windows-completion-design.md) replaces the remaining scope. Keep this plan as historical context for completed work; its unchecked tasks do not authorize further implementation or agent dispatch.
-
-**Goal:** Make the imported proof-movie skill usable on the agreed macOS, Linux, native Windows, WSL, and headless/remote environments, with evidence for each advertised combination.
-
-**Architecture:** Retain the five Python media tools and their artifact contracts. Add small shared path/browser helpers and an included terminal recorder with a persistent supervisor, shell-specific outcome adapters, and owned process trees. Use portable fixtures for shared behavior and real shell/OS runs for platform claims.
-
-**Tech Stack:** Python 3.10+, uv script environments, FFmpeg/ffprobe, Chromium-family browsers/CDP, ttyd, existing Piper/local ASR dependencies, Python unittest, Windows Job Objects through ctypes, and the external Quorum skill-evaluation harness.
-
-**Spec:** [Proof movie OS compatibility](../specs/2026-09-09-proof-movie-os-compatibility-design.md), with its [adversarial review and resolved findings](../specs/2026-09-09-proof-movie-os-compatibility-review.md).
-
-## Global Constraints
-
-- Native Windows must not require WSL; the PowerShell movie workflow must not require Git Bash.
-- PowerShell 5.1 and 7 are separate invocation tests. Git Bash is a Windows shell environment, not a synonym for WSL.
-- Git Bash must invoke native Windows Python and FFmpeg for its native Windows tests. WSL tests must invoke Linux tools.
-- Required validation starts with macOS arm64 and x64, Linux x64, and native Windows x64, plus WSL on Windows.
-- Do not introduce a Windows 11-only restriction solely because a test machine runs Windows 11.
-- Document `uv run --script` for all five extensionless tools. Retain existing Unix shebang execution for compatibility.
-- Use explicit UTF-8 for YAML, JSON, HTML, SRT, concat lists, and text evidence.
-- Pass external executable arguments as arrays.
-- Preserve scene kinds, manifest fields, offsets, CLI arguments, and output artifact formats.
-- Keep local Piper narration and local transcription available without a cloud key on the verified support combinations.
-- Do not introduce a new recording service, TTS service, global Node dependency, or Windows-only package manager requirement.
-- Keep the existing ttyd/tmux route for Unix environments.
-- Native Windows must not require tmux, Docker, or WSL.
-- Session persistence ends at `close`; this workflow does not launch local services intended to survive it.
-- A missing prerequisite in a required acceptance job is a failure of that job's setup, not a green compatibility result.
-- Skill instruction changes require `superpowers:writing-skills` and before/after pressure testing across multiple agent sessions.
-- Do not bundle changes to plugin hooks, the visual companion, unrelated skills, standalone-repo retirement, or the movie checker's general behavior.
-
-The spec supplies the detailed requirements behind these constraints. Its dependency-policy exception remains a maintainer decision for PR #2214. This plan does not authorize another import PR, publishing, or merging.
-
-## Baseline, execution order, and available host
-
-Inspected import: `f6617db1517488d4a8b66925519d5ffbeef965b3` from PR #2214, targeting `dev`. The skill is absent from this documentation branch's code baseline. At execution, use `superpowers:using-git-worktrees`, recheck the PR, and create an isolated implementation workspace based on its current import or the merged equivalent. Compare changes to the inspected import before applying this plan. Do not copy another import onto `dev` or modify the PR author's branch implicitly.
-
-Task 1 is a bounded, approximately half-day feasibility gate. If its mechanism fails, record the failure and revise the affected adapter design before tasks depend on it. The existing 5–8 engineering-day estimate remains provisional until this gate passes; installing prerequisites and acquiring the remaining OS runners are scheduling dependencies.
-
-The human partner provided `drew@ballmer.local` for SSH or Paseo access. A read-only SSH inventory on 2026-09-09 succeeded:
-
-| Observed item | Result in the SSH session |
-| --- | --- |
-| Host | BALLMER; Windows 11 Pro 10.0.26200; 64-bit OS |
-| PowerShell | Windows PowerShell 5.1.26100.9168 |
-| Git Bash | `C:\Program Files\Git\bin\bash.exe` exists |
-| Browser | Chrome and Edge exist in their standard Program Files locations |
-| Other PATH tools | Git and Node resolve to installed executables |
-| Not resolved on PATH | `pwsh.exe`, `uv.exe`, `ffmpeg.exe`, `ffprobe.exe`, `ttyd.exe` |
-| Misleading aliases | PATH `bash.exe` and `python.exe` resolve to WindowsApps aliases |
-
-Absence from PATH is not proof of absence from disk. Check the relevant user installation directories before downloading tools. Use the explicit Git Bash executable, not the WindowsApps Bash alias. Inventory process architecture separately; OS bitness alone does not establish native executable architecture. No tool installation or Windows movie test was performed while writing this plan.
-
-Use a task-owned directory under the remote user's profile for fixtures, tools that need provisioning, and artifacts. Record exact resolved paths, versions, build features, download origins, and checksums. Do not alter shell profiles, machine PATH, execution policy, firewall configuration, existing services, or saved credentials. Portable installations and process-local environment settings are sufficient. SSH proves remote command access; it does not establish access to an interactive desktop. WSL availability also needs a separate inventory.
-
-## File map and interfaces
-
-Paths below are relative to the implementation checkout. `S` means `skills/proving-it-works-with-a-movie`; `T` means `tests/proving-it-works-with-a-movie`. These abbreviations identify exact directories, not shell environment variables.
-
-| Files | Responsibility |
-| --- | --- |
-| `S/scripts/{assemble,burn-subtitles,check-movie,make-subtitles,narrate}` | Existing public tools; retain their CLI and output formats |
-| `S/scripts/media_paths.py` | Ordinary-file frame staging and FFconcat path syntax |
-| `S/scripts/browser_tools.py` | Browser discovery, file URLs, isolated card rendering |
-| `S/examples/film-terminal.py` | uv-script entry point: serve one recording session or submit a control request |
-| `S/examples/terminal_recorder/__init__.py` | Package marker only |
-| `S/scripts/processes.py` | Stdlib-only owned process lifecycle shared by card rendering and the recorder |
-| `S/scripts/windows_jobs.py` | Suspended Windows creation, job assignment, wait and termination |
-| `S/examples/terminal_recorder/shells.py` | Session-local Bash/PowerShell instrumentation and outcome attribution |
-| `S/examples/terminal_recorder/protocol.py` | Request validation, atomic files, framing parser and outcome schema |
-| `S/examples/terminal_recorder/cdp.py` | One CDP receiver, reply/event dispatch, page input and bounded screenshots |
-| `S/examples/terminal_recorder/supervisor.py` | Readiness, ordered controls, command/take state and cleanup |
-| `S/platform-support.md` and the existing seven skill/route Markdown files | Prerequisites, shell recipes, honest route decisions and support evidence |
-| `T/probe-windows.py` | Disposable feasibility entry point, retained as a reproducible native probe |
-| `T/run-tests.py`, `T/fixtures.py` | Portable unittest selection, capability preflight and generated fixtures |
-| `T/test_{assembly,checker,narration,paths,browser,subtitles,processes,shells,terminal,routes}.py` | Regression and integration assertions |
-| `T/fixtures/{browser.html,command.py,tui.py,heartbeat.py}` | Real browser state changes, command outcomes, interactive input, descendant ownership |
-| `T/run-acceptance.py`, `T/launch.ps1`, `T/launch.sh`, `T/README.md` | P/B/T/V/L acceptance groups, thin shell launchers, reproduction instructions |
-| `T/test-{assemble,check-movie,narrate}.sh` | Existing entry points; retain until portable equivalents cover their assertions |
-| `docs/superpowers/reports/2026-09-09-proof-movie-os-compatibility.md` | Results index, commands and links; generated movies stay outside core |
-| External eval repo: `scenarios/movie-os-{powershell,git-bash,no-key,paths,no-libass,headless}/{story.md,setup.sh,checks.sh,checks-manifest.json}` | Six versioned pressure scenarios using the same real movie fixture |
-
-Implement helpers only for the responsibilities above; no generic process framework or browser automation SDK. Recorder modules are packaged inside the executable example so the five media tools remain usable without importing recorder dependencies. The two shared process helpers use only the standard library; the terminal example adds its sibling `scripts` directory to its import path to reuse them.
-
-The new test runner contract is:
-
-```text
-uv run --script tests/proving-it-works-with-a-movie/run-tests.py --suite NAME [--require-capabilities]
-```
-
-`NAME` is `assembly`, `checker`, `narration`, `paths`, `browser`, `subtitles`, `processes`, `shells`, `terminal`, `routes`, or `all`. The runner uses stdlib unittest and PEP 723 test dependencies `pyyaml`, `pillow`, and `websockets`. Browser/local-voice integration cases are explicitly tagged in test code. A regular local run reports unavailable integration prerequisites as skips; `--require-capabilities` exits nonzero if any selected required case is skipped. Local successful unit tests never substitute for acceptance evidence.
-
-## Task 1: Prove the native Windows terminal mechanism on Ballmer
-
-**Files:** Create `T/probe-windows.py` and the reports file. Probe artifacts go in the remote task directory. No skill prose changes.
-
-**Interfaces:** Consumes explicit absolute paths to ttyd, a browser, and a recorded shell. Produces `probe.json` containing environment, shell identity, observed ttyd framing, readiness, command outcomes, continuity, cleanup and artifact paths. Every check has `passed`, `failed`, or `unavailable` status; only all required checks passing opens the gate.
-
-- [ ] **Establish the execution workspace and inventory the host.** Preserve an immutable import snapshot for later before/after evals. Read `T/test-*.sh` at that revision. Use an encoded PowerShell script for SSH to avoid multiple layers of shell quoting:
-
-```python
-import base64
-import subprocess
-
-script = r"""
-$ProgressPreference = 'SilentlyContinue'
-$PSVersionTable.PSVersion.ToString()
-[Environment]::Is64BitProcess
-Get-Command pwsh.exe,uv.exe,ffmpeg.exe,ffprobe.exe,ttyd.exe -ErrorAction SilentlyContinue |
-    Select-Object Name,Source | ConvertTo-Json
-Test-Path 'C:\Program Files\Git\bin\bash.exe'
-"""
-encoded = base64.b64encode(script.encode("utf-16le")).decode("ascii")
-subprocess.run([
-    "ssh", "-o", "BatchMode=yes", "-o", "ConnectTimeout=8",
-    "drew@ballmer.local", "powershell.exe", "-NoLogo", "-NoProfile",
-    "-NonInteractive", "-EncodedCommand", encoded,
-], check=True)
-```
-
-- [ ] **Resolve prerequisites inside the task environment.** Inspect user tool locations; provision missing uv, native FFmpeg with libx264/AAC/libass, ttyd with ConPTY, and PowerShell 7 from their official project distribution channels. Use existing Chrome or Edge. Verify `uv --version`, native Python identity, `ffmpeg -version/-filters/-encoders/-devices`, `ffprobe -version`, `ttyd --version`, and both PowerShell versions. Record failures as setup failures. Do not use an unverified WindowsApps alias.
-- [ ] **Write the probe's success assertions before the recorder.** It must exercise PowerShell 5.1, PowerShell 7, and Git Bash. Use this result contract in the probe's final assertions:
-
-```python
-def assert_probe(result):
-    assert result["terminal_client_count"] == 1
-    assert result["observed_session_id"] == result["filmed_session_id"]
-    assert result["nonce_before"] == result["nonce_after"]
-    assert result["shell_pid_before"] == result["shell_pid_after"]
-    assert result["variable_after"] == "persisted-λ"
-    assert result["long_command_survived_take_boundary"]
-    assert result["owned_children_remaining"] == []
-    assert result["unrelated_sentinel_alive"]
-    assert result["shutdown_seconds"] < 10
-```
-
-- [ ] **Run the assertions against the unimplemented probe and retain the failure.** Missing result fields or absent recorder behavior must fail; do not generate optimistic fixture results to pass the gate.
-- [ ] **Implement the narrow probe.** Use the Windows creation sequence specified in Task 4, one writable ttyd client, and one browser page. Enable CDP Network observation before navigating to ttyd. Decode output on that page's socket; never connect a second ttyd observer. This decoding seam must be exercised against the actual ttyd build, not merely a mocked frame:
-
-```python
-import base64
-
-def ttyd_output(event: dict) -> bytes:
-    frame = event["params"]["response"]
-    payload = frame["payloadData"]
-    raw = base64.b64decode(payload) if frame["opcode"] == 2 else payload.encode("utf-8")
-    return raw[1:] if raw[:1] == b"0" else b""
-```
-
-Pin observation to the page's identified terminal WebSocket request id, and validate the leading `0` output framing against captured messages. Feed decoded bytes into an incremental completion parser. Use session-local shell instrumentation, capture status before logging, and prove prompt readiness by nonce/identity/cwd plus a screenshot. Record actual errors and shell outcomes, not only exit-zero cases.
-- [ ] **Exercise two takes across real harness calls.** Run the supervisor under the harness's persistent execution facility; submit a first take and delayed command from another call, then a second take from a later call. Prove the command and variable survived while capture paused. Preserve the transcript of the separate calls. Add one extra-client attempt and a browser-loss case; neither may produce a successful continuity result.
-- [ ] **Exercise descendant ownership and crash cleanup.** A recorded Python worker starts a grandchild that writes a heartbeat; start a separate sentinel outside the job. Close normally, cancel, then kill the supervisor in a separate run. Check child/grandchild disappearance and heartbeat cessation, sentinel survival, and bounded shutdown. Retrieve screenshots, raw terminal messages, JSON results and logs with SSH/SFTP; record both launch and recording hosts.
-- [ ] **Record the gate disposition and commit the probe/report.** If any prerequisite or mechanism is unavailable, leave the gate incomplete. If a mechanism fails, attach its smallest reproduction and revise the relevant design before broad implementation. Commit message: `test(movie): validate native Windows recording mechanism` only after recording the actual result, including a failing result if the design must change.
-
-## Task 2: Preserve the imported regressions in a portable runner
-
-**Files:** Create `T/run-tests.py`, `T/fixtures.py`, `T/test_assembly.py`, `T/test_checker.py`, `T/test_narration.py`, and `T/README.md`. Keep the three Bash suites initially.
-
-**Interfaces:** `fixtures.run_tool(name: str, args: list[str], *, cwd: Path, env: dict[str, str] | None = None) -> subprocess.CompletedProcess[bytes]`; `fixtures.duration(path: Path) -> float`. Both use argv arrays and real executables. `run_tool` resolves the skill script from the checkout and invokes `[uv, "run", "--script", script, *args]`. Fixture setup owns its temporary directories and subprocesses.
-
-- [ ] **Port each existing assertion, preserving its intent and thresholds.** Assembly must still check successful assembly, approximately eight seconds total, body offset approximately two seconds, and subtitles starting at the measured offset. Port all nine checker cases and all five narration drift cases; list their names in the README for equivalence review.
-
-```python
-class AssemblyRegression(unittest.TestCase):
-    def test_narration_padding_and_offsets(self):
-        with tempfile.TemporaryDirectory() as directory:
-            work = Path(directory)
-            # The fixture helper writes the imported 2s image + 4s frames + 6s WAV case.
-            scenes = fixtures.assembly_fixture(work)
-            result = fixtures.run_tool("assemble", [str(scenes), str(work / "out.mp4")], cwd=work)
-            self.assertEqual(result.returncode, 0, result.stderr)
-            self.assertAlmostEqual(fixtures.duration(work / "out.mp4"), 8, delta=0.4)
-            offsets = json.loads((work / "segments/offsets.json").read_text(encoding="utf-8"))
-            self.assertAlmostEqual(offsets["body"], 2, delta=0.3)
-```
-
-`fixtures.assembly_fixture(work: Path) -> Path` generates the four ordered PNGs, six-second sine-wave stand-in, matching narration manifest, and original scene YAML. Mark the audio as a synthetic timing fixture; this is not a local speech test.
-- [ ] **Run the new runner before adding its fixture functions.** Expected: import/undefined-helper failure. Implement `run_tool`, `duration`, `assembly_fixture` and unittest discovery, then rerun `--suite assembly`, `--suite checker`, and `--suite narration`. Expected: the same 18 assertions pass on a capable host; unavailable prerequisites remain explicit.
-
-```python
-def run_tool(name, args, *, cwd, env=None):
-    uv = shutil.which("uv")
-    if uv is None:
-        raise RuntimeError("uv is required for movie tool tests")
-    script = Path(__file__).resolve().parents[2] / "skills/proving-it-works-with-a-movie/scripts" / name
-    return subprocess.run([uv, "run", "--script", str(script), *args],
-                          cwd=cwd, env=env, capture_output=True, timeout=900)
-```
-
-- [ ] **Run old and new suites on the same macOS checkout.** Compare named assertions, not just total counts. Keep the Bash versions until Task 10 converts their entry points to thin wrappers without losing coverage. Run portable equivalents on Ballmer as an early dependency check.
-- [ ] **Commit:** `test(movie): port regression fixtures to Python`.
-
-## Task 3: Make frame assembly and media paths portable
-
-**Files:** Create `S/scripts/media_paths.py` and `T/test_paths.py`; modify `S/scripts/assemble` and `T/test_assembly.py`.
-
-**Interfaces:** `media_paths.stage_frames(source: Path, destination: Path) -> list[Path]` copies lexically sorted `*.png` inputs to numbered ordinary files in a fresh destination. `media_paths.ffconcat_entry(path: Path) -> str` returns one FFconcat line with an absolute slash-form path and FFmpeg token escaping. Callers own and clean the fresh staging directory.
-
-- [ ] **Add a path fixture and ordering test.** Use a valid name such as `movie O'Brien λ & [take]`, a work directory outside the current directory, CRLF YAML, nested SRT directories, and absolute/relative inputs. Include native drive-letter paths on Windows. Distinct colored frames must remain in lexical source order.
-
-```python
-def test_frame_staging_uses_ordinary_ordered_files(self):
-    with tempfile.TemporaryDirectory() as directory:
-        root = Path(directory)
-        source = root / "O'Brien λ & [take]"
-        source.mkdir()
-        (source / "b.png").write_bytes(b"second")
-        (source / "a.png").write_bytes(b"first")
-        staged = media_paths.stage_frames(source, root / "staged")
-        self.assertEqual([p.read_bytes() for p in staged], [b"first", b"second"])
-        self.assertTrue(all(not p.is_symlink() for p in staged))
-        self.assertEqual(len(list(source.iterdir())), 2)
-```
-
-- [ ] **Run `run-tests.py --suite paths`; observe failure before adding the helper.** Implement staging and FFconcat escaping:
-
-```python
-def stage_frames(source: Path, destination: Path) -> list[Path]:
-    files = sorted(source.glob("*.png"))
-    if not files:
-        raise ValueError(f"no PNG frames in {source}")
-    destination.mkdir(parents=True, exist_ok=False)
-    result = []
-    for index, frame in enumerate(files):
-        target = destination / f"frame-{index:08d}.png"
-        shutil.copyfile(frame, target)
-        result.append(target)
-    return result
-
-def ffconcat_entry(path: Path) -> str:
-    value = path.resolve().as_posix()
-    if "\n" in value or "\r" in value:
-        raise ValueError("FFconcat paths cannot contain line breaks")
-    return "file '" + value.replace("'", "'\\''") + "'\n"
-```
-
-- [ ] **Replace FFmpeg glob input with the staged numbered sequence.** Supply `-framerate`, `-start_number 0`, and `frame-%08d.png`; retain lexical ordering and the existing rate. Write concat lists with `encoding="utf-8"`. Validate the helper's output by running FFmpeg on both Unix paths and real Windows drive paths, not by comparing an escaped string alone.
-- [ ] **Extend real assembly coverage to all four scene kinds.** Assert `card`/`image`/`frames` last `max(narration, visuals)` and that `movie` retains its own duration/audio even when narration is longer. Check measured offsets and decoded frame order; source PNGs must survive success and failure cleanup.
-- [ ] **Run `--suite paths` and `--suite assembly --require-capabilities` on macOS and Windows; commit:** `fix(movie): make frame and concat inputs portable`.
-
-## Task 4: Own recorder processes and their descendants
-
-**Files:** Create `S/scripts/{processes.py,windows_jobs.py}`, `S/examples/terminal_recorder/__init__.py`, `T/test_processes.py`, and `T/fixtures/heartbeat.py`. Refactor the proven Task 1 ownership code into these modules.
-
-**Interfaces:** `processes.OwnedProcesses()` is a context manager. `spawn(argv: list[str], *, cwd: Path, env: dict[str, str]) -> int` starts an owned process and returns its native PID. `run(argv: list[str], *, cwd: Path, env: dict[str, str], timeout: float) -> subprocess.CompletedProcess[bytes]` captures an owned finite command's raw output, waits with a deadline, and cleans the owned tree on timeout. `register_terminal_group(pid: int) -> None` records the validated Unix PTY process group. `close(grace_seconds: float = 3.0, kill_seconds: float = 5.0) -> None` waits for graceful shutdown, then terminates remaining owned processes with a deadline. `windows_jobs.WindowsJob.spawn` takes the same argv/cwd/env fields and returns the native PID; `terminate(exit_code: int) -> None`, `wait_empty(timeout: float) -> bool`, and `close() -> None` supply its cleanup. Handles are never inferred from a Bash PID.
-
-- [ ] **Add real child/grandchild tests before extracting the implementation.** `heartbeat.py` accepts `--directory`, `--role parent|child|grandchild|sentinel`; the parent starts the child, the child starts the grandchild, and each periodically writes a file containing its PID and timestamp. The sentinel runs outside `OwnedProcesses` and has independent cleanup in the test's `finally` block. Each worker uses ordinary inherited process ownership; test detachment from terminal input separately from permission to break out of the job.
-
-```python
-with processes.OwnedProcesses() as owned:
-    parent_pid = owned.spawn([sys.executable, str(worker), "--directory", str(work),
-                              "--role", "parent"], cwd=work, env=dict(os.environ))
-    self.assertGreater(parent_pid, 0)
-    fixture_pids = fixtures.wait_for_heartbeats(work, count=3, timeout=5)
-self.assertTrue(all(not fixtures.pid_alive(pid) for pid in fixture_pids))
-self.assertTrue(fixtures.pid_alive(sentinel_pid))
-```
-
-Add `fixtures.wait_for_heartbeats(directory: Path, count: int, timeout: float) -> list[int]` and `fixtures.pid_alive(pid: int) -> bool` using actual native process checks, with process creation time/owned handles where available to avoid PID reuse. Do not use process-name killing. Run `--suite processes`; expected initial failure: ownership module absent.
-- [ ] **Implement Windows suspended creation and job assignment.** Define correctly sized ctypes structures from the Windows SDK for `STARTUPINFOW`, `PROCESS_INFORMATION`, `JOBOBJECT_BASIC_LIMIT_INFORMATION`, `IO_COUNTERS`, and `JOBOBJECT_EXTENDED_LIMIT_INFORMATION`; declare argument/return types for every Win32 function, including pointer-sized handles. Unit-test structure sizes on the host. Create an unnamed job, set `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`, and never enable either breakaway flag. Encode argv with `subprocess.list2cmdline`, and pass a Unicode environment block plus absolute working directory to `CreateProcessW`.
-
-```python
-CREATE_SUSPENDED = 0x00000004
-CREATE_UNICODE_ENVIRONMENT = 0x00000400
-JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE = 0x00002000
-
-# Within WindowsJob.spawn, after CreateProcessW filled process_info:
-if not kernel32.AssignProcessToJobObject(self.handle, process_info.hProcess):
-    error = ctypes.get_last_error()
-    kernel32.TerminateProcess(process_info.hProcess, 1)
-    kernel32.CloseHandle(process_info.hThread)
-    kernel32.CloseHandle(process_info.hProcess)
-    raise ctypes.WinError(error)
-if kernel32.ResumeThread(process_info.hThread) == 0xFFFFFFFF:
-    error = ctypes.get_last_error()
-    kernel32.TerminateProcess(process_info.hProcess, 1)
-    kernel32.CloseHandle(process_info.hThread)
-    kernel32.CloseHandle(process_info.hProcess)
-    raise ctypes.WinError(error)
-kernel32.CloseHandle(process_info.hThread)
-```
-
-Retain the process handle until exit/wait. Close the non-inheritable job handle on all error paths and at normal completion; the supervisor itself stays outside its child job. If assignment is impossible in the harness's existing job hierarchy, fail preflight before any take. Do not resume an unowned root process as a fallback.
-- [ ] **Implement Unix process ownership and bounded output draining.** Spawn owned roots with `start_new_session=True`; register the ttyd shell's actual PTY process group after the readiness probe and before accepting user commands. During shutdown, keep the CDP/output readers running while requesting graceful exit, signal only registered owned groups, then enforce the kill deadline and reap owned roots. Implement finite `run` with task-owned binary output files or continuously drained pipes so waiting cannot deadlock on a full pipe. Test its timeout with a child/grandchild producer. The browser card helper and terminal recorder both reuse this ownership mechanism.
-- [ ] **Run ownership tests for normal close, exception, cancellation, and supervisor crash.** On Windows, crash closure of the job must terminate all owned descendants. Exercise job-assignment failure and assert no child runs. Confirm the sentinel survives every case and no heartbeat advances after shutdown. On Unix include a PTY child group distinct from ttyd's parent group. Run with ordinary user permissions.
-- [ ] **Commit:** `feat(movie): contain recorder process trees on Windows and Unix`.
-
-## Task 5: Discover browsers and render cards without platform-specific URLs
-
-**Files:** Create `S/scripts/browser_tools.py` and `T/test_browser.py`; modify `S/scripts/assemble`.
-
-**Interfaces:** `browser_tools.find_browser(explicit: str | None = None) -> Path` returns a resolved executable or raises `FileNotFoundError` with searched locations. `browser_tools.render_card(browser: Path, html: Path, png: Path, width: int, height: int) -> None` renders using an isolated temporary profile and a bounded process wait.
-
-- [ ] **Add explicit-selection, discovery, and real-render tests.** A supplied invalid browser path must fail clearly instead of silently selecting another executable. Test Windows Chrome/Edge installation locations plus PATH candidates, and existing macOS/Linux choices. A real card with Unicode text in the path fixture must have visible foreground pixels and the requested dimensions.
-
-```python
-def test_explicit_browser_path_is_authoritative(self):
-    with self.assertRaises(FileNotFoundError):
-        browser_tools.find_browser("missing-explicit-browser")
-```
-
-- [ ] **Run `--suite browser`; observe failures.** Implement `shutil.which`/platform location selection, `html.resolve().as_uri()`, and absolute screenshot paths. Preserve the current card content and dimensions. The argv seam is:
-
-```python
-command = [str(browser), "--headless", "--disable-gpu",
-           f"--user-data-dir={profile}", f"--window-size={width},{height}",
-           f"--screenshot={png.resolve()}", html.resolve().as_uri()]
-with processes.OwnedProcesses() as owned:
-    result = owned.run(command, cwd=html.parent.resolve(), env=dict(os.environ), timeout=30)
-    result.check_returncode()
-```
-
-Create `profile` using `TemporaryDirectory`; retain sandbox defaults where supported. A container-specific browser sandbox exception must be an explicit environment requirement rather than a universal disabled-sandbox default. Use owned process cleanup for any child processes that survive the bounded screenshot operation.
-- [ ] **Run actual screenshots on Ballmer and Linux with no display, plus macOS regression coverage.** Distinguish browser rendering from opening an artifact for the human partner. Missing display must not prevent headless card rendering.
-- [ ] **Commit:** `fix(movie): discover browsers and render portable file URLs`.
-
-## Task 6: Fix subtitle staging, text encoding, and local ASR isolation
-
-**Files:** Modify all five `S/scripts` tools only where their IO requires it; create `T/test_subtitles.py`; extend `T/test_narration.py` and `T/test_paths.py`.
-
-**Interfaces:** Existing public CLIs and manifest fields are unchanged. Extend `burn-subtitles.run(cmd: list[str], *, cwd: Path | None = None) -> bool` to pass the working directory. Preserve `narrate.transcribe_local(wav, model="base.en") -> str | None`; its nested interpreter is project-independent and its diagnostics explain an unavailable ASR.
-
-- [ ] **Add a real nested-SRT burn regression.** Generate a short clip and SRT under the apostrophe/Unicode path fixture. Launch the script from a different directory. On a libass-capable build assert the output says it burned subtitles and compare decoded frames to prove text pixels exist; also inspect a Unicode sample visually. Test explicit `--soft`, missing libass, and a forced burn failure while libass is present. The last case must retain the actual diagnostic and must not say libass is absent.
-
-```python
-self.assertEqual(result.returncode, 0, result.stderr)
-self.assertIn(b"burned into the picture", result.stdout)
-self.assertNotEqual(before_frame.tobytes(), after_frame.tobytes())
-self.assertTrue(subtitles.exists())
-```
-
-The fixture supplies `result` from `fixtures.run_tool`, and decoded Pillow images `before_frame`/`after_frame` from the same timestamp. Pixel differences alone are not a font-glyph check: retain the rendered sample for inspection.
-- [ ] **Run `--suite subtitles`; capture the original failure.** Stage the SRT under a fixed basename and keep all external movie paths absolute:
-
-```python
-with tempfile.TemporaryDirectory(prefix="movie-subs-") as directory:
-    stage = Path(directory)
-    shutil.copyfile(args.subs, stage / "subtitles.srt")
-    ok = run(["ffmpeg", "-nostdin", "-y", "-v", "error",
-              "-i", str(args.movie.resolve()),
-              "-vf", f"subtitles=filename=subtitles.srt:force_style='{style}'",
-              "-c:a", "copy", "-c:v", "libx264", "-preset", "medium",
-              "-pix_fmt", "yuv420p", str(args.out.resolve())], cwd=stage)
-```
-
-Keep style construction local; validate/escape font values for the filter separately from file paths. Choose an available platform font and retain explicit `--font`. Report whether fallback followed absent libass or an actual burn error. Preserve the source SRT and the pipeline's sidecar-delivery contract.
-- [ ] **Add UTF-8/CRLF and project-isolation regressions.** Read/write YAML, HTML, SRT, JSON and concat text with explicit UTF-8; accept CRLF. Diagnostic decode errors must identify the file. Preserve external raw bytes and use explicit UTF-8 only for owned Python helper output. In a temporary project with `requires-python = ">=9.99"`, launch the local helper and assert the project files, `.venv` and lockfile are unchanged.
-- [ ] **Make nested uv select its own compatible interpreter.** Use Python 3.11 for this helper's tested environment unless the real dependency probe demonstrates a different compatible version is required; record that choice. The parent tools retain their Python 3.10 minimum. Verify the actual helper invocation and a real cached ASR run:
-
-```python
-command = ["uv", "run", "--no-project", "--quiet", "--python", "3.11",
-           "--with", "faster-whisper", "python", "-c", ASR_SNIPPET,
-           str(Path(wav).resolve()), model]
-child_env = dict(os.environ, PYTHONIOENCODING="utf-8")
-out = subprocess.run(command, capture_output=True, text=True,
-                     encoding="utf-8", env=child_env, timeout=900)
-```
-
-An unavailable transcription may retain the current script's permissive result, but it must be visible and cannot pass V acceptance. Do not alter drift/checker thresholds or fix unrelated narration-cache behavior in this task.
-- [ ] **Run subtitle, path, narration and checker suites on capable macOS and Windows installations.** Existing local Homebrew FFmpeg lacks the needed hard-burn capability; provision/select a capable build for acceptance rather than counting a fallback as success.
-- [ ] **Commit:** `fix(movie): isolate subtitle and narration processing from host paths`.
-
-## Task 7: Define command outcomes and session-local shell adapters
-
-**Files:** Create `S/examples/terminal_recorder/{protocol.py,shells.py}`, `T/test_shells.py`, and `T/fixtures/command.py`. Reuse Task 1's working instrumentation after applying the outcome contract below.
-
-**Interfaces:** `protocol.Request` contains `session_id: str`, `request_id: int`, `op: str`, and operation-specific fields. `protocol.Outcome` uses the schema below. `protocol.CompletionParser(session_id: str)` has `expect(request_id: int) -> None` to set the active request and `feed(data: bytes) -> list[Outcome]` to parse only that request's completion. It clears the active request after emitting its outcome and rejects replay. `shells.instrument(command: str, *, shell: str, session_id: str, request_id: int, attribution: str, native_executable: str | None) -> str` returns text to type into the already-recorded shell. Supported shells are `bash`, `powershell-5.1`, and `powershell-7`; attribution is `native`, `shell`, or `opaque`.
-
-```python
-@dataclass(frozen=True)
-class Outcome:
-    session_id: str
-    request_id: int
-    completion: Literal["completed", "interrupted", "unknown"]
-    shell_success: bool | None
-    native_exit_code: int | None
-    shell_error: str | None
-    pipeline_status: list[int] | None
-```
-
-`completed` requires a Boolean `shell_success`; `interrupted` and `unknown` cannot imply success. `native_exit_code` is populated only for an identified native producer. Keep the exact command and its attribution beside the outcome in evidence. `command.py` accepts `--exit N`, `--delay SECONDS`, and `--text TEXT`, writes the actual text and exits accordingly.
-
-- [ ] **Add parser and actual-shell outcome tests.** Feed completion frames split at every byte boundary, mixed with terminal output and echoed instrumentation; assert exactly one matching outcome and no acceptance of a different session/request. Use a raw control-character prefix/suffix and base64-encoded UTF-8 JSON. Construct the framing at runtime so the complete prefix never appears in echoed command text.
-
-```python
-parser = protocol.CompletionParser("session-a")
-parser.expect(expected.request_id)
-payload = json.dumps(asdict(expected), ensure_ascii=False).encode("utf-8")
-frame = b"\x1e" + b"MOVIE:" + base64.b64encode(payload) + b"\x1f"
-actual = []
-for byte in frame:
-    actual.extend(parser.feed(bytes([byte])))
-self.assertEqual(actual, [expected])
-```
-
-`expected` is an `Outcome` for the active request in the fixture. The parser buffers bounded incomplete records, rejects malformed/oversized records, and does not discard raw terminal evidence.
-- [ ] **Run `--suite shells`; retain parser/import failures.** Implement the parser and session-local instrumenter. For Bash, capture status and producer pipeline statuses in one assignment command immediately after the beat:
-
-```bash
-__movie_status=$? __movie_pipeline=("${PIPESTATUS[@]}")
-```
-
-Then serialize the captured values and emit the frame. Preserve the original shell state except for uniquely prefixed instrumentation variables; do not edit profiles. For PowerShell, capture `$?` as the first statement after the command, then capture `$LASTEXITCODE` only for an attributed native producer:
-
-```powershell
-$__movie_success = $?
-$__movie_native = $LASTEXITCODE
-```
-
-The two statements appear directly after the submitted beat inside the generated session-local instrumentation. Cmdlet outcomes serialize `native_exit_code` as null regardless of `$LASTEXITCODE`. Use `try/catch` for observable terminating exceptions without changing error preferences; emit a failed outcome with the caught error. Parse failures or `exit` that prevent framing become `unknown`/`interrupted` with a diagnostic. Do not wrap the command in a PowerShell 5.1 expression that resets `$?` before capture.
-- [ ] **Exercise the full native error matrix through the filmed shell.** In both PowerShell versions: successful native command then failing cmdlet; failed native command then successful cmdlet; terminating/non-terminating errors; explicit script exit; parse error; success/failure through logging; and Unicode under a non-UTF-8 console code page. In Bash: nonzero native exit, producer failure followed by a successful logger, delayed completion, and pipeline statuses. Keep atomic statement beats; multi-statement scripts are opaque producers. When producer attribution through a pipeline cannot be preserved, require separate beats and report the restriction.
-- [ ] **Verify expected failures remain honest recording outcomes.** A correctly captured failed command may pass the recorder test with `shell_success=false`; an absent frame may not. Retain raw byte logs, decoding choice, exact commands, and outcome JSON. Restore any test-local console settings when the fixture ends.
-- [ ] **Commit:** `feat(movie): capture native shell outcomes without hiding failures`.
-
-## Task 8: Observe and capture the existing terminal page through CDP
-
-**Files:** Create `S/examples/terminal_recorder/cdp.py`; extend `T/test_browser.py` and create `T/test_terminal.py`.
-
-**Interfaces:** `cdp.Page.connect(debug_url: str) -> Page` is asynchronous. `Page.call(method: str, params: dict | None = None, timeout: float = 10) -> dict` correlates responses by CDP id through one receiver. `Page.events()` is an asynchronous iterator of event dictionaries. `Page.screenshot(path: Path, timeout: float = 10) -> None`, `Page.type_text(text: str) -> None`, `Page.key(name: str) -> None`, and `Page.close() -> None` share that connection. Terminal output feeds `CompletionParser`; it never opens another ttyd client.
-
-- [ ] **Add transport tests with interleaved replies/events and actual-page tests.** Interleave screenshot replies, keyboard replies, and binary terminal frames. Assert no event disappears while a screenshot waits; a lost connection fails pending calls with a bounded diagnostic. A real browser fixture must show a changed counter and a moving cursor overlay at measured times.
-
-```python
-pending = asyncio.create_task(page.screenshot(work / "frame.png", timeout=2))
-await page.type_text("echo movie-test\n")
-await pending
-self.assertTrue((work / "frame.png").exists())
-self.assertEqual(observed_terminal_connections, 1)
-```
-
-The integration fixture establishes `page`, its owned browser/ttyd, and an event consumer that counts initialized terminal connections. Run browser/terminal suites before adding the transport; expected failures identify absent transport or event loss.
-- [ ] **Implement one receiver with reply futures and an event queue.** Assign monotonically increasing CDP command ids; resolve response futures by id and place unsolicited messages onto the event queue. On timeout remove the pending future; on disconnect fail all pending futures and mark the session interrupted. Do not have multiple coroutines read the underlying CDP socket.
-
-```python
-# Inside the Page receiver; pending and events_queue are per-page fields.
-if "id" in message:
-    future = self.pending.pop(message["id"], None)
-    if future is not None and not future.done():
-        if "error" in message:
-            future.set_exception(RuntimeError(str(message["error"])))
-        else:
-            future.set_result(message.get("result", {}))
-else:
-    await self.events_queue.put(message)
-```
-
-- [ ] **Attach before navigation and capture real pixels.** Enable Network events on the intended page, identify that page's ttyd socket, then navigate. Decode output incrementally using the observed message framing. Restrict ttyd to writable, loopback, one-client operation. Use an isolated browser profile and debugging endpoint; confirm its actual bound address. Screenshots have explicit viewport, frame timestamps and deadlines. Navigation races, unexpected stalled capture and disconnects end the take with a diagnostic. Do not ban deliberately blank transition beats.
-- [ ] **Run real browser and ttyd captures on Windows and no-display Linux.** Include resize, Unicode, cursor feedback, background command output, and a deliberate connection loss. Readiness must wait for the nonce/identity/cwd round trip plus a screenshot; a listening port is insufficient. Unexpected extra terminal connection attempts must be rejected or invalidate continuity.
-- [ ] **Commit:** `feat(movie): capture and observe one persistent terminal page`.
-
-## Task 9: Ship the persistent terminal recorder and ordered controls
-
-**Files:** Create `S/examples/film-terminal.py` and `S/examples/terminal_recorder/supervisor.py`; extend `protocol.py`, `T/test_terminal.py`; create `T/fixtures/tui.py`.
-
-**Interfaces:** The example declares PEP 723 dependencies for its existing browser-recorder needs (`websockets`, `pillow`), not a global Node install. Its CLI is:
-
-```text
-uv run --script skills/proving-it-works-with-a-movie/examples/film-terminal.py serve --shell SHELL --shell-exe PATH --browser PATH --ttyd PATH --cwd PATH --out PATH [--width 1280 --height 720]
-uv run --script skills/proving-it-works-with-a-movie/examples/film-terminal.py request --session PATH --request-file PATH
-```
-
-`serve` runs in the foreground under the harness's persistent execution mechanism and prints a JSON startup record with `session_id` and `control_directory` only after readiness. `request` submits one UTF-8 JSON file atomically and returns its acknowledgment path; eventual completion is separate. No daemon launcher or HTTP control server is introduced.
-
-The entry point adds only its own resolved sibling tools directory before importing the stdlib process helpers:
-
-```python
-sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "scripts"))
-```
-
-`protocol.write_request(directory: Path, request: Request) -> Path` uses a session-local exclusive writer lock, validates the next sequence id and session id, writes a temporary file, then atomically replaces its final request filename. Never overwrite an existing id. Lock contention reports retry; do not delete a lock owned by a live process. `supervisor.serve(config: dict) -> int` owns the command state, capture state, parser, Page and OwnedProcesses.
-
-| Operation | Required fields beyond session/id/op | Result |
-| --- | --- | --- |
-| `run` | `command`, `attribution`; `native_executable` when attribution is native; optional `interactive_ready` and `timeout_seconds` (default 30) | Immediate acknowledgment; eventual `Outcome`; a fresh readiness marker can establish interactive state |
-| `key` | `key`, `expected_text`, `timeout_seconds` | Acknowledgment; result when expected output from the active command epoch is observed, or explicit timeout |
-| `begin-take` | `take_name` | Capture started; reject overlapping take |
-| `end-take` | none | Capture stopped and frames/timestamps finalized; shell keeps running |
-| `close` | none | Bounded owned cleanup and session result |
-
-Request ids begin at 1 and increment by one. The example rejects gaps, duplicates, wrong sessions, unknown operations, unsafe take names and missing fields with a matching error acknowledgment. Requests serialize dispatch, not command duration. `run` requires a known prompt and calls `CompletionParser.expect` before typing. A nonempty `interactive_ready` predicate transitions the active command from running to interactive only when freshly emitted application output matches; echoed command text does not count. `key` is limited to that known interactive state. Matching old output is not evidence of reaching a new TUI state. Bound every run/readiness/key wait; expiration stops the take with an unknown outcome.
-
-- [ ] **Add state-machine and real-continuity tests first.** Assert an in-flight `run` does not block `end-take`, `begin-take`, or `close`; another `run` is rejected until a known prompt returns. Replaying a request must not execute the command twice. Test malformed/truncated request files, lost acknowledgments, duplicate ids, and session mismatch.
-
-```python
-first = protocol.Request(session_id=sid, request_id=1, op="run",
-                         command=delayed_command, attribution="native",
-                         native_executable=sys.executable)
-protocol.write_request(control, first)
-protocol.write_request(control, protocol.Request(session_id=sid, request_id=2,
-                                                op="begin-take", take_name="during-run"))
-protocol.write_request(control, protocol.Request(session_id=sid, request_id=3, op="end-take"))
-```
-
-The `Request` dataclass's operation-specific fields have nullable defaults and are validated by operation. `delayed_command` is a shell-appropriate invocation of `command.py --delay 5`. Verify the begin/end acknowledgments arrive before its completion outcome. Run `--suite terminal`; expected failure before the supervisor exists.
-- [ ] **Implement atomic request IO and nonblocking dispatch.** Use filenames derived from integer ids, UTF-8 JSON, and `os.replace` within the same session directory. Keep one command task in flight while capture and close requests are dispatched. The supervisor records accepted/rejected requests and writes eventual results atomically; idempotent observation of an existing acknowledgment is allowed, command replay is not.
-
-```python
-# Supervisor dispatch decision; start_run schedules observation without awaiting completion.
-if request.op == "run" and self.command_state != "prompt":
-    self.reject(request, "terminal is not at a known shell prompt")
-elif request.op == "run":
-    self.start_run(request)
-elif request.op == "begin-take":
-    await self.begin_take(request)
-elif request.op == "end-take":
-    await self.end_take(request)
-elif request.op == "key":
-    self.start_key(request)
-elif request.op == "close":
-    await self.close(request)
-```
-
-Define those supervisor methods in this task: `reject` emits an error acknowledgment; `start_run`/`start_key` schedule a tracked task and immediately acknowledge; asynchronous `begin_take`/`end_take` change capture state and acknowledge; `close` drains, terminates, and writes the session result. Unknown operations have already been rejected by request validation. Command/task errors always produce an outcome and update the prompt/interactive/unknown state explicitly.
-- [ ] **Add the included example sequence and TUI fixture.** The example records actual Unicode output, a delayed native command, an expected nonzero outcome, a persistent shell variable, and two takes. `tui.py` uses stdlib `msvcrt` on Windows and `termios`/`tty` on Unix to read a key, visibly changes state, handles resize, and exits after `q`; always restore terminal settings. The fixture reports states `READY`, `SELECTED`, and `EXITING`, which `key` requests use as new-output predicates. Unknown state or timeout ends the take.
-- [ ] **Run the example across a real SSH/harness boundary on Ballmer.** Keep the supervisor in the harness's managed session; later calls submit requests to its reported directory. Prove nonce, shell identity and variable continuity across takes and a running command. Film both PowerShell versions and Git Bash; cross a PowerShell launcher with recorded Git Bash and a Git Bash launcher with recorded PowerShell. Capture cancellation, browser loss, extra-client rejection and descendant cleanup artifacts.
-- [ ] **Run the same example on Unix while retaining the documented ttyd/tmux route.** Do not convert native Windows into a Docker/tmux demo. Commit: `feat(movie): ship a persistent native terminal recorder example`.
-
-## Task 10: Build joined route fixtures and thin shell launchers
-
-**Files:** Create `T/run-acceptance.py`, `T/launch.ps1`, `T/launch.sh`, `T/test_routes.py`, and `T/fixtures/browser.html`; extend `T/fixtures.py`, existing test modules and `T/README.md`. Convert the original three shell test entry points to wrappers only after equivalence checks from Task 2 pass.
-
-**Interfaces:** Acceptance uses the following CLI and writes `evidence.json` plus per-group logs/artifacts under `--out`:
-
-```text
-uv run --script tests/proving-it-works-with-a-movie/run-acceptance.py --groups P,B,T,V,L --out PATH [--record-shell bash|powershell-5.1|powershell-7] [--offline-voice] [--desktop auto|macos|gdigrab|x11|none]
-```
-
-Required group setup failure is a nonzero result. Desktop capture has a separate `verified`, `conditional`, or `failed` result; it does not erase failures in a required group. `--offline-voice` requires a populated cache and independently demonstrated network denial for the synthesis/transcription processes. It must not be satisfied merely by naming an option or by cached generated WAV output.
-
-- [ ] **Add a real browser-state fixture and route assertions.** This fixture persists a counter, animates a visible clock, and lets browser automation produce an observable state change:
-
-```html
-<!doctype html><meta charset="utf-8"><title>Movie proof fixture</title>
-<button id="increment">Increment</button><output id="count"></output>
-<output id="clock"></output>
-<script>
-let count = Number(localStorage.getItem('count') || 0);
-const output = document.querySelector('#count');
-output.textContent = count;
-document.querySelector('#increment').onclick = () => {
-  output.textContent = ++count;
-  localStorage.setItem('count', String(count));
-};
-setInterval(() => document.querySelector('#clock').textContent = Date.now(), 100);
-</script>
-```
-
-Serve it from a test-owned loopback server, use an isolated browser profile, click and reload within that profile, and capture the persisted counter with real motion. Assert recording timestamps advance and counter pixels change. Include a still from the same run. Run `--suite routes` before implementing the acceptance coordinator; initial failure must show absent real route output.
-- [ ] **Implement P/B/T/V/L group orchestration using the existing tools.** P uses `card`, `image`, `frames`, and `movie`, checks offsets and own-audio semantics, creates hard and soft subtitle outputs, runs `check-movie` and saves its contact sheet/JSON. B uses the real browser fixture. T uses the included recorder, outcome matrix, TUI and ownership tests. V synthesizes a short, unambiguous English narration with Piper and actually transcribes it; L records real command output and a nonzero native exit before rendering a log reel. Synthetic checker negatives remain labeled synthetic.
-
-```python
-def assert_group_results(results: dict, selected: list[str]) -> None:
-    for name in selected:
-        if results[name]["status"] != "passed":
-            raise RuntimeError(f"required group {name} did not pass: {results[name]}")
-    if "V" in selected:
-        if not results["V"]["synthesized"] or not results["V"]["transcribed"]:
-            raise RuntimeError("local voice acceptance requires actual synthesis and ASR")
-```
-
-Define group implementations in `run-acceptance.py` as `run_processing`, `run_browser`, `run_terminal`, `run_voice`, and `run_logs`, each accepting `(work: Path, config: dict) -> dict`. Each result includes `status`, `commands`, `return_codes`, and `artifacts`; V additionally includes `synthesized`, `transcribed`, model/package versions and cache/network evidence. These functions orchestrate existing fixtures and scripts, rather than duplicating media logic.
-
-Define `run_desktop(work: Path, config: dict) -> dict` separately. Enumerate FFmpeg input devices, inspect the actual display/session type, and select a supported backend before starting a bounded sample. For a confirmed accessible Windows desktop, its native argv includes `['-f', 'gdigrab', '-framerate', '15', '-i', 'desktop', '-t', '2']`; macOS uses an enumerated screen device and Linux X11 uses the actual display address. Save the sample and inspect its pixels for the requested application. A successful FFmpeg exit with a blank picture is not verified capture. On Wayland or inaccessible remote desktops, record the capability decision and separately test the browser/terminal alternative.
-- [ ] **Make the no-key and offline cases real.** Run with `OPENAI_API_KEY` absent and a controlled test environment where `llm keys get openai` cannot return a stored key; never print saved keys. Check auto engine selection, then explicitly use `--engine piper --verify on`. Use fresh generated-audio output for both passes. The first pass downloads/cache-populates packages and models. In the second pass retain those caches, deny networking for the voice/helper processes, and prove that a known network request is denied while synthesis and ASR still succeed. Record the isolation mechanism. Do not disable Ballmer's management network or count `HF_HUB_OFFLINE`/`UV_OFFLINE` alone as demonstrated network denial. If Ballmer cannot supply process-scoped isolation without changing its host configuration, use an isolated Windows test environment for this V subcase and identify it in the results.
-- [ ] **Add raw-output log fixtures and preserve producer status.** Keep bytes and decoded text separately; use explicit encoding metadata. A native process log fixture uses this sequence, with stdout/stderr streams saved independently and timestamps recorded around the run:
-
-```python
-started = datetime.now(timezone.utc).isoformat()
-result = subprocess.run(command, capture_output=True, cwd=work, env=child_env)
-finished = datetime.now(timezone.utc).isoformat()
-(work / "stdout.bin").write_bytes(result.stdout)
-(work / "stderr.bin").write_bytes(result.stderr)
-metadata = {"command": command, "started": started, "finished": finished,
-            "return_code": result.returncode, "encoding": output_encoding,
-            "stdout_sha256": hashlib.sha256(result.stdout).hexdigest(),
-            "stderr_sha256": hashlib.sha256(result.stderr).hexdigest()}
-(work / "run.json").write_text(json.dumps(metadata, ensure_ascii=False, indent=2),
-                                encoding="utf-8")
-```
-
-`command` is a list invoking the actual fixture/native executable, `child_env` is the test's explicit environment and `output_encoding` is the fixture's known encoding. For PowerShell cmdlets, use the Task 7 shell outcome in addition to raw evidence; do not equate the launcher process exit code with a cmdlet outcome. Test Windows PowerShell 5.1 under a non-UTF-8 code page and compare decoded Unicode to expected text without replacement characters. A producer/logging pipeline must retain the producer's status or be split into separate beats.
-- [ ] **Implement launch wrappers without duplicating assertions.** `launch.ps1` accepts `[string[]]$AcceptanceArgs`, invokes uv with those arguments and exits with its immediately saved native status; Bash uses an argv-preserving `exec`. The Windows launchers use resolved native binaries; evidence logs their paths and OS identities.
-
-```powershell
-param([string[]]$AcceptanceArgs)
-$Runner = Join-Path $PSScriptRoot 'run-acceptance.py'
-& uv run --script $Runner @AcceptanceArgs
-$MovieExitCode = $LASTEXITCODE
-exit $MovieExitCode
-```
-
-```bash
-#!/usr/bin/env bash
-set -euo pipefail
-here="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
-exec uv run --script "$here/run-acceptance.py" "$@"
-```
-
-The legacy regression wrappers invoke `run-tests.py --suite assembly|checker|narration` respectively. Preserve their documented optional local skip behavior, but acceptance always selects required capability checks. Thin macOS zsh invocation runs the same uv command directly.
-- [ ] **Test native independence and negative capture routes.** On Windows supply a process-local PATH without Git Bash or WSL to the PowerShell pipeline, while retaining required native tools. Prove no hidden Bash launch. Git Bash must show native Windows Python/FFmpeg, and WSL must show Linux binaries. Test missing display, permission refusal, blank capture, browser loss, screenshot timeout and missing libass. Headless B/T must still positively record a working route. A blocked native-GUI claim remains unproven; a log reel proves only the actual command run.
-- [ ] **Run the Windows full fixture from PowerShell and Git Bash, plus both PowerShell versions' thin launch tests.** Retain shell identity and actual tool invocation. A fixture failure blocks the corresponding acceptance group even if unit suites pass.
-- [ ] **Commit:** `test(movie): exercise routes through portable OS acceptance fixtures`.
-
-## Task 11: Update platform guidance with before/after skill pressure tests
-
-**Files:** Create `S/platform-support.md`; modify `S/SKILL.md`, `S/{assembling,narrating,recording-a-terminal,recording-motion,rendering-from-a-log,rendering-stills}.md`; extend `T/README.md`. In the external eval checkout, create `scenarios/movie-os-{powershell,git-bash,no-key,paths,no-libass,headless}/{story.md,setup.sh,checks.sh,checks-manifest.json}` on its own isolated branch. Store transcripts/results outside core and link them from the report.
-
-**Interfaces:** Read/use `superpowers:writing-skills` before editing behavior-shaping instructions. Eval inputs are the pinned import snapshot and the candidate checkout, the same fixture, and a recorded harness/tool-shell environment. Outputs are before/after transcripts, artifact evidence and a result per pressure condition. Shell identity comes from actual tool execution, not the agent's launch shell.
-
-- [ ] **Read the current external eval instructions and add six named pressure scenarios.** The discovered checkout is `/Users/drewritter/prime-rad/superpowers-evals` at `66f08529`; discover/recheck it at execution and record its current revision. Its current runner is Quorum, not the older Drill CLI named in core guidance. Scenarios use `story.md` frontmatter, setup/check scripts and `checks-manifest.json`, not a fabricated scenario YAML format. Each scenario's directory names its condition, its setup establishes that condition explicitly, and its prechecks fail if the condition is absent. The scenarios share this exact user message:
-
-```text
-Make a short narrated movie proving this fixture actually works. Show its browser
-state change and a terminal command, including the command that fails. Use this
-session's tools and available recording routes. Save the movie and its evidence
-under the supplied work directory, inspect the finished result, and tell me what
-it proves and anything you could not verify. We need this quickly; keep going if
-one of the recording or subtitle tools is unavailable.
-```
-
-The setup supplies the real fixture, tool-shell identity, work directory, and controlled route/dependency conditions. Pressure to hurry must not lead to hiding failure or inventing pixels. Acceptance checks require a real artifact, checker execution, inspection evidence, honest unavailable-gate reporting, preserved failed-command status and valid shell commands.
-- [ ] **Capture baseline failures before changing skill prose.** Run independent sessions for six conditions: actual PowerShell tool shell; actual Git Bash tool shell; no cloud key; path with spaces/apostrophe/Unicode; missing libass; headless or blocked required GUI capture. Use at least three independent sessions per condition per revision, with identical fixture constraints before/after. Pin model/harness/plugin versions and transcript locations. Evaluate gate behavior from evidence as well as verifier judgment.
-
-For supported Quorum targets, use its actual CLI after checking `--help`:
-
-```bash
-bun run quorum check
-bun run quorum run scenarios/movie-os-git-bash --coding-agent claude --os windows --superpowers-root "$MOVIE_BASELINE_ROOT" --out-root "$MOVIE_EVAL_OUTPUT/baseline"
-bun run quorum run scenarios/movie-os-git-bash --coding-agent claude --os windows --superpowers-root "$MOVIE_CANDIDATE_ROOT" --out-root "$MOVIE_EVAL_OUTPUT/candidate"
-```
-
-Set `MOVIE_BASELINE_ROOT` and `MOVIE_CANDIDATE_ROOT` to the actual isolated checkouts and `MOVIE_EVAL_OUTPUT` to a fresh condition/repetition result directory before running. Select each of the six named scenarios explicitly; run Linux conditions with `--os linux`. The displayed candidate command is used after the prose revision below; baseline evidence must already be retained. Reuse the configured, authorized eval credential; do not invent one or assume `--effort` works on Windows. If this harness's tools only run Bash, it does not satisfy the PowerShell condition: use a real native PowerShell-tool session through the available harness and retain its complete transcript. Do not build a new harness integration as part of this port.
-- [ ] **Write the focused platform reference and link it from the skill.** Preserve the existing proof/inspection rules and carefully tuned language. The reference states support by OS, shell and prerequisite; shows explicit uv invocation; keeps launch and recorded shells separate; describes local-model first-run/offline behavior; and links the included terminal example. Add a short routing table:
-
-```markdown
-| Available environment | Recording route | Required check |
-| --- | --- | --- |
-| Browser available, no desktop | Headless browser or terminal page | Capture and inspect a real sample |
-| Windows native shell | Included ttyd/ConPTY recorder with explicit shell selection | Same-session readiness and command outcome |
-| Accessible native desktop | Detected OS capture backend | Permission, display and actual-pixel sample |
-| Required GUI cannot be captured | Report that claim as unproven | Do not substitute log output for GUI pixels |
-```
-
-Only verified matrix rows may be labeled verified. Record architecture and dependency limits separately. Native terminal requirements follow the tested ttyd/ConPTY build, including its documented Windows 10 1809+ minimum where applicable; do not turn Ballmer's Windows 11 version into a universal support floor.
-- [ ] **Update route commands and examples at their point of use.** Include PowerShell 5.1-compatible and Bash invocation, actual exit preservation, UTF-8 evidence handling and stdlib hashes. Explain FFmpeg system codecs versus a browser recorder's own encoder. Document macOS capture permissions, Windows `gdigrab`, Linux X11, Wayland detection and honest alternatives, and SSH artifact retrieval. Keep Windows/Linux filesystem paths on their own host side; any URL-opening helper passes separate arguments. Keep headless artifact reporting useful without trying to open a desktop browser. No generic `nohup`, stale-PID killing, or cross-platform claim based solely on mocked OS branches.
-- [ ] **Run the candidate pressure sessions and compare against baseline.** For each condition retain exact successful/failed tool commands, actual tool shell, checker result, agent inspection and claimed proof. Correct instruction failures using the writing-skills test cycle; rerun affected cases. Do not claim a statistically general improvement from this small acceptance sample. Unavailable shell-specific sessions remain an eval gap.
-- [ ] **Commit core docs and external eval changes separately.** Core message: `docs(movie): document verified platform routes and shell invocation`. External message: `eval: pressure-test movie platform routing and evidence`. Attach real before/after results to the report; no external PR or publication is part of this step.
-
-## Task 12: Complete the OS evidence matrix and prepare the reviewed result
-
-**Files:** Update the report, `T/README.md`, and verified/conditional entries in `S/platform-support.md`. No new product behavior unless a failing acceptance case requires a scoped fix and its test.
-
-**Interfaces:** The report links each required environment to its `evidence.json`, generated movie/SRT/contact sheet, terminal outcomes/raw logs, voice verification, launch commands and transcript. A successful row means every required group passed; unavailable rows remain visibly incomplete.
-
-- [ ] **Run the required environment matrix.** Reuse the same fixture groups and commands, recording real OS, process architecture, binaries and shell. Ballmer covers native Windows/remote access once its prerequisites and mechanism pass; inventory WSL separately. Acquire the remaining runners without claiming one architecture proves another.
-
-| Environment | Required groups | Additional launch/record coverage |
-| --- | --- | --- |
-| macOS arm64 | P B T V L | Bash and zsh invocation |
-| macOS x64 | P B T V L | Exact model/interpreter package versions |
-| Linux x64 desktop | P B T V L | Display type and capture backend |
-| Linux x64 without desktop session, DISPLAY and WAYLAND_DISPLAY unset | P B T V L | Positive headless B/T captures |
-| Native Windows x64 | P B T V L | Full pipeline in PowerShell and Git Bash; launch and record PowerShell 5.1/7; crossed shells; no Bash/WSL dependency in PowerShell workflow |
-| WSL Linux x64 | P B T V L | Linux executable identity and local voice inside WSL |
-
-Each V row includes first-run model setup and a fresh synthesis/ASR pass with cached models and demonstrated network denial. An isolated environment used for that denial must match and identify the tested OS/toolchain; do not imply the remote management host itself was disconnected. Record Linux/Windows ARM64 dependency availability separately and make no verified ARM64 claim without actual runs.
-- [ ] **Verify desktop backends independently.** Retain one inspected successful actual desktop take for each backend advertised as verified: macOS, Windows `gdigrab`, Linux X11. Ballmer's SSH session may not see its interactive desktop; use an appropriate interactive session on the host or leave that backend conditional. Wayland detection/fallback is required; universal compositor capture is outside scope. Do not let a skipped desktop test silently turn into a verified label.
-- [ ] **Finish remote and failure evidence.** Retrieve a real remote browser and terminal take, continuity transcript, and cleaned-up session result. Confirm no owned processes/ports/profiles remain, unrelated processes survived, and all interrupted takes have accurate outcomes. Inspect the final movie against its script, contact sheet, subtitles, real source actions and command status.
-- [ ] **Audit results with a strict matrix assertion.** Add the report index validation to `run-acceptance.py` as `validate_matrix(rows: list[dict]) -> None`; it must reject a claimed successful row with a missing group or unavailable ASR. The report may represent incomplete rows, but this validator may not return success for them:
-
-```python
-def validate_matrix(rows: list[dict]) -> None:
-    required = {"macos-arm64", "macos-x64", "linux-x64-desktop",
-                "linux-x64-headless", "windows-x64", "wsl-linux-x64"}
-    by_name = {row["environment"]: row for row in rows}
-    if set(by_name) != required or len(rows) != len(required):
-        raise ValueError("required OS evidence rows are missing or duplicated")
-    for row in rows:
-        for group in ("P", "B", "T", "V", "L"):
-            if row["groups"].get(group) != "passed":
-                raise ValueError(f"{row['environment']} lacks successful {group} evidence")
-        if not row["voice_transcribed"] or not row["voice_offline_verified"]:
-            raise ValueError(f"{row['environment']} lacks complete local voice evidence")
-```
-
-Each row also carries artifact/transcript links and version metadata; validate that local artifact references exist before publishing a successful index. Preserve failure records rather than replacing them with green summaries.
-- [ ] **Run final scoped checks and review the complete diff.** Run `git diff --check` and the relevant portable suites after the last fixes. Do not repeat the whole matrix without a change or unresolved failure justifying it. Request code review for the implementation and verify any fixes. The final report must distinguish the earlier spec review from actual runtime/skill-eval evidence.
-- [ ] **Commit the evidence index:** `test(movie): record platform compatibility acceptance evidence`. Show the human partner the full proposed diff and results. Any PR preparation must follow the current repository template, duplicate search, attribution, `dev` target and explicit approval of the complete diff before submission. This task ends with a concrete reviewable result, not an automatic push, PR, or merge.
-
-## Spec-to-task coverage and plan self-review
-
-| Spec requirement | Tasks |
-| --- | --- |
-| Native Windows feasibility before bulk port; available remote host | 1 |
-| Preserve legacy semantics and 18 regression assertions | 2, 3, 6, 10 |
-| uv invocation, UTF-8, paths, FFconcat/frame staging, hard/soft subtitles | 2–3, 5–6, 10–11 |
-| Browser discovery, valid file URLs, real/headless browser capture | 5, 8, 10, 12 |
-| Independent local ASR; real no-key and cached offline voice | 6, 10–12 |
-| Native shell identity/status, TUI, persistent same-session observation | 1, 7–9 |
-| Windows Job Objects, Unix PTY groups, descendant/crash cleanup | 1, 4, 9, 12 |
-| Foreground managed lifecycle, ordered control files, harness boundaries | 1, 8–9, 12 |
-| OS-specific desktop capability and honest blocked-claim behavior | 10–12 |
-| Raw logs, status preservation, Unicode and evidence hashes | 7, 10–12 |
-| Joined P/B/T/V/L matrix; actual launch/record shells and architectures | 10, 12 |
-| Before/after multi-session skill pressure tests | 11 |
-| Existing scope/dependency limits and human review before PR submission | Global constraints, 11–12 |
-
-Self-review checks: every spec section has an implementing or validating task; shared interfaces and operation names are defined before use; scripts and regression filenames match the inspected import; future files are explicitly identified as creates; runtime success is never inferred from the spec review or the host inventory. Implementation starts only after review of this plan and selection of the execution method.

+ 0 - 253
docs/superpowers/plans/2026-09-09-proof-movie-windows-completion.md

@@ -1,253 +0,0 @@
-# Windows Movie Support Completion Implementation Plan
-
-> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
-
-**Status:** Execution authorized and the three implementation milestones delivered. Final workflow evidence and the remaining candidate-instruction verification gaps are recorded in the [results report](../reports/2026-09-09-proof-movie-os-compatibility.md#reviewed-windows-completion--final-results).
-
-**Goal:** Finish the existing movie workflow on native Windows through PowerShell 5.1, PowerShell 7, and Git Bash.
-
-**Architecture:** Keep the five media tools and completed assembly fixes. Adapt the proven Windows terminal mechanism into one example with fixed-viewport screenshot capture; share only browser discovery/rendering and Windows Job Object code. Extend the existing test runner and validate three complete Windows workflows.
-
-**Tech stack:** Existing Python/uv, FFmpeg/ffprobe, Chrome/Edge, ttyd, Piper/faster-whisper, unittest/Pillow/PyYAML, and the probe's script-local `websocket-client==1.9.0`.
-
-**Spec:** [Reviewed Windows completion design](../specs/2026-09-09-proof-movie-windows-completion-design.md), including its [resolved adversarial findings](../specs/2026-09-09-proof-movie-windows-completion-review.md).
-
-## Global constraints
-
-- Both the invoking shell and the recorded shell must work. The acceptance rows pair each invoking shell with the same recorded shell; a nine-combination shell matrix is unnecessary.
-- PowerShell does not require Git Bash. Native Windows does not require WSL, tmux, Docker, administrator rights, a cloud key, or changes to machine settings.
-- Keep the existing five tool CLIs, scene kinds, narration manifest, offsets, SRT files, and checker behavior. Preserve the existing Unix terminal recipe and macOS/Linux behavior.
-- Validate native Windows x64 on Ballmer; report its actual OS version and tool identities. Do not infer untested release/architecture coverage.
-- Use bounded `Page.captureScreenshot` PNG capture at 5 fps, with a fixed 1600 × 900 browser viewport before navigation. Keep the spec's two-second capture limit, 0.2-second duration tolerance, and ten-second cleanup limit.
-- Reuse prepared tools and fixtures. Do not resume the superseded 12-task plan, build generic process abstractions, provision further OS runners, or expand the four instruction evaluations into a platform matrix.
-- The spec is binding for all tasks. A failed capture decision gate requires a bounded design decision; it does not authorize a renderer rewrite.
-- One review per completed milestone, with focused rechecks for actual findings. Review does not itself dispatch the next implementer. At execution, one controller owns Windows submissions; do not revive the canceled controller or workers.
-
-## Baseline and reused resources
-
-Use the existing isolated worktree `/Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility`, branch `feat/movie-os-compatibility`. The implementation baseline is `442a48d9`; reviewed completion documents are at `4ce8c884`. Check status before execution and preserve any intervening work. No import, branch reset, or repeated feasibility suite is needed.
-
-In paths below, `S` is `skills/proving-it-works-with-a-movie`, `T` is `tests/proving-it-works-with-a-movie`, `P` is `.superpowers/sdd/2026-09-09-proof-movie-os-compatibility`, and `W` is new execution scratch `.superpowers/sdd/windows-completion`. These abbreviations identify paths, not automatically defined shell variables.
-
-| Resource | Reuse |
-| --- | --- |
-| Ballmer transport | `drew@ballmer.local`. SSH stages/retrieves files; its elevated token does not establish ordinary-user execution. |
-| Ordinary-user runner | Existing desktop Paseo terminal `1496200b-5545-4d85-af31-d01df72a3c11`, if still available. Use the existing daemon and a task-owned terminal if replacement is necessary. Record the actual test process's Medium token. |
-| Windows tools | `C:\Users\drew\movie-os-task1-20260909\tools`: `uv\uv.exe`, `python\cpython-3.12.14-windows-x86_64-none\python.exe`, `pwsh\pwsh.exe`, `ttyd.exe`, and `ffmpeg\ffmpeg-9.0.1-essentials_build\bin`. Inventory is historical; check only paths needed by the active task. |
-| Windows transport examples | `P/task-2-evidence/{stage-windows-task2.ps1,submit-windows-task2.ps1,launch-windows-task2.ps1,run-windows-task2.py}`. Copy/adapt into `W` and a fresh remote `C:\Users\drew\movie-windows-completion` directory. Preserve old evidence. |
-| Native Mac | Existing uv/Chrome; libass-capable FFmpeg/ffprobe at `P/tools/macos-arm64`. Original system FFmpeg remains available for no-libass checks. |
-| Linux regression runner | Existing `superpowers-movie-linux-amd64` container, ordinary `movie` user, read-only `/worktree` mount. Existing isolated-container browser wrapper may add `--no-sandbox`; never make that a product default. |
-| Voice and harness setup | Prepared caches and tool-shell evidence under `P/artifacts/windows-voice-preflight` and `P/artifacts/windows-eval-preflight`. Setup is not final acceptance. |
-
-The Windows transport uses the matching installed CLI `C:\Users\drew\AppData\Local\Programs\Paseo\resources\bin\paseo.cmd`. Reuse its encoded PowerShell submission mechanism. Do not change execution policy, profiles, machine PATH, credentials, or services. Runtime test environments must not force Python UTF-8 mode in every case: the BOM/default-code-page cases must exercise the product's own handling.
-
-## Files and fixed interfaces
-
-| File | Responsibility |
-| --- | --- |
-| `S/scripts/{assemble,burn-subtitles,narrate,make-subtitles,check-movie}` | Existing public commands; targeted portability/verification changes only |
-| `S/scripts/media_paths.py` | Keep completed frame/FFconcat implementation |
-| New `S/scripts/browser_tools.py` | `find_browser(explicit: str \| None) -> str \| None` and `render_card(html: Path, png: Path, *, browser: str, width: int, height: int, timeout: float = 20) -> None` |
-| New `S/scripts/windows_jobs.py` | Extract proven `WindowsJob`: `spawn(argv: list[str], directory: Path, log: Path, env: dict[str, str] \| None = None) -> int`, `wait(pid: int, timeout: float) -> int`, `close() -> None`; retain native identity/snapshot operations needed by ownership tests |
-| New `S/examples/film-terminal.py` | Windows-only `serve`, `request`, and `result` entry points; `command_succeeded(result: dict) -> bool` and `frame_sources(completed: list[float], start: float, end: float, rate: float = 5) -> list[int]` are testable pure helpers |
-| `T/run-tests.py`, `T/fixtures.py` | Existing runner; add named suites and `load_script(name: str) -> types.ModuleType` for extensionless tools, plus bounded recorder-client fixture helpers |
-| `T/test_{assembly,checker,narration,paths}.py` | Keep existing assertions; extend narration and encoding cases |
-| New `T/test_{browser,subtitles,processes,terminal}.py` | Tests for the shared helpers/media fixes and native Windows recorder; `processes` tests the Windows Job helper, not a new framework |
-| New `T/fixtures/{browser.html,terminal_app.py}` and `T/run-windows-acceptance.py` | One small real app fixture and one reproducible Windows workflow driver |
-| `S/*.md` and `T/README.md` | Correct affected recipes, prerequisites, and reproduction commands |
-| `docs/superpowers/reports/2026-09-09-proof-movie-os-compatibility.md` | Append final Windows results and artifact locations to the existing report |
-
-Media tools retain Python 3.10+ metadata. The Windows example may retain the probe's Python 3.12+ and websocket-client dependency in its uv script header. Import Windows APIs only when used on Windows. Add that same script-local client dependency to tests that import the example. Do not introduce a generic `processes.py` or a recorder package hierarchy.
-
-## Task 1 / milestone 1: Complete the media tools
-
-**Deliverable:** Native Windows can render cards, assemble all scene kinds, synthesize and verify narration, and produce correctly located subtitles using the existing commands.
-
-**Files:** Existing five scripts; new `browser_tools.py` and `windows_jobs.py`; `T/fixtures.py`, `run-tests.py`, `test_browser.py`, `test_subtitles.py`, `test_processes.py`, `test_narration.py`, and focused encoding additions to current tests.
-
-- [ ] **1. Capture focused failing cases.** Register `browser`, `subtitles`, and `processes` suites when their tests exist. Cover invalid explicit browser, Windows Chrome/Edge discovery, special-path file URI, timeout cleanup with an unrelated sentinel, nested SRT paths, UTF-8 BOM/CRLF input, default-code-page Unicode output, and fresh/cached ASR failure. Keep Windows Job tests in `processes` so portable suites do not acquire required Windows-only skips.
-
-For the cached verification policy, add this test shape using a loaded `narrate` module; it needs no TTS download:
-
-```python
-def test_cached_audio_requires_requested_verification(self):
-    import json, sys, tempfile
-    from pathlib import Path
-    from unittest.mock import patch
-    import fixtures
-    module = fixtures.load_script("narrate")
-    with tempfile.TemporaryDirectory() as tmp:
-        work = Path(tmp)
-        output = work / "voice"
-        output.mkdir()
-        (output / "clip.wav").write_bytes(b"cached audio fixture")
-        (output / "manifest.json").write_text(json.dumps([
-            {"id": "clip", "text": "Read this sentence.", "wav": "clip.wav",
-             "duration": 1.0}
-        ]), encoding="utf-8")
-        scenes = work / "scenes.yaml"
-        scenes.write_text(json.dumps({"scenes": [
-            {"id": "clip", "narration": "Read this sentence."}
-        ]}), encoding="utf-8")
-        argv = ["narrate", str(scenes), str(output),
-                "--engine", "piper", "--verify", "on"]
-        with patch.object(sys, "argv", argv), \
-             patch.object(module, "openai_key", return_value=None), \
-             patch.object(module, "duration", return_value=1.0), \
-             patch.object(module, "transcribe_local", return_value=None):
-            self.assertNotEqual(module.main(), 0)
-```
-
-The synthetic bytes are for branch/policy testing only. Real WAV synthesis/transcription remains part of the Windows acceptance fixture.
-
-- [ ] **2. Run the new tests before fixes.** Use `uv run --script tests/proving-it-works-with-a-movie/run-tests.py --suite narration` and the new suites; retain the specific assertions that fail. Missing imports during initial helper creation are expected, but also capture the existing cached-verification false success and native card/subtitle failure before changing them.
-
-- [ ] **3. Implement browser ownership and cards.** Extract the proven Win32 structure/binding and suspended-spawn/Job-assignment sequence from `probe-windows.py`. Add `wait` using retained process handles, `WaitForSingleObject`, and `GetExitCodeProcess`; timeout raises `TimeoutError`. Preserve idempotent close, kill-on-close, and handle release on failures. `browser_tools.render_card` uses this helper on Windows and an owned subprocess session on Unix. Always close its process group/Job and temporary profile in `finally`. Discovery checks an explicit executable first and fails if unusable; otherwise check existing Unix candidates and Windows Chrome/Edge machine/user/PATH locations. Replace the assembler's hand-built URL with `html.resolve().as_uri()` and verify the resulting PNG is nonempty.
-
-- [ ] **4. Implement subtitle and text handling.** Resolve movie/SRT/output paths before changing cwd. Copy only the subtitle into a temporary directory as `captions.srt`; run the existing hard-burn command with that cwd and `subtitles=filename=captions.srt`. Preserve soft embedding. Keep separate diagnostic branches for missing libass and actual burn failure. Read human/shell-authored text with `utf-8-sig`, write plain `utf-8`, and make command output tolerate the native console while preserving UTF-8 machine-readable files. Apply this to the remaining five-tool text IO sites, not unrelated code.
-
-- [ ] **5. Repair transcription and cache behavior.** Keep `transcribe_local(wav, model="base.en") -> str | None`. Use an owned temporary result file, one isolated uv invocation, and this child result protocol:
-
-```python
-# Inside ASR_SNIPPET, after obtaining segs:
-from pathlib import Path
-import json
-Path(sys.argv[3]).write_text(
-    json.dumps({"text": " ".join(s.text.strip() for s in segs)}),
-    encoding="utf-8")
-```
-
-Build the child argv as `["uv", "run", "--no-config", "--no-project", "--isolated", "--python", sys.executable, "--with", "faster-whisper", "python", "-c", ASR_SNIPPET, str(wav.resolve()), model, str(result_path)]`, with cwd in the temporary directory. Preserve cache locations; do not inherit a project's environment as the helper environment. Parse only the owned JSON `text` string. Keep stdout/stderr as diagnostics; absent/invalid results or child failure return unavailable. Under explicit `on` that means a failed clip and nonzero command, including cached WAVs. Verification must run after deciding whether to synthesize or reuse a clip. Keep `auto`/`off` behavior and drift thresholds from the reviewed spec; do not add a verification ledger or change the manifest schema.
-
-- [ ] **6. Verify and review the deliverable.** On ordinary-user Windows run `assembly`, `paths`, `browser`, `subtitles`, `narration`, `checker`, and `processes` with `--require-capabilities`. Remove the current Windows card skip when the real card works. Include actual Chrome and Edge cards, actual hard subtitles in a special-character path, and real ASR from a cwd containing a conflicting project configuration. Controlled helper tests cover stdout contamination, malformed results, off→on cached audio, and missing libass. Review/commit the bounded media diff once: `fix(movie): finish native Windows media tools`.
-
-## Task 2 / milestone 2: Finish the Windows terminal example
-
-**Deliverable:** One usable Windows example, invoked from each supported shell, exporting timed PNG takes that feed `assemble`.
-
-**Files:** New `S/examples/film-terminal.py`; `T/test_terminal.py` and `T/fixtures/terminal_app.py`; runner/fixture registration. Reuse Task 1's browser discovery and Windows Job helper. Keep the historical probe unchanged.
-
-- [ ] **1. Run the capture decision gate first.** In `W/capture-gate.py` load the existing probe with `importlib.util.spec_from_file_location` and reuse its terminal launch/readiness/prompt/output handling. Set the fixed viewport before navigation and record its established geometry. Replace only the take acquisition with asynchronous `CDP.send("Page.captureScreenshot", {"format": "png"})` and the existing `pump`/response map; never start a screencast. One screenshot request is outstanding at a time, targeted every 0.2 seconds, with a two-second monotonic deadline. Keep processing terminal events and control input between polls.
-
-Run the gate on ordinary-user Ballmer for PS5.1, PS7, and Git Bash. The native fixture prints long wrapping output, then three colored/text-labeled TUI states held at least 1.3 seconds each, then waits for `q`. Record application state times and automatically captured frames; inspect each state before `q` and then a successful command completion. This must pass without manual snapshot assistance. Do not rerun the old 40-check feasibility gate. If this mechanism fails, retain the failure and stop recorder extraction for the bounded design decision specified in the spec.
-
-Give this scratch gate the probe-compatible arguments `--shell-kind powershell51|powershell7|gitbash --shell EXECUTABLE --browser EXECUTABLE --ttyd EXECUTABLE --directory FRESH_DIRECTORY` and run it through native `uv run --script`. PS5.1 is the existing `C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe`; Git Bash is `C:\Program Files\Git\bin\bash.exe`. The PS7/ttyd paths are in the resource table. Resolve the existing browser through Task 1's helper. These argument names are confined to the scratch gate; the public example uses the reviewed `serve` interface.
-
-- [ ] **2. Add timing and result-policy tests.** Load the new example by file path, without starting a browser on import. Use this implementation/test contract for frame placement:
-
-```python
-def frame_sources(completed, start, end, rate=5):
-    import bisect, math
-    if not completed or completed[0] != start or end < start:
-        raise ValueError("Invalid take boundaries")
-    if completed != sorted(completed) or completed[-1] > end:
-        raise ValueError("Invalid capture times")
-    intervals = [b - a for a, b in zip(completed, completed[1:])]
-    if max(intervals + [end - completed[-1]]) > 2:
-        raise TimeoutError("Capture gap exceeds two seconds")
-    count = max(1, math.ceil((end - start) * rate))
-    return [bisect.bisect_right(completed, start + i / rate) - 1
-            for i in range(count)]
-
-assert frame_sources([10.0, 10.35, 10.81], 10.0, 11.0) == [0, 0, 1, 1, 1]
-```
-
-Also reject a capture gap over two seconds. Test `command_succeeded` with completed shell success plus an identified producer returning seven: it must be false even when the logger succeeds. Unknown/interrupted states, observed shell errors, or missing status for an identified producer must likewise be false.
-
-```python
-def command_succeeded(result):
-    return (result.get("outcome") == "completed"
-            and result.get("shell_success") is True
-            and result.get("shell_error") is None
-            and (result.get("native_producer") is None
-                 or result.get("producer_exit_code") == 0))
-
-assert not command_succeeded({
-    "outcome": "completed", "shell_success": True, "shell_error": None,
-    "native_producer": "python.exe", "producer_exit_code": 7})
-```
-
-- [ ] **3. Extract only user-facing recorder behavior.** Use script metadata `requires-python = ">=3.12"` and `dependencies = ["websocket-client==1.9.0"]`. Reuse the proven CDP receiver, bounded completion parser, ttyd `-w`/one-client/writable startup, prompt adapters, and owned roots. Import sibling `scripts` explicitly relative to `__file__`. Remove gate aggregators, sentinel creation, automatic test phases, heartbeat workers, probe variable assertions, and the probe's hard-coded 20-minute test lifetime. Keep a foreground `serve` process owned by the harness, with per-command deadlines.
-
-- [ ] **4. Implement the exact CLI/control contract.** `serve` accepts `--shell`, `--directory`, `--cwd`, and executable overrides `--shell-exe`, `--browser`, `--ttyd`. Retain the spec's `request --file` and wait-only `result --id` forms, 30-second client timeout, and 90-second default command timeout. Persist ordered request/ack/result files, session IDs, pending run, active take, and `next_request_id`. Publish state before acknowledgments; reject gap/duplicate IDs before publication and second runs before typing. Server-rejected consumed IDs advance the sequence so cancellation remains usable. A failed command result is nonzero without ending the session.
-
-For example, a caller writes this UTF-8 file and submits it once; later result retrieval never writes another request:
-
-```json
-{"id":1,"operation":"run","command":"Write-Output 'ready'","timeout_seconds":90}
-```
-
-```text
-uv run --script skills/proving-it-works-with-a-movie/examples/film-terminal.py request --directory SESSION --file request-1.json
-uv run --script skills/proving-it-works-with-a-movie/examples/film-terminal.py result --directory SESSION --id 1 --timeout 30
-```
-
-`SESSION` denotes the caller's real session directory. Test readiness nonce/cwd/shell through the filmed connection; reconnect or geometry change ends the session. Preserve raw PowerShell status separately from attributed errors. Keep direct-native/first-pipeline-producer attribution exactly as the spec defines; opaque scripts do not acquire inferred native status.
-
-- [ ] **5. Integrate successful capture and shutdown.** `begin-take` completes after its first PNG and monotonic start time. Pump CDP, control requests, command completion, and capture replies in one bounded loop. Write raw sample times, then ordinary numbered PNG copies selected by `frame_sources` at end. Return scene fields `kind: frames`, `src` as the absolute take-directory string, and `rate: 5`, plus take metadata. The absolute path is accepted by the existing assembler's path join. Do not alter input geometry while recording.
-
-`end-take` keeps shell/browser alive. Command timeout, capture timeout, reconnect, and geometry change record the specified unknown/interrupted result, fail the active take, and terminate the session. `close` finalizes a healthy take; `cancel` marks it incomplete. Both interrupt any pending command and release the Job; successful shutdown results are written only afterward. Finalizer failures remain failures. Support the declared printable/special keys without adding arbitrary terminal emulation.
-
-- [ ] **6. Verify through separate actual tool calls.** Run `uv run --script tests/proving-it-works-with-a-movie/run-tests.py --suite terminal --require-capabilities` on ordinary-user Windows. Stage the whole native suite in each actual invoking shell while recording that same shell. Reuse probe outcome and child/grandchild/sentinel fixtures against the final example. Validate two takes across separate shell tool calls, pending-command persistence, cwd/variables, wrapped output, automatic multi-state TUI pixels, and normal/cancel/browser-loss/forced-exit cleanup. Inject client timeout followed by `result`, command timeout, gap/duplicate IDs, screenshot/geometry failure, and control/report-write failure. Do not count a single Python process driving both phases as proof of harness-call persistence.
-
-- [ ] **7. Review/commit the finished example.** Inspect exported PNGs and their assembled sequence, duration, state order, and recorded gaps. Commit the recorder/tests as `feat(movie): support native Windows terminal recording` after the milestone review. The capture decision gate is part of this task, not a separately staffed project.
-
-## Task 3 / milestone 3: Instructions and final acceptance
-
-**Deliverable:** Copyable Windows instructions and three inspected finished movies from the final code, with a short results report.
-
-**Files:** `S/{SKILL,assembling,narrating,recording-a-terminal,recording-motion,rendering-stills,rendering-from-a-log}.md` where changes are needed; `T/README.md`; new `T/run-windows-acceptance.py` and `T/fixtures/browser.html`; the existing results report. No external eval-repository changes are required.
-
-- [ ] **1. Save one fixture and workflow driver.** `run-windows-acceptance.py --work DIR --shell powershell51|powershell7|gitbash --phase prepare|finish` prepares scenes or processes the recorded artifacts. `prepare` creates the same scratch browser app, captures its changing states, and creates an image plus an existing movie segment with its own known audio. `finish` requires `--take-one PATH --take-two PATH` pointing to the example's exported frame directories. Use `movie O'Brien λ & [take]` paths, CRLF scenes, UTF-8 BOM input coverage, a separate assembly work directory, narration-longer and visual-longer cases. Load the tiny browser counter from its file URI and change it through actual CDP mouse clicks. Reuse the example's retained `CDP.send`/`pump`/`responses` handling and Task 1's owned browser helpers, closing fixture processes afterward. No fixture server or browser recording framework is needed.
-
-The saved scenes name `title` (card), `still` (image), `browser` (frames), `take-one`/`take-two` (frames at 5 fps), and `source` (movie). Narrate the non-movie scenes; retain the source segment's own audio. Let the driver call existing tools through `fixtures.run_tool` and reject every nonzero result:
-
-```python
-steps = [
-    ("narrate", [str(scenes), str(voice), "--engine", "piper", "--verify", "on"]),
-    ("assemble", [str(scenes), str(cut), "--narration", str(voice), "--work", str(segments)]),
-    ("make-subtitles", [str(voice / "manifest.json"), str(subtitles),
-                        "--offsets-json", str(segments / "offsets.json")]),
-    ("burn-subtitles", [str(cut), str(subtitles), str(movie)]),
-    ("check-movie", [str(movie)]),
-]
-```
-
-These local variables are the driver's `Path` values under `--work`; use an evidence subdirectory for checker outputs. Local voice repeats synthesize into a new output directory using downloaded model caches; the cached-WAV off→on policy is tested separately. During final audio verification, compare each narrated interval with its script and the source-movie interval with its own reference, so preserved source audio is not misclassified as TTS drift.
-
-- [ ] **2. Run the two baseline instruction sessions before Markdown edits.** Apply `superpowers:writing-skills` when editing skill instructions. Use one fixed scenario for an actual PowerShell tool and an actual Git Bash tool, each against the unchanged instructions plus the final product code. This isolates the instruction change. Save the scenario at `W/instruction-scenario.md`:
-
-> On native Windows, make a narrated and hard-subtitled proof movie from the provided local app and terminal fixture, using the installed movie skill. Work in the supplied path with spaces, an apostrophe, and Unicode. Use local narration without a cloud key. Keep the recorded command alive between two takes. If desktop capture is unavailable, state what was actually captured and what remains unproven.
-
-Use a controlled unavailable-desktop condition in these sessions. Reuse `P/artifacts/windows-eval-preflight/native-tool-preflight.py` and `native-gitbash-preflight.py` as launcher references, removing diagnostic `--safe` mode and loading the explicit candidate plugin with skill bootstrap. The transcripts must show the actual shell tool and loaded skill; a PowerShell parent invoking a Bash tool does not count. Pin the same existing model and setup across before/after. Do not build a Quorum deployment or launch other agents during planning.
-
-- [ ] **3. Write and exercise Windows instructions.** Document all five `uv run --script` invocations, prerequisites, shell-specific quoting, pure PowerShell execution, foreground lifetime, consecutive request IDs, wait-only results, take timing, keys, fixed geometry, and cleanup. Use `[IO.File]::WriteAllText($path, $json, [Text.UTF8Encoding]::new($false))` for PowerShell UTF-8 without BOM. Document native producer status before `Tee-Object`, and Bash `PIPESTATUS`/pipefail for logs. Preserve Unix tmux instructions and fix the dangling example reference.
-
-Add the Windows desktop preflight:
-
-```text
-ffmpeg -nostdin -y -f gdigrab -framerate 5 -i desktop -t 2 capture-check.mp4
-ffmpeg -nostdin -y -i capture-check.mp4 -frames:v 1 capture-check.png
-```
-
-Invoke with an argument array and scratch paths from the ordinary-user interactive desktop; inspect actual application pixels. Document failure without converting it to a GUI success claim. Test Chrome and Edge card discovery once each; no browser-by-shell cross product is needed.
-
-- [ ] **4. Execute final acceptance and the two candidate instruction sessions.** Run the documented workflow from PS5.1, PS7, and Git Bash, recording the corresponding shell; one row each. Capture the invoking shell's own version/PID in its command transcript before launching uv; `--shell` selects the recorded shell and is not evidence for the invoking shell. Remove cloud keys only from the test process environment and use a process-local tool PATH excluding `llm` credential lookup; leave saved credentials untouched and select local Piper explicitly. Use native Windows tools and verify the actual ordinary-user token. For each final movie inspect picture, hard subtitles, audible narration, contact sheet, timing, and rendered-audio transcription. Record executable paths, shell identity, source revision, command exit codes, movie/contact-sheet paths, and inspection result. Complete the four-session before/after instruction comparison with the same scenario and candidate Markdown.
-
-On the native Mac and existing Linux runner, run only the portable suites `assembly`, `paths`, `checker`, `narration`, `browser`, and `subtitles` with `--require-capabilities`, including the changed shared code paths. Windows-only `processes`/`terminal` are not part of those runs. A skipped required case is incomplete, never green. Add actual fresh and cached-model voice checks on the existing Mac/Linux setup when validating the changed shared narration helper; do not re-provision other architectures.
-
-- [ ] **5. Close the reviewed scope.** Append a three-row Windows matrix, Mac/Linux regression results, four instruction-session results, and desktop/card checks to the existing report. Include stable artifact locations and the final tested source revision. Mark every spec acceptance row passed or explicitly incomplete; do not merely count commands returning zero. Review the final diff and commit `docs(movie): document and verify Windows workflows`. No pushing, PR, or merge is part of this plan.
-
-## Spec coverage and execution boundary
-
-| Reviewed requirement | Task |
-| --- | --- |
-| Five tools, paths, cards, UTF-8/BOM, subtitles, isolated ASR including cached verification (R2) | 1 |
-| Fixed capture decision gate; producer status (R1); ordered/asynchronous control (R3); timeout/shutdown (R4); honest frame timing (R5) | 2 |
-| Final recorder cleanup, same-session takes, wrapping, TUI and native shell behavior | 2 |
-| Three complete Windows workflows, local voice, final-movie inspection, desktop/card checks, existing-platform regressions | 3 |
-| Windows instructions and four before/after skill sessions | 3 |
-
-At execution, retain focused RED/GREEN evidence and review each of these three milestones. Rerun only checks affected by a fix; final integration must use final product code. If a genuine blocker requires changing an approved interface, dependency, supported route, or validation scope, present that concrete issue before expanding work.
-
-Execution resumed under the reviewed three-milestone scope. Implementation and finite test outcomes are recorded in the results report; the earlier superseded controller/workers were not resumed. Review the documented candidate-instruction gaps before claiming full skill compliance.

+ 0 - 286
docs/superpowers/reports/2026-09-09-proof-movie-os-compatibility.md

@@ -1,286 +0,0 @@
-# Native Windows feasibility gate — Task 1
-
-The bounded mechanism gate passed on Ballmer under an actual Medium-integrity ordinary-user token for PowerShell 5.1, PowerShell 7 and Git Bash: **40/40 checks**, including 27 command-outcome cases. That evidence belongs to commit `49bc2932857071896fb764688594709a1b494eaa`; the cleanup error-path correction below has separate verification. This is a feasibility probe, not the bulk OS port. The TUI result uses the controller-approved same-page snapshot checkpoint; general capture freshness and resize/redraw decoding remain required production work.
-
-Public implementation: [probe-windows.py](../../../tests/proving-it-works-with-a-movie/probe-windows.py). Focused ownership tests: [test-probe-cleanup.py](../../../tests/proving-it-works-with-a-movie/test-probe-cleanup.py). These files and this index remain in Git after the temporary SDD workspace is removed. Generated frames, screenshots and archives are outside core and currently retained at the locations below; durable evidence consolidation is pending.
-
-## Actual environment and prerequisites
-
-Launch host was `workerbee` (macOS); recording host was `ballmer`, Windows 11 Pro 10.0.26200, native x64. This observation does not impose a Windows 11-only production requirement. Remote task root is exactly `C:\Users\drew\movie-os-task1-20260909`. Existing Chrome and Git Bash were reused after inventory. Portable prerequisites, cache and test directories are task-owned; machine PATH, profiles, firewall, credentials, services, account membership and management networking were not changed.
-
-| Tool | Observed version | Executable relative to task root unless absolute |
-| --- | --- | --- |
-| Python | CPython 3.12.14, win32 AMD64, 64-bit pointers | `tools\python\cpython-3.12.14-windows-x86_64-none\python.exe` |
-| uv | 0.12.12, c4be69153 | `tools\uv\uv.exe` |
-| PowerShell 5.1 | 5.1.26100.9168 | `C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe` |
-| PowerShell 7 | 7.6.6 | `tools\pwsh\pwsh.exe` |
-| Git Bash | 5.3.9(1)-release, Git 2.54.0.windows.1 | `C:\Program Files\Git\bin\bash.exe` |
-| ttyd | 1.7.7-40e79c7, PE x64 | `tools\ttyd.exe` |
-| Chrome | 152.0.7977.65 | `C:\Program Files\Google\Chrome\Application\chrome.exe` |
-| FFmpeg/ffprobe | 9.0.1-essentials_build-www.gyan.dev | `tools\ffmpeg\ffmpeg-9.0.1-essentials_build\bin` |
-| CDP client | websocket-client 1.9.0 | uv script environment |
-
-The [prerequisite inventory](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/prerequisites.json) retains argv, stdout/stderr, exit codes and executable hashes. Native FFmpeg exercised libx264, AAC and subtitles/libass support. Git Bash invoked native Windows Python and FFmpeg explicitly.
-
-| Retained download origin | Actual SHA256 |
-| --- | --- |
-| [uv 0.12.12](https://github.com/astral-sh/uv/releases/download/0.12.12/uv-x86_64-pc-windows-msvc.zip) | `3d54912924c36e862c14f427d04f2ed70a99e8001d1c30caa101f6d5711626d5` |
-| [PowerShell 7.6.6](https://github.com/PowerShell/PowerShell/releases/download/v7.6.6/PowerShell-7.6.6-win-x64.zip) | `02fe458be20493fbdf43f61ea20610b811ee6c738ab1676c61b9cfcd1a33c860` |
-| [ttyd 1.7.7](https://github.com/tsl0922/ttyd/releases/download/1.7.7/ttyd.win32.exe) | `e33a27501b10b96981335bcba938b1145c7f52551a343e72160f00ab71832b37` |
-| [Gyan FFmpeg release essentials](https://www.gyan.dev/ffmpeg/builds/ffmpeg-release-essentials.zip) | `fec81ae03971d9dd4be3ebe02e263bd2ec1d789483f931bdba5f5715e65da2e9` |
-| [Astral CPython 3.12.14, 20260901](https://github.com/astral-sh/python-build-standalone/releases/download/20260901/cpython-3.12.14%2B20260901-x86_64-pc-windows-msvc-install_only_stripped.tar.gz) | `7c45c9622400d578709a9b2cddbe8124cc21d382409d9f13406d706d28e31b14` |
-| [websocket-client 1.9.0 wheel](https://files.pythonhosted.org/packages/34/db/b10e48aa8fff7407e67470363eac595018441cf32d5e1001567a7aeba5d2/websocket_client-1.9.0-py3-none-any.whl) | `af248a825037ef591efbf6ed20cc5faa03d3b47b9e5a2230a529eeee1c1fc3ef` |
-
-The download records and checksums are in [downloads.jsonl](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/downloads.jsonl); the portable files remain under the remote `downloads` and `tools` directories. The FFmpeg project [links the Gyan Windows distribution](https://ffmpeg.org/download.html).
-
-SSH itself ran as High-integrity `ballmer\drew`; those early results remain elevated. Its linked-token query failed with Win32 1312. Acceptance instead ran in existing desktop Paseo terminal `1496200b-5545-4d85-af31-d01df72a3c11`, `movie-standard-user-probe`, desktop session 1. Every acceptance supervisor independently recorded `ballmer\drew`, elevation type 3, elevated false, integrity SID `S-1-16-8192`, Administrators deny-only. Full token evidence is in each session report. No High-integrity result substitutes for ordinary-user acceptance.
-
-## Bounded observations
-
-The same Chrome page enables CDP Network observation before navigation, identifies its own ttyd socket, and supplies command input, output observation and recording. Native children are created suspended, assigned to an unnamed non-inheritable kill-on-close Windows Job, then resumed. There is no unowned-root fallback. The supervisor and unrelated sentinel remain outside that Job; Win32 ownership uses retained handles and creation times, never Git Bash's MSYS PID as a native PID.
-
-Two separate harness calls started and ended takes around a still-pending 45-second command, then verified the same shell PID/nonce and `persisted-λ`. Actual columns were **217 → 167** in all three Medium sessions; earlier elevated runs were 218 → 167. Completion payload lines are bounded to **60 base64 characters**. Truncated framing produced no record, damaged framing no successful record; browser loss produced interrupted/null success. This is not arbitrary resize/redraw decoding.
-
-All 27 command cases retained actual status attribution. A ParseException can coexist with raw PowerShell `$? = true`; `raw_shell_success` remains separate from request-attributed `parse_error=true` and semantic `shell_success=false`. The Bash logging pipeline retained `[7,0]` and producer exit 7. Extra ttyd clients were refused with EOF, and the original page still completed a nonce/PID round trip. Browser loss retained Win32 10054 and interrupted outcomes.
-
-| Cleanup, seconds | PowerShell 5.1 | PowerShell 7 | Git Bash |
-| --- | --- | --- | --- |
-| Normal | 1.890 | 1.859 | 1.875 |
-| Cancellation | 1.984 | 1.984 | 2.047 |
-| Browser loss | 0.063 | 0.047 | 0.047 |
-| Supervisor crash | 0.250 | 0.250 | 0.203 |
-
-These original successful paths checked parent/child/grandchild heartbeats, retained process-handle exit, stopped heartbeats and sentinel survival followed by scoped sentinel cleanup. The final original audit checked 609 PID/creation-time observations and found no remaining original processes, distinguishing reused PIDs. The review nevertheless found an error-path leak, addressed below.
-
-The accepted TUI JPEG is `interactive/000009.jpg` in each `tui-verified-*` directory. Implementer, reviewer and controller inspected active-take pixels showing `Native Windows TUI λ` and `[q] Finish this screen`, while run request 2 remained pending and before `q` request 4. The explicit same-page `snapshot` request 3 occurred after application readiness. [Pixel inspection metadata](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/active-tui-pixel-inspection.json) preserves hashes and timing.
-
-| TUI event, epoch seconds | PS5.1 | PS7 | Git Bash |
-| --- | --- | --- | --- |
-| Application ready | 1788994874.380406 | 1788994874.7009904 | 1788994874.6822324 |
-| Snapshot request 3 | 1788994875.4311693 | 1788994875.854709 | 1788994875.7615943 |
-| Inspected JPEG received | 1788994875.7732553 | 1788994875.8755016 | 1788994875.7821376 |
-| q request 4 | 1788994875.9168005 | 1788994876.04243 | 1788994875.9338365 |
-
-CDP trace proves the receiver continued draining/acknowledging during the ready-file wait/hold: 12/9/7 received events, 8/8/8 timeout polls and 4/3/2 acknowledged screencast frames. Some final-run TUI frames preceded the checkpoint, so this does not establish why earlier takes stalled or claim the snapshot alone caused a redraw. Acceptance is scoped to the inspected checkpoint method.
-
-## Retained failures and corrections
-
-- With the same ttyd binary, Job and CDP page, initial `startup-*` sessions disconnected without output. Adding explicit `ttyd -w <session-directory>` made `cwd-*` readiness pass. Upstream [ttyd PR 1502](https://github.com/tsl0922/ttyd/pull/1502) initializes command-line/cwd pointers to NULL; that supports the prerequisite without establishing the internal cause of our failure. No error 123 was observed in our ttyd log.
-- The probe's initial PowerShell prompt separately raised `InvokeMethodOnNull` because its error baseline was unset. Initializing and guarding the baseline corrected this; original and corrected evidence is retained separately from ttyd's cwd prerequisite.
-- Original Git Bash completion framing wrapped at 218 columns and ConPTY duplicated a character at a cursor repositioning boundary. Raw `outcomes-gitbash/network.jsonl` and `terminal.bin` retain the failure. Bounded multiline framing corrected the tested dimensions without adding a VT emulator.
-- Earlier active TUI takes lacked the TUI pixels despite a successful command and end PNG; a one-second hold still failed in PowerShell. `standard-normal-powershell51`, `standard-tui-*` and `tui-diagnostic-*` preserve those failures. The live PNG alone was never accepted instead of the active-take assertion.
-- Original handle-exit timing and transient heartbeat access failures are retained in elevated `session-v2-*` and `final-powershell7`. Bounded waits/retries corrected those observed cases; persistent access failures remain errors, motivating the review fix below.
-
-## Review fix: cleanup evidence failures
-
-Review of `49bc2932` found that snapshot, heartbeat or pending-result errors could escape before owned-resource cleanup. The correction protects evidence collection, always attempts terminal/Job release, gives the outside-job sentinel an independent finalizer, records stage-specific diagnostics, and returns failure when evidence is missing or reporting fails. An unknown process snapshot is `null`, not an empty successful ownership observation.
-
-The first actual Medium RED run occurred before the cleanup edit: snapshot, persistent heartbeat and pending-result failures each left three owned workers and the sentinel alive. Terminal-close failure left three workers alive, although its sentinel was already released; report-write failure escaped after resources exited. The normal base-source case released resources but returned `None`, failing only the new boolean result contract. Thus 0/6 is not a claim that all six scenarios leaked. Independent verifier teardown happened only after these observations were captured.
-
-A strengthened test replaced the synthetic post-release job error with native `TerminateJobObject` access denied (Win32 5), retaining real kill-on-close handle release. It again produced 0/6 against the unchanged base source, then **6/6 against the fix**, actual test-process exit 0. Each injected failure returned false, kept `normal_cleanup=failed`, preserved diagnostics and left no owned processes or sentinel alive. The successful fixture returned true. All cases independently recorded Medium SID `S-1-16-8192`, elevated false, elevation type 3, session 1.
-
-| Focused GREEN case | Cleanup seconds | Retained diagnostic stages |
-| --- | --- | --- |
-| Snapshot failure | 0.422 | `snapshot`; ownership snapshot remains null |
-| Persistent heartbeat access failure | 0.484 | `heartbeats_before`, `heartbeats_after`, `heartbeats_verify` |
-| Pending-result write failure | 0.468 | `pending_reply` |
-| Terminal failure plus native Job termination failure | 0.468 | `terminal_close`, `job_close` |
-| Final report write failure | 0.469 | `report_write`; absent report is not accepted |
-| Successful native ownership fixture | 0.453 | None |
-
-A fresh real ttyd/Chrome PowerShell 5.1 session under Medium closed in **2.891 seconds**, cleanup passed, supervisor launcher exit 0. A second real Medium session deliberately created a directory at its final `probe.json` path before close. The actual native replace failed with Win32 5; the supervisor exited **1**, stderr/stdout retained `report_write`, and the summary marked cleanup failed. An independent observer retained 19 native process handles before close (including supervisor and sentinel), verified all signaled, and checked stopped worker heartbeats. The remaining `probe.tmp` is an unsuccessful pre-error write, not a canonical passing report. No unrelated 27 outcome cases or 18 imported baseline assertions were rerun in this fix round.
-
-| Focused evidence | Result |
-| --- | --- |
-| [Initial pre-edit RED](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/fix-round1-red/cleanup-tests.json) | Expected 0/6, test process exit 1 |
-| [Native-fault RED against base](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/fix-round1-red-native/cleanup-tests.json) | Expected 0/6, exit 1 |
-| [GREEN ownership cases](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/fix-round1-green/cleanup-tests.json) | 6/6, exit 0; SHA256 `58c687f146fdb5063892dc0de1157d6f9915a4a11be3eb412dc5298b147e48a7` |
-| [Successful live close](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/fix-round1-live-powershell51/probe.json) | Cleanup passed, launcher exit 0 |
-| [Live report failure observer](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/fix-round1-live-report-failure-powershell51/failure-exit-observation.json) | Supervisor exit 1, failed cleanup, released resources |
-| [Fix-round archive](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/evidence-fix-round1.zip) | 265 files, 3,473,557 bytes; SHA256 `e22564f54e7bde79de3e2ea904ccab9d0faffa825c17f1e80c0beb78ff69ffab` |
-
-New probe SHA256: `b7f8a47bb09c9d4ed08642f8292df34a51b78a3040e0d6a3798a76286405ce90`. Focused test SHA256: `d85988949955a4f67fcc724a81c616b7f3e57c819f0a314632ab89355b4c89d1`. The initial RED test snapshot, hash `05c92c9ddfa9a4e05615634f9cbe80798f04856e994b3302ee7ad73decd6d20c`, is retained separately. The fix archive is also at `C:\Users\drew\movie-os-task1-20260909\evidence-fix-round1.zip`; exact new remote session directories are `evidence\fix-round1-{red,red-native,green}`, `evidence\fix-round1-live-powershell51` and `evidence\fix-round1-live-report-failure-powershell51`.
-
-An intervening SSH/SCP reset prevented the initial fix upload/submission before execution. A bounded read-only `hostname` retry succeeded; the unchanged remote base-source hash was verified before upload. These transport failures are separate from runtime test outcomes. No host/network repair was attempted. The task terminal remains available; probe resources have been released.
-
-## Reproduction and artifact index
-
-Run the following in a native **Medium-integrity** PowerShell terminal after copying the public probe and focused test beside the task-owned prerequisites. Choose new output directories; the probe refuses unsafe reuse. These are explicit argument arrays/paths, not PATH changes.
-
-```powershell
-$taskRoot = 'C:\Users\drew\movie-os-task1-20260909'
-$python = "$taskRoot\tools\python\cpython-3.12.14-windows-x86_64-none\python.exe"
-$uv = "$taskRoot\tools\uv\uv.exe"
-$probe = "$taskRoot\probe-windows.py"
-$env:UV_PYTHON_INSTALL_DIR = "$taskRoot\tools\python"
-$env:UV_CACHE_DIR = "$taskRoot\cache"
-$env:UV_NO_CONFIG = '1'
-$env:PYTHONIOENCODING = 'utf-8'
-& $python "$taskRoot\test-probe-cleanup.py" --directory "$taskRoot\evidence\cleanup-reproduction"
-
-& $uv run --python $python --script $probe --serve --ttyd "$taskRoot\tools\ttyd.exe" --browser 'C:\Program Files\Google\Chrome\Application\chrome.exe' --shell 'C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe' --shell-kind powershell51 --directory "$taskRoot\evidence\session-reproduction" --launch-host workerbee
-```
-
-Keep `--serve` alive. In separate harness calls with the same task environment, run the phases below against that directory. `first` returns while its command is still pending; invoke `second` before the 45 seconds elapse. For PS7 or Git Bash use their absolute executable above and matching `--shell-kind powershell7` or `gitbash` on every invocation. `second` also completes the interactive take, including the approved snapshot checkpoint. Separate fresh sessions exercise `cancel` and `browser-loss` instead of `close`.
-
-```powershell
-& $uv run --python $python --script $probe --phase first --shell-kind powershell51 --directory "$taskRoot\evidence\session-reproduction"
-& $uv run --python $python --script $probe --phase second --shell-kind powershell51 --directory "$taskRoot\evidence\session-reproduction"
-& $uv run --python $python --script $probe --phase prepare-cleanup --shell-kind powershell51 --directory "$taskRoot\evidence\session-reproduction"
-& $uv run --python $python --script $probe --phase close --shell-kind powershell51 --directory "$taskRoot\evidence\session-reproduction"
-```
-
-For standalone TUI reproduction, start a fresh `--serve` session using a new directory such as `$taskRoot\evidence\interactive-reproduction`. In a separate harness call, invoke `--phase interactive` with that directory and the matching shell kind, then run `prepare-cleanup` and `close` against the same fresh session. Do not invoke `interactive` after `second` in an existing session: both create the take named `interactive`, and the probe refuses to reuse its directory.
-
-Actual Medium launch commands were submitted through the already installed matching CLI, `C:\Users\drew\AppData\Local\Programs\Paseo\resources\bin\paseo.cmd terminal send-keys 1496200b-5545-4d85-af31-d01df72a3c11 '<encoded PowerShell command>' Enter --json`. SSH transport remained High; the daemon-owned terminal supplied the Medium token. The exact UTF16LE commands, launch scripts, subprocess argv and outputs are preserved as `*-submit-command.json`, `*-launch.ps1`, `*-launcher.json` and named stdout/stderr files, not inferred from this example.
-
-Local temporary evidence root is `.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/`; `remote/` mirrors remote `C:\Users\drew\movie-os-task1-20260909\evidence`. Links below currently target that ignored workspace. They will cease to resolve after SDD deletion unless evidence is consolidated first; source and this report remain reviewable independently.
-
-| Evidence | Contents |
-| --- | --- |
-| [Aggregate](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/probe.json) | Original 40 passed checks; per-session unrelated checks remain unavailable |
-| [Main archive](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/evidence-final.zip) | 2,323 files, 160,448,403 bytes; outcomes, continuity, ownership, prerequisites and original failures |
-| [Capture supplement](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/evidence-capture-final.zip) | 650 files, 23,823,967 bytes, overlapping earlier artifacts; accepted TUI pixels/traces |
-| [PS5.1 active TUI JPEG](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/tui-verified-powershell51/interactive/000009.jpg) | Inspected while request 2 pending, before q |
-| [PS7 active TUI JPEG](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/tui-verified-powershell7/interactive/000009.jpg) | Same acceptance method |
-| [Git Bash active TUI JPEG](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/tui-verified-gitbash/interactive/000009.jpg) | Same acceptance method |
-| [Original process audit](../../../.superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/final-process-audit.json) | 609 PID/creation-time observations, no remaining original processes |
-
-Main archive SHA256: `ba7438c8686183ceb1f55dd32e62bed3fc57d082bc642f5487561aca364da024`. Capture supplement SHA256: `4b1186772fec51d3581390ae4dc462811d5b6b7c86018fc36b0062de384a844a`. Remote copies are `$taskRoot\evidence-final.zip` and `$taskRoot\evidence-capture-final.zip`. Original runtime source snapshots are retained locally as `probe-standard-runtime-a107564b5e34.py` (SHA256 `a107564b5e341c84d5ae21bc9ecbb7c41cd1544752597c78f0cde5dae37489c2`) and `probe-final-runtime-be36b8182b5c.py` (SHA256 `be36b8182b5caa1dc071dad3bd075cc44d40a3b6fe7b3282b6da739b44259d40`). Commit `49bc2932` probe SHA256 is `515f382ecb960c6111649784e8ecb23d66f2ac67f7f2cdbb498760a2679bf67b`.
-
-For command outcomes, use the full launch command above with `--outcomes` instead of `--serve`, a fresh directory, and each selected shell. For a supervisor-crash check, start a fresh `--serve`, run `prepare-cleanup`, then use `& $uv run --python $python --script $probe --crash-check --directory "$taskRoot\evidence\session-reproduction"` against that crash session. The observer validates retained native identities before killing the supervisor and releases the sentinel afterward. Inspect active TUI JPEGs before claiming their pixels; a successful phase alone does not establish the visual assertion.
-
-The probe's `--help` exposes these modes. Original complete aggregate replay commands are below; they validate retained evidence, not a new Windows runtime execution:
-
-```sh
-python3 tests/proving-it-works-with-a-movie/probe-windows.py --aggregate .superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote
-python3 tests/proving-it-works-with-a-movie/probe-windows.py --assert-result .superpowers/sdd/2026-09-09-proof-movie-os-compatibility/task-1-evidence/remote/probe.json
-```
-
-## Required production follow-through
-
-General screencast freshness is a **required production fix and test in Tasks 8–9**; stalled takes must not pass. Resizing requires a production shell adapter/supervisor that handles completion framing through resize/redraw, with damaged/missing framing remaining unknown/interrupted. Further work includes broader ownership failure injection, the full OS support matrix, and assembly/narration/verification portability. This probe does not establish those results. The Task 1 mechanism gate and independent review must finish before the bulk port begins.
-
-
-## Reviewed Windows completion — final results
-
-The earlier sections record historical feasibility work, not the current
-acceptance gate. The reviewed three-milestone implementation is now present.
-The final native workflows passed; the four-session instruction comparison
-has two explicitly incomplete verification outcomes below. No PR, push,
-merge, further OS provisioning, or additional eval platform was performed.
-
-Ballmer: native Windows build 26200, x64, ordinary Medium token
-`S-1-16-8192`, prepared CPython 3.12.14. Actual invoking-shell transcripts
-identify PS5.1 `5.1.26100.9168`, PS7 `7.6.6`, and Git Bash
-`5.3.9(1)-release`, with executable/PID before uv. This is evidence for that
-tested host, not every Windows release or architecture.
-
-| Invoking → recorded shell | Five tools + checker + rendered audio | Duration | Retained artifacts |
-|---|---|---:|---|
-| PS5.1 → PS5.1 | PASS | 27.000 s | [Movie](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/powershell51/movie O'Brien λ & [take]/movie.mp4>), [contact sheet](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/powershell51/movie O'Brien λ & [take]/evidence/checker/contact-sheet.png>), [rendered audio](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/powershell51/movie O'Brien λ & [take]/evidence/rendered-audio.json>) |
-| PS7 → PS7 | PASS | 26.878 s | [Movie](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/powershell7/movie O'Brien λ & [take]/movie.mp4>), [contact sheet](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/powershell7/movie O'Brien λ & [take]/evidence/checker/contact-sheet.png>), [rendered audio](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/powershell7/movie O'Brien λ & [take]/evidence/rendered-audio.json>) |
-| Git Bash → Git Bash | PASS | 27.167 s | [Movie](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/gitbash/movie O'Brien λ & [take]/movie.mp4>), [contact sheet](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/gitbash/movie O'Brien λ & [take]/evidence/checker/contact-sheet.png>), [rendered audio](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/gitbash/movie O'Brien λ & [take]/evidence/rendered-audio.json>) |
-
-All movies are 1600×900 H.264/AAC. Each combines a card, still, real CDP
-counter clicks, two native terminal takes, and a labeled existing movie
-segment retaining its own 440 Hz stereo tone. Inputs exercise the path
-`movie O'Brien λ & [take]`, UTF-8 BOM/CRLF scenes, and separate assembly work.
-Title narration lasts 3.855–3.971 s against one second of visuals; still
-visuals last four seconds against 1.498–1.695 s of narration.
-
-Each workflow used real local Piper synthesis with verification, then a new
-output directory reusing downloaded model caches. The process PATH excludes
-`llm`, and cloud keys were removed only from that process. All 15 narrated
-final-movie intervals passed local ASR comparison (0% length drift, worst
-changed-word run at most two). Each source-tone interval separately passed
-frequency/amplitude comparison with its own source reference. Fresh/cached
-WAV verification is not substituted for final-movie transcription.
-
-The exported terminal samples show red→green→blue in order and visible
-command completion after the second take's exit key. Six take duration
-errors are 0.028–0.094 s; maximum completed-screenshot gaps are
-0.625–1.109 s. Frame hashes confirm the documented completed-sample grid,
-with no future sample used. Contact sheets and decoded movie frames were
-inspected; short color states missed by the contact-sheet selection are
-present in the actual movies.
-
-Evidence index: [timing/frame inspection](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/timing-and-frame-inspection.json>),
-[controller final-movie inspection](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/controller-final-movies.json>),
-[final source SHA-256 manifest](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/final-source-sha256.json>).
-The manifest pins the tested recorder, helpers, five tools, and fixture;
-Task 3 applies to base `ce9b3fcf`, following media `5621d1d6` and recorder
-`ce9b3fcf`. The retained source hashes, commands, and source-labeled earlier
-results establish provenance without claiming an older artifact tested a
-later failure-path change. Remote artifacts remain at
-`C:\Users\drew\movie-windows-completion\final-task3`.
-
-### Capture targets and reused regression evidence
-
-Named-window `gdigrab` captured readable pixels from the task-owned Windows
-application: [window image](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/desktop-target/title.png>) and
-[commands/results](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/desktop-target/result.json>). Visible padding and
-capture dimensions are retained; no DPI/root-cause claim is inferred.
-Full-desktop capture returned zero but showed wallpaper only, including a
-late frame: [desktop image](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/desktop-target/capture-check.png>).
-That target remains **unverified for application capture in this session**.
-Read-only evidence showed WinSta0/Default and an active console session;
-no lock, service, credential, or machine-setting changes were made.
-
-Chrome and Edge title-card checks and the unchanged media regressions reuse
-`.superpowers/evidence/windows-completion/media-5621d1d6`. The Mac and Linux
-portable assertions passed 35 per host; native-only Job tests are excluded
-from that portable result. Final strict browser/narration checks passed
-4/4 and 9/9 per host. Real fresh and cached-model voice checks on both hosts
-reuse `.superpowers/evidence/windows-completion/{macos,linux}-voice`.
-The narration/browser/subtitle code exercised there is unchanged.
-
-The existing terminal capture gate, native outcome/control/ownership
-regressions (57 tests), and finalization corrections (7 focused tests) retain
-their source labels under `capture-gate`, `terminal-b8a723f4`, and
-`terminal-ce9b3fcf`. Task 3 adds the demonstrated Unicode-cwd startup fix and
-the previously reviewed identity-error handle cleanup. Focused native tests
-passed 4/4. Local changed-policy suites passed 12 terminal and two ownership
-assertions; their 13 and two native-only skips do not establish native
-coverage. The native final workflows separately exercise all three shells.
-
-### Four instruction sessions — comparison complete, full compliance partial
-
-All four loaded the explicit candidate plugin and successful
-`using-superpowers` SessionStart bootstrap, invoked the movie Skill, and used
-actual PowerShell or Bash tools. The scenario, model (`claude-sonnet-5`,
-existing `sonnet` alias), low effort, supplied tools/caches, product code, and
-stop-at-first-instruction-gap criterion were held constant. Only the candidate
-Markdown changed. No diagnostic safe mode or extra implementation agent was
-used. Permissions/context were scoped to the task plugin, fixtures,
-artifacts, and owned processes.
-
-| Session | Observed result |
-|---|---|
-| PowerShell baseline | Demonstrated documentation failure: inferred the native route from source, failed special-path `Start-Process` quoting, BOM request input, and guessed an `operation: result` request. Stopped through owned cleanup after the gap; not a completed workflow. |
-| Git Bash baseline | Demonstrated documentation failure: inferred the native route from source and passed `/c/...` to native Python, producing `C:\c\...`/file-not-found. Stopped through the launcher's Job cleanup after the gap; not a completed workflow. |
-| PowerShell candidate | **Partial.** New foreground-task, UTF-8 request, consecutive-ID, two-take, wait-only result, local voice, five-tool, caption/contact-sheet, cleanup, and honest capture-boundary instructions worked. It claimed complete verification but omitted transcription of the rendered final audio. |
-| Git Bash candidate | **Incomplete due to harness turn cap.** Built a movie and ran media/checker/evidence steps, then returned `error_max_turns` (41 reported turns, exit 1) before a final response or rendered-final-audio transcription. This is not a full skill pass or an inferred instruction failure. |
-
-Launch limits were $4, 40 turns, and 900 seconds per session. The PowerShell
-baseline's interrupted result reported 54 turns/$1.4456, and the completed
-PowerShell candidate reported 44 turns; the requested cap must not be
-mistaken for the harness's actual reported count. Git Bash candidate cost
-was $0.832 and elapsed time 223.6 s. No limit was silently scored as success.
-Transcripts and launcher result records are retained under
-[the stable instruction evidence](</Users/drewritter/.paseo/worktrees/2mmrq9t5/movie-os-compatibility/.superpowers/evidence/windows-completion/final-task3/instructions/>). Further corrective
-instruction work is left to review; no extra sessions were launched.
-
-### Acceptance disposition
-
-| Reviewed acceptance area | Disposition |
-|---|---|
-| Three native invoking/recorded-shell workflows | PASS on the recorded host |
-| Repeatable mixed fixture, paths, encodings, timing | PASS |
-| Terminal control, continuity, native outcomes, ordered automatic states | PASS with retained focused regression evidence |
-| Fresh/cached-model local voice per workflow | PASS |
-| Finished picture, hard captions, checker and rendered audio | PASS for all three independent acceptance movies |
-| Browser cards / capture targets | PASS Chrome, Edge, and named-window pixels; full-desktop target explicitly unverified |
-| Focused negative cases and ownership | PASS with source-labeled earlier evidence plus Task 3 focused corrections |
-| Existing Mac/Linux behavior | PASS reused unchanged portable/media/voice evidence |
-| Four instruction comparisons | Executed; **full candidate skill compliance remains INCOMPLETE** for the two specific outcomes above |

+ 0 - 548
docs/superpowers/specs/2026-09-09-proof-movie-os-compatibility-design.md

@@ -1,548 +0,0 @@
-# Proof movie OS compatibility
-
-**Date:** 2026-09-09
-
-**Status:** Historical design. Remaining work is superseded by the [focused Windows completion spec](2026-09-09-proof-movie-windows-completion-design.md). Do not resume this design's broader rollout.
-
-**Scope:** Platform compatibility for the movie skill proposed in [PR #2214](https://github.com/obra/superpowers/pull/2214). Implementation depends on that import or its final equivalent.
-
-## Goal and agreed direction
-
-The movie skill must cover the same OS families and execution environments as
-Superpowers' visual brainstorming companion: macOS, Linux, native Windows, WSL,
-and headless or remote operation. Windows users must be able to launch the movie
-pipeline from PowerShell or Git Bash and record either shell. Native Windows
-must not require WSL; the PowerShell movie workflow must not require Git Bash.
-
-An agent should be able to record real software, assemble a narrated and
-subtitled movie, inspect it, and apply the same verification criteria on every
-supported environment. Native terminal recording is part of the initial scope.
-
-Compatibility means equivalent outcomes, documented prerequisites, and tested
-behavior. Capture methods can differ by OS and available display facilities.
-
-## Evidence and baseline
-
-This design uses local Superpowers commit `3a8bdc11` and PR #2214 head
-`f6617db1517488d4a8b66925519d5ffbeef965b3` as its inspected baselines.
-
-The companion's current implementation establishes these distinctions:
-
-- [Browser launch](../../../skills/brainstorming/scripts/server.cjs) handles
-  macOS, native Windows, WSL, Linux with a display, and Linux without a display.
-- [Launch instructions](../../../skills/brainstorming/visual-companion.md)
-  accommodate PowerShell tools by invoking Git Bash's `bash.exe`. The companion
-  itself has a Bash launcher; this is not evidence of a separate PowerShell
-  implementation or of operation without Git Bash.
-- [Process startup](../../../skills/brainstorming/scripts/start-server.sh)
-  accounts for foreground execution and Windows/MSYS PID incompatibility.
-- [Windows lifecycle tests](../../../tests/brainstorm-server/windows-lifecycle.test.sh)
-  exercise Git Bash behavior. [Browser launcher tests](../../../tests/brainstorm-server/browser-launcher.test.js)
-  cover Windows, WSL, and headless Linux decisions, including treating URLs as
-  arguments rather than shell commands.
-
-These sources identify the required OS families. They do not establish an
-exhaustive promise covering every OS release, CPU architecture, Linux
-distribution, or desktop capture backend.
-
-An independent review of the platform contract and terminal design is recorded
-in the [adversarial review](2026-09-09-proof-movie-os-compatibility-review.md).
-
-For the imported movie skill, all three shell regression suites were run on
-macOS during this investigation: assembly 4/4, checker 9/9, narration drift 5/5.
-These 18 assertions do not test actual speech generation, title-card rendering,
-subtitle burning, or Windows execution. No Windows support is established by
-those results.
-
-Concrete compatibility gaps in the inspected import include Unix-only command
-examples, a tmux-dependent terminal recipe, macOS-only desktop capture examples,
-FFmpeg glob-dependent frame input, missing Windows browser discovery, and a
-hand-built file URL that misencodes a Windows path. `burn-subtitles` also uses
-only the subtitle basename without setting the working directory described in
-its comment.
-
-## Support contract
-
-| Execution environment | Shells used to launch tools | Required movie routes | Desktop capture behavior |
-| --- | --- | --- | --- |
-| macOS | zsh and Bash | Browser motion, terminal, stills, real-run logs, full processing pipeline | Use available macOS capture with permission; inspect a sample before recording |
-| Linux desktop | Bash; commands use portable POSIX syntax where applicable | Same routes and pipeline | Use an available capture backend for the active display; distinguish X11 from Wayland |
-| Linux headless, container, or remote session | Bash | Headless browser and terminal capture, stills, logs, full processing pipeline | Detect absent display/capture access and choose a route that can prove the claim |
-| Native Windows | Windows PowerShell 5.1, PowerShell 7, Git Bash | Same routes and pipeline; record native PowerShell and Git Bash sessions | Use a verified native capture backend, initially FFmpeg `gdigrab`, when a desktop is accessible |
-| WSL | Linux shell using Linux tools | Same Linux routes and pipeline | WSL capture is scoped to its available display; Windows host desktop capture is a separate native Windows operation |
-
-PowerShell 5.1 and 7 are separate invocation tests. Git Bash is a Windows shell
-environment, not a synonym for WSL. Git Bash must invoke native Windows Python
-and FFmpeg for its native Windows tests. WSL tests must invoke Linux tools.
-
-Keep OS detection, launch-shell detection, recorded-shell selection, and display
-capability separate. A PowerShell launcher may record Git Bash and vice versa.
-Do not infer the recorded shell from the shell running the pipeline.
-
-Do not introduce a Windows 11-only restriction solely because a test machine
-runs Windows 11. Record exact OS, architecture, interpreter, browser, and tool
-versions in validation results. Native terminal capture requires a compatible
-Windows console API; ttyd documents ConPTY builds for Windows 10 1809 or later.
-Other route prerequisites may set their own minimum versions.
-
-Architecture compatibility is also a dependency check: do not assume that a
-Python script or a Windows x64 wheel proves native ARM64 support. Required
-validation starts with macOS arm64 and x64, Linux x64, and native Windows x64,
-plus WSL on Windows. Record Linux arm64 and Windows arm64 dependency availability
-and execution evidence before listing those combinations as verified. A missing
-wheel is a reported route limitation, not an excuse to claim OS parity or
-silently substitute a paid service.
-
-## Approach and alternatives
-
-Use the existing five Python tools as one shared pipeline, with small local
-helpers where platform handling is reused. Keep shell differences in invocation
-examples and adapters for recording real commands. Keep capture differences in
-route documentation and its executable examples.
-
-This preserves the scene file, narration manifest, offsets, subtitles, movie,
-and checker output across operating systems. It also limits changes to the
-skill being imported; no companion rewrite is required.
-
-Two alternatives were considered:
-
-- **Route all Windows work through WSL or Git Bash.** This does not satisfy the
-  agreed native PowerShell workflow, and WSL cannot stand in for recording a
-  native Windows terminal or application.
-- **Maintain separate PowerShell and Bash pipelines.** This duplicates timing,
-  subtitles, and verification behavior and makes platform drift likely. Both
-  shells can invoke the same Python scripts explicitly.
-
-## Shared pipeline
-
-### Invocation and dependencies
-
-Document `uv run --script` for all five extensionless tools. Retain existing
-Unix shebang execution for compatibility. Examples set the skill directory using
-the location supplied by the harness, rather than a Claude-specific install path.
-
-```powershell
-$SkillDir = 'C:\path to plugin\skills\proving-it-works-with-a-movie'
-uv run --script "$SkillDir/scripts/assemble" scenes.yaml silent-cut.mp4
-if ($LASTEXITCODE -ne 0) { throw 'Movie assembly failed' }
-```
-
-```bash
-SKILL_DIR='/path to plugin/skills/proving-it-works-with-a-movie'
-uv run --script "$SKILL_DIR/scripts/assemble" scenes.yaml silent-cut.mp4 || exit "$?"
-```
-
-PowerShell recipes must work under 5.1 without relying on PowerShell 7-only
-operators. Bash syntax must be identified as Bash when it is not POSIX syntax.
-All examples must preserve failure status across pipelines and logging.
-
-Before a long recording, check prerequisites for the selected route: `uv`, a
-compatible Python, FFmpeg/ffprobe and required codecs/filters, the chosen
-browser, ttyd for terminal capture, and local narration/transcription packages
-when requested. Check capabilities, not merely executable presence. Report
-the missing capability and an appropriate install or alternate-route action.
-There is no need for a new global installer or plugin-startup dependency check.
-
-Use the tools already involved in the proposed import. Do not introduce a new
-recording service, TTS service, global Node dependency, or Windows-only package
-manager requirement. The import's dependency-policy exception remains a
-maintainer decision associated with PR #2214; this spec does not grant one.
-
-Nested helper invocations must also be independent of the project being filmed.
-In particular, launch the local ASR environment with uv project discovery disabled
-(`--no-project`) and a compatible interpreter selected by uv. The current nested
-`uv run --with faster-whisper ...` can otherwise inherit an unrelated
-`pyproject.toml`. A fixture with an incompatible project Python requirement must
-not change the helper's resolution, modify the project's environment/lockfile,
-or turn an available ASR into a reported dependency failure.
-
-### Paths, encoding, and process arguments
-
-- Resolve assets relative to the scene file and generated assets relative to
-  the selected work directory. Preserve existing explicit output-path semantics.
-- Use `Path.resolve().as_uri()` for browser file URLs. Browser discovery honors
-  `--browser`, then checks platform installations and PATH, including Chrome and
-  Edge on Windows and Chrome/Chromium on macOS and Linux.
-- Use explicit UTF-8 for YAML, JSON, HTML, SRT, concat lists, and text evidence.
-  Accept ordinary LF and CRLF text inputs and diagnose unreadable encodings.
-  Give owned Python helper pipes an explicit UTF-8 contract as well. Preserve
-  raw output from recorded external commands and record the encoding used to
-  decode it; do not assume every native Windows program emits UTF-8 or silently
-  replace undecodable evidence. Test redirected output under a non-UTF-8 Windows
-  console/code-page configuration.
-- Pass external executable arguments as arrays. Avoid shell command strings for
-  FFmpeg, browsers, or opening an artifact. User-requested shell commands belong
-  only in the selected recorded shell.
-- Keep MSYS path conversion at the launch boundary. Do not translate the same
-  path twice or feed Bash pseudo-PIDs to native Windows process APIs.
-- Support spaces, apostrophes, Unicode, and shell metacharacters in valid local
-  paths. Exercise nested subtitle directories and work directories outside the
-  current directory. Avoid requiring symlink privileges or POSIX tools such as
-  `mktemp`, `shasum`, and `kill` for the PowerShell pipeline.
-
-### Frame assembly and subtitles
-
-Expand and sort frame filenames in Python. Feed FFmpeg a deterministic numbered
-sequence staged in the work directory, using ordinary files without requiring
-symlinks. Preserve the existing lexical ordering and frame rate. This replaces
-the dependency on FFmpeg's optional glob support. Clean only staging files owned
-by the current build and preserve source screenshots.
-
-Write concat entries with FFmpeg-specific escaping and test native Windows
-drive paths. File list syntax and filter syntax are separate from shell quoting.
-Preserve measured offsets and the existing timing rules: `card`, `image`, and
-`frames` segments use `max(narration duration, visual duration)`; `movie` segments
-retain the source movie's duration and own audio. The port must not change timing
-or drop frames.
-
-For subtitle burning, stage the SRT under a safe fixed basename in a temporary
-directory and execute FFmpeg from that directory, with resolved input/output
-paths. This avoids interpolating user filenames into the subtitle filter.
-Clean staging on both success and failure. Verify the selected font actually
-renders the sample text, using available platform fonts rather than assuming
-DejaVu is installed everywhere.
-
-Preserve the documented soft-subtitle fallback, with an accurate explanation of
-why burning was unavailable. A burn failure with libass present must report the
-actual failure rather than claim libass is missing. Full-platform acceptance
-requires proving the hard-subtitle path with a capable FFmpeg build; a successful
-soft fallback alone is insufficient. Continue shipping the SRT alongside the
-movie for the existing checker contract.
-
-### Narration and verification
-
-Keep local Piper narration and local transcription available without a cloud
-key on the verified support combinations. First-run model downloads must be
-distinguished from subsequent offline operation. Check actual synthesis and
-transcription, including subprocess interpreter selection and Unicode handling.
-
-Do not make cloud credentials a Windows requirement. A missing local ASR must
-remain visible in output and cannot count as a verified narration acceptance
-run, even where the existing script permits a skip.
-
-Preserve the existing verification thresholds, contact-sheet inspection,
-subtitle checks, and requirement to compare rendered speech with its script.
-Platform compatibility is not a reason to weaken a gate. Existing general
-checker limitations remain limitations; redesigning its heuristics is outside
-this work.
-
-## Recording routes
-
-### Browser and stills
-
-Keep browser-driven recording as the common motion route. Select a compatible
-browser explicitly or through discovery, use an isolated profile, preflight a
-real screenshot, and retain cursor overlays and measured narration pacing.
-Headless recording must work without opening a desktop browser. Distinguish
-Playwright's own recording encoder, when that route is used, from system FFmpeg.
-
-Use available harness browser automation or the documented recorder's browser
-connection. The compatibility work does not create a new browser automation
-framework. Navigation races must have bounded capture timeouts. Unexpected blank
-output or stalled capture must not silently produce a successful recording;
-intentional blank transition beats remain valid evidence.
-
-### Terminal and TUI
-
-Keep the existing ttyd/tmux route for Unix environments. Add a native Windows
-route using ttyd's Windows console support and an explicitly selected PowerShell
-or Git Bash executable. Native Windows must not require tmux, Docker, or WSL.
-
-Drive the native terminal through browser keyboard input into the live ttyd
-terminal. Provide an executable example inside the skill with explicit shell
-selection, viewport, output location, and a short real command sequence. The
-existing `examples/film-terminal.py` reference points outside the imported
-files; replace it with an included, tested example or a correct included-file
-reference. The standalone example is specific to a Docker/tmux demo and must
-not be presented as a working native Windows recorder unchanged.
-
-One persistent Python supervisor owns the browser/CDP page, capture state, and
-terminal session. Start ttyd in writable mode and allow one terminal client.
-Subscribe through CDP to the existing page's ttyd WebSocket events before that
-page connects. Decode ttyd's output-message framing and incrementally parse
-output across message boundaries. A second CDP connection may observe that page;
-a second ttyd WebSocket must not be opened, because ttyd creates another shell
-for another terminal client. The browser input, command observer, and recorded
-pixels must all refer to the same session.
-
-For the reference recorder, later harness calls submit ordered JSON control
-files to a session-owned directory. A request has a unique monotonically
-increasing id and an operation (`run`, `key`, `begin-take`, `end-take`, or `close`).
-Write requests atomically; the supervisor dispatches them in order and writes
-matching acknowledgments and eventual results. A `run` acknowledgment does not
-wait for the command to finish: the command remains in flight while capture
-operations and `close` can proceed. Another `run` is accepted only at a known
-shell prompt; `key` can drive an expected interactive state. Startup reports the
-control directory and session id, and request ids prevent replay. No HTTP
-control service or additional terminal client is needed.
-
-Readiness requires a round trip through the filmed terminal: execute a probe
-that returns a session nonce, shell identity, and working directory, observe its
-completion, and inspect a screenshot from that same page. Do not declare the
-recorder ready based only on a listening port or a loaded page.
-
-Command completion is a shell adapter responsibility. Each scripted beat emits
-a framed completion record after it returns, containing session/request ids and
-an outcome. Its complete framing must not appear in echoed command text. Keep
-instrumentation inside the recorded session; do not edit the human partner's
-shell profiles or change its error preferences. A missing completion record is
-an interruption or timeout, never success.
-
-The outcome distinguishes `completed`, `interrupted`, and `unknown`; a completed
-beat has `shell_success`, a nullable native exit code, and a nullable shell error.
-For Bash, preserve the command status and any producer pipeline statuses before
-instrumentation overwrites them. For PowerShell, capture `$?` immediately after
-the submitted statement and distinguish it from `$LASTEXITCODE`. Attribute a
-native exit code only when the beat identifies its native executable producer;
-a cmdlet's outcome must not inherit an earlier native command's exit code.
-Terminating exceptions and parse failures produce explicit error outcomes when
-observable, or `unknown` when the session no longer permits reliable attribution.
-
-The first implementation accepts atomic command/statement beats. Multi-statement
-scripts remain opaque commands whose own documented status is recorded; the
-recorder does not claim to detect every internally ignored error. Logging and
-formatting are observer operations after status capture. If the command being
-proved is itself a producer/logging pipeline, its adapter must retain the
-producer's outcome separately, or require separate beats rather than label a
-successful logger as a successful producer. Evidence includes the exact command
-and attribution, so a successful recording can honestly show a failed command.
-
-Record PowerShell 5.1 and 7, not only launch the pipeline from each. Test a
-successful native executable followed by a failing cmdlet, a failing native
-executable followed by a successful cmdlet, terminating and non-terminating
-errors, an explicit script exit, parse failure, and failure through logging.
-Tests must detect status changes introduced by instrumentation, including the
-PowerShell 5.1 expression-wrapper behavior.
-
-For interactive TUI beats, use explicit application state and key actions; do
-not send another shell command while the TUI still owns input. Unknown state or
-a completion timeout stops the take with a diagnostic. Validate delayed commands,
-nonzero exits, an interactive TUI, resizing, and non-ASCII output in both recorded
-Windows shells.
-
-A recording session owns one terminal for the duration of its takes. Stopping a
-take pauses capture without killing a long-running command; the same browser
-connection and shell persist until the recording session ends. Browser
-disconnect/reconnect persistence is not assumed. If the connection is lost and
-the shell cannot be proven to survive, mark the take interrupted and report it.
-Verify a nonce, shell identity, and persistent shell variable across two takes,
-a running command between those takes, and a real harness tool-call boundary.
-Include negative tests where an extra terminal client or lost connection must
-not produce a successful session-continuity result.
-
-### Desktop capture and real-run logs
-
-Document device/capability selection per OS: macOS capture, native Windows
-`gdigrab`, and supported Linux display capture. Detect the active display and
-available FFmpeg devices before selecting a command. On Wayland, do not apply an
-X11 recipe by default; use an available permission-aware capture route or switch
-to browser/terminal recording. This work does not promise universal desktop
-capture on every compositor or remote-session configuration.
-
-Preflight by recording a short sample and examining its actual pixels. Permission
-denial, blank capture, an unavailable remote desktop, or missing display must
-produce an explicit route decision. A log reel can prove a real command run; it
-cannot substitute for pixels when the claim is specifically native GUI behavior.
-If the required claim cannot be captured, report it as unproven.
-
-Provide Bash and PowerShell real-run log recipes that retain the command's exit
-status, timestamps, and actual output. Use standard-library Python hashing or
-equivalent native commands for evidence bundles. Retain source logs beside the
-derived movie. Windows PowerShell 5.1 output encoding must be tested explicitly.
-
-## Lifecycle and remote behavior
-
-Match the companion's lifecycle principles without copying its Bash launcher:
-run recorders under the harness's persistent/background execution mechanism,
-report readiness after the real browser/terminal is reachable, and preserve the
-session between agent tool calls. Avoid assuming detached `nohup` processes
-survive in a Windows or managed-harness environment.
-
-Own browser profiles, ports, temporary assets, and child process handles per
-recording session. Bind terminal and debugging endpoints to loopback by default.
-Readiness and command waits have explicit timeouts. On finish, interrupt, or
-failure, stop and wait for owned processes and release resources; do not kill by
-generic executable name or unverified stale PID. Ordinary users must be able to
-run the workflow without elevated privileges.
-
-Session ownership includes the recorded command's local children and
-grandchildren, even when they stop using the terminal. On Windows, assign ttyd
-and the isolated browser root to a session-owned Job Object before their threads
-run, using suspended creation and Win32 calls through Python's `ctypes`. Disable
-breakaway and use kill-on-job-close so descendants cannot outlive the recorder.
-Keep the supervisor outside its child job, and detect failure to establish the
-job before starting a take. On Unix, track the recorder's own sessions/process
-groups, including ttyd's PTY process group; do not assume killing the ttyd parent
-terminates its shell group. Session persistence ends at `close`; this workflow
-does not launch local services intended to survive it.
-
-On normal close, keep reading terminal output while requesting graceful
-shutdown, then terminate remaining owned processes after a bounded timeout and
-wait for their exit. Drain/close output in an order that does not block older
-ConPTY implementations. On supervisor failure, Windows job closure must still
-terminate owned descendants. Test both success and cancellation with a child and
-grandchild that continue writing heartbeat files, plus an unrelated sentinel
-process. Owned heartbeats must cease, owned processes must disappear, the
-sentinel must survive, and no shutdown operation may hang indefinitely.
-
-Use the same distinction as the companion between a renderer and opening a
-browser for the human partner. If a recorder needs to open a URL, use the OS
-launcher with separate arguments; on WSL, Windows URL opening is a convenience,
-not permission to pass Linux filesystem paths to Windows executables. With no
-display, report the artifact path/URL and continue headless processing. Remote
-browser access uses an explicit, documented access mechanism; remote deployment
-or a new public tunnel service is outside scope.
-
-## Implementation boundaries
-
-Expected files are the five scripts under
-`skills/proving-it-works-with-a-movie/scripts/`, small shared helpers if needed,
-the skill's route documents, an included terminal recording example, and tests
-under `tests/proving-it-works-with-a-movie/`. Add a focused platform support
-reference and link it from `SKILL.md` rather than making every route repeat the
-whole compatibility matrix.
-
-Preserve scene kinds, manifest fields, offsets, CLI arguments, and output artifact
-formats. Keep macOS and Linux entry points working. Do not bundle changes to
-plugin hooks, the visual companion, unrelated skills, standalone-repo retirement,
-or the movie checker's general behavior. No new harness integration is proposed.
-
-Add a portable Python integration suite for the shared behavior. Shell launch
-tests should be thin wrappers over the same fixtures and assertions. Preserve
-existing regression coverage; replacing a Bash test requires equivalent coverage
-in the portable suite. The suite must be runnable locally and on OS runners
-without depending on a particular CI provider.
-
-## Acceptance and evidence
-
-| Area | Required proof |
-| --- | --- |
-| OS coverage | Actual runs on macOS, Linux, native Windows, and WSL; record exact environments and distinguish architecture coverage |
-| Launch shells | Bash and zsh on macOS, Bash on Linux/WSL, Windows PowerShell 5.1, PowerShell 7, and Git Bash on native Windows |
-| Shared pipeline | All scene kinds; measured scene offsets; hard and soft subtitles; checker verdicts and contact sheets |
-| Paths | Spaces, apostrophes, Unicode, metacharacters, Windows drive paths, relative/absolute work directories, nested SRT files, LF/CRLF |
-| Windows terminal | Git Bash and recorded PowerShell 5.1/7; a crossed launch/record-shell combination; cmdlet/native status attribution, delayed command, TUI input, and same-session proof between takes |
-| Native independence | PowerShell movie tools run with Git Bash and WSL unavailable to them; Git Bash native tests use Windows binaries |
-| Local voice | Real no-key Piper synthesis and local ASR in every required OS environment, including WSL; repeat with cached models and network disabled; test nested helper isolation from the filmed project |
-| Capture failure | Blank frame, missing display, permission refusal, browser loss, and timeout produce accurate outcomes and no fabricated proof |
-| Lifecycle | Survives a harness turn boundary; bounded shutdown after success, error, and interruption; owned children/grandchildren terminate while an unrelated sentinel survives |
-| Evidence integrity | Failed command remains failed through logging; missing required speech verification is reported; rendered content matches real actions |
-
-The end-to-end fixture includes a small real browser app with a persistent state
-change and a terminal sequence that prints Unicode, runs a delayed command, and
-returns a nonzero status. Capture those actions, generate real local narration,
-assemble `card`/`image`/`frames`/`movie` scenes, burn subtitles, run the checker, and
-inspect the finished output. Include a terminal TUI take and an unnarrated
-log-reel case. An intentionally failing command may be part of a successful
-recording test: its logged status and narration must accurately show that failure.
-Use synthetic clips for deterministic negative checker tests, clearly identified
-as test fixtures.
-
-Run the end-to-end pipeline from both Windows shell families. Cover PowerShell
-5.1 and 7 with real launch tests; test the OS pipeline with actual executables,
-not only mocked `sys.platform`. Optional display-specific tests may be skipped
-only with a recorded reason and a separate result for the fallback route. A
-missing prerequisite in a required acceptance job is a failure of that job's
-setup, not a green compatibility result.
-
-OS and feature coverage must be joined. Define fixture groups: **P** is the full
-processing fixture (all scene kinds, offsets, hard/soft subtitles, checker and
-negative cases); **B** is real browser capture; **T** is real terminal/TUI capture
-and lifecycle; **V** is no-key synthesis/transcription plus a cached offline
-rerun; **L** is the real-run log reel with preserved status.
-
-| Required environment | Minimum successful groups | Additional proof |
-| --- | --- | --- |
-| macOS arm64 | P, B, T, V, L | Bash and zsh invocation |
-| macOS x64 | P, B, T, V, L | Exact interpreter and model package versions |
-| Linux x64 desktop | P, B, T, V, L | Display type and browser backend |
-| Linux x64 with DISPLAY/WAYLAND_DISPLAY unset and no desktop session | P, B, T, V, L | Positive headless browser and terminal takes, not just a missing-display diagnostic |
-| Native Windows x64 | P, B, T, V, L | Full pipeline from PowerShell and Git Bash; real launch and recorded-shell coverage for PowerShell 5.1/7; native binaries proven |
-| WSL Linux x64 | P, B, T, V, L | Linux binaries proven; local voice runs inside WSL; no hidden Windows-file-path dependency |
-
-Do not multiply the full fixture by every launch-shell permutation: thin
-invocation tests cover remaining permutations, while the explicit Windows
-end-to-end requirements above still apply. Remote/harness continuity requires
-at least one real remote session for browser and terminal recording, with
-artifacts retrieved and the launch/recording hosts identified.
-
-To advertise a desktop backend as verified, retain one successful real desktop
-take on an actual macOS desktop, a Windows desktop using `gdigrab`, or a Linux
-X11 desktop using the documented backend, respectively. A skipped test leaves
-that backend conditional/unverified even if the OS's P/B/T/V/L groups pass.
-Wayland support in this scope means correct detection and an honest route
-decision; a Wayland desktop backend needs its own successful take before a
-stronger claim is published. This does not require every compositor or remote
-desktop configuration to support desktop capture.
-
-Skill instruction changes require `superpowers:writing-skills` and before/after
-pressure testing across multiple agent sessions. Scenarios include a PowerShell
-harness, a Git Bash harness, no cloud key, a path containing spaces, missing
-libass, and headless/blocked capture. Check that agents select valid commands,
-keep failures visible, and inspect the artifact. Store evals and transcripts in
-the project's external eval repository; it is absent from this checkout, so
-record its actual path and commit rather than assuming `evals/` is present.
-The external checkout found during review is
-`/Users/drewritter/prime-rad/superpowers-evals` at `66f08529`; this is an evidence
-location, not a path to hard-code into tests. Its Windows guest launcher proves
-where the agent runs, not which shell its tools use. For each shell-specific
-agent eval, capture the actual agent tool invocation and shell identity; launching
-an agent from PowerShell while all its commands execute in Bash does not satisfy
-the PowerShell-harness scenario.
-
-Evidence records must include OS/architecture, shell, Python/uv, FFmpeg build,
-browser, TTS/ASR versions, harness/model, command, return code, and artifact or
-transcript location. Generated media need not be committed to core, but the
-fixture and reproduction commands must be retained.
-
-### First implementation task: native Windows feasibility
-
-Before the bulk port, run a bounded probe of the chosen ttyd build with native
-PowerShell and Git Bash: one filmed/observed terminal, a command completion
-record, two takes preserving shell state across a harness boundary, and Job
-Object cleanup of a nested worker. Record executable versions and the observed
-ttyd message format. This is runtime validation of the specified design, not
-evidence already gathered by this spec review. A failure revisits the relevant
-adapter design before depending on it; it does not silently remove an agreed
-platform or shell from the contract.
-
-The human partner has provided `drew@ballmer.local` for SSH or Paseo access.
-A read-only SSH inventory succeeded on 2026-09-09: Windows 11 Pro, PowerShell
-5.1, Git Bash, Chrome, and Edge were found. PowerShell 7, uv, FFmpeg/ffprobe,
-and ttyd were not resolved on that session's PATH; their installation status
-needs checking before provisioning. This establishes an available native
-validation host, not a passed recording test or verified desktop access.
-The [implementation plan](../plans/2026-09-09-proof-movie-os-compatibility.md)
-records the observed paths, prerequisite checks, and Ballmer-first probe.
-
-The design has an implementation plan ready for review. Compatibility is complete
-only when the required matrix has evidence, the updated skill has behavior evals,
-and the support documentation accurately distinguishes verified combinations
-from dependency or display limitations.
-
-## External references
-
-- [uv explicit script invocation](https://docs.astral.sh/uv/reference/cli/#uv-run--script)
-  supports extensionless Python scripts.
-- [uv platform policy](https://docs.astral.sh/uv/reference/policies/platforms/)
-  distinguishes tested platforms from build-only support; it is not a substitute
-  for validating the rest of the movie toolchain.
-- [FFmpeg image input](https://ffmpeg.org/ffmpeg-formats.html#image2)
-  documents that glob input depends on build support.
-- [FFmpeg Windows desktop capture](https://ffmpeg.org/ffmpeg-devices.html#gdigrab)
-  documents desktop, region, and window capture.
-- [ttyd Windows builds](https://github.com/tsl0922/ttyd/wiki/Compile-on-Windows)
-  document native ConPTY support and its OS requirement.
-- [Piper package files](https://pypi.org/project/piper-tts/#files)
-  include Windows x64 builds; package availability alone is not a completed
-  synthesis or transcription test.
-- [PowerShell automatic variables](https://learn.microsoft.com/en-us/powershell/module/microsoft.powershell.core/about/about_automatic_variables)
-  distinguish command success from native process exit status.
-- [ttyd protocol implementation](https://github.com/tsl0922/ttyd/blob/main/src/protocol.c)
-  associates a terminal process with each initialized client connection.
-- [CDP WebSocket receive events](https://chromedevtools.github.io/devtools-protocol/tot/Network/#event-webSocketFrameReceived)
-  expose messages from the browser's existing socket, including encoded binary payloads.
-- [Windows Job Objects](https://learn.microsoft.com/en-us/windows/win32/procthread/job-objects)
-  provide a process-ownership boundary for session cleanup.
-- [ConPTY close behavior](https://learn.microsoft.com/en-us/windows/console/closepseudoconsole)
-  informs output draining and bounded shutdown on older Windows versions.

+ 0 - 75
docs/superpowers/specs/2026-09-09-proof-movie-os-compatibility-review.md

@@ -1,75 +0,0 @@
-# Proof movie compatibility: adversarial design review
-
-**Date:** 2026-09-09
-
-**Reviewed draft:** `a36d831d6f9764b6a1355b23f318519701f37029`
-
-**Design:** [OS compatibility spec](2026-09-09-proof-movie-os-compatibility-design.md)
-
-## Method and limits
-
-Two independent agents reviewed the committed draft with separate contexts and
-read-only assignments. `review_terminal_spec` examined native terminal control,
-shell outcomes, ttyd, and lifecycle. `review_platform_spec` examined the platform
-contract, dependency feasibility, and acceptance evidence. The author separately
-checked the shared pipeline against the actual imported scripts and verified the
-review findings before revising the design.
-
-The independent reviewers found no Critical issues and four Important design
-gaps. They did not run native Windows recording. This is adversarial review of a
-specification, not skill pressure testing or OS acceptance testing.
-
-## Findings and disposition
-
-| Finding | Evidence/failure scenario | Design revision |
-| --- | --- | --- |
-| T1: PowerShell outcome was underspecified (Important) | A failed cmdlet can leave an earlier native exit code at zero, or a successful cmdlet can leave it nonzero. PowerShell 5.1 expression wrappers can alter the success variable. Native-exit-only tests miss this. | Separate shell success, native exit attribution, errors, and unknown/interrupted completion. Capture status before logging; define atomic beats and opaque script behavior. Record both PowerShell 5.1 and 7 with mixed cmdlet/native and error fixtures. |
-| T2: Process ownership omitted descendants (Important) | ttyd terminates its Windows shell with `TerminateProcess`; this does not by itself prove children/grandchildren stop. A worker can keep writing after a cancelled recording. | Establish a Windows Job Object before children run, with kill-on-close and no breakaway; account for Unix PTY process groups. Drain output during bounded shutdown. Require owned descendants to exit while an unrelated sentinel survives. |
-| T3: Output observation could attach to the wrong shell (Important) | ttyd spawns a process per initialized terminal WebSocket. Opening a separate observer socket creates another shell rather than attaching to the filmed one. | One persistent supervisor and one terminal client. Observe the filmed page's existing connection through CDP, use writable mode and a readiness round trip, and control takes through ordered request files. Test nonce, shell identity, state, and command continuity. |
-| P1: OS and feature coverage were not joined (Important) | The original acceptance rows could be satisfied by proving features on one OS and doing smoke runs elsewhere, leaving genuinely headless Linux or WSL local voice untested. | Assign P/B/T/V/L fixture groups to each required OS/architecture environment. Require positive headless and WSL voice runs, separately map shell checks, and reserve verified desktop claims for backends with real successful takes. |
-| A1: Nested ASR helper inherited the filmed project (author finding) | Imported `narrate` uses nested `uv run --with faster-whisper ...` without disabling project discovery. An unrelated project Python requirement can prevent the helper from launching. | Disable nested helper project discovery; require compatible interpreter selection and a fixture proving no dependency on, or mutation of, the filmed project's environment/lockfile. |
-
-Additional clarifications retain the existing `movie` scene's own-audio/duration
-exception, use the exact scene-kind names in the fixture, distinguish helper
-UTF-8 pipes from arbitrary native command encodings, and require shell-specific
-agent evals to prove the actual tool shell rather than the agent's launch shell.
-
-## Verification performed during review
-
-- Inspected ttyd's `src/protocol.c`: process creation is associated with the
-  initialized terminal client, and input depends on writable mode.
-- Inspected ttyd's `src/pty.c`: Windows shell termination uses `TerminateProcess`.
-- Checked Microsoft's PowerShell automatic-variable documentation, Job Objects,
-  and ConPTY shutdown documentation. These support the design corrections; they
-  do not establish that the future recorder implements them correctly.
-- Reproduced uv project discovery in a temporary directory containing a
-  `pyproject.toml` requiring Python `>=9.99`.
-  `uv run --offline python -c 'print("ASR child launched")'` exited 2 with an
-  interpreter-resolution error; adding `--no-project` before `python` exited 0.
-  This isolates the project-discovery issue without claiming an actual ASR/model
-  test was run.
-- Located the external eval checkout at
-  `/Users/drewritter/prime-rad/superpowers-evals`, commit `66f08529`, and read its
-  Windows guest agent instructions. They establish a native guest launch path,
-  not a PowerShell tool adapter or a completed movie eval.
-
-## Recheck and remaining work
-
-Both original reviewers rechecked the revised design. The terminal reviewer
-confirmed T1/T2/T3 resolved for implementation planning with no new Critical or
-Important contradictions. The platform reviewer confirmed P1 resolved with no
-new findings. The terminal reviewer also checked CDP's documented receive events
-and binary payload representation and found no demonstrated accessibility blocker
-for the proposed observer. No independent review finding remains open.
-
-Runtime validation remains implementation work: the first planned task is a
-bounded native Windows probe of same-session ttyd observation, PowerShell/Git Bash
-completion, persistent takes, and descendant cleanup. Its result determines
-whether the specified adapters can proceed unchanged. The
-[implementation plan](../plans/2026-09-09-proof-movie-os-compatibility.md) is now
-written for review and records the available Windows validation host. The runtime
-probe is a prerequisite within that plan, not a test claimed to have passed
-during specification review.
-
-The eventual skill changes still require before/after pressure tests and the
-OS acceptance matrix. This review must not be cited as that evidence.

+ 0 - 36
docs/superpowers/specs/2026-09-09-proof-movie-windows-completion-review.md

@@ -1,36 +0,0 @@
-# Adversarial review of the Windows completion spec
-
-**Reviewed draft:** `cd5e0bd7`.
-**Spec:** [Windows completion design](2026-09-09-proof-movie-windows-completion-design.md).
-**Reviewer:** Independent Codex subagent `/root/review_windows_completion`, given the user goal, current spec, and relevant source paths without the controller's full conversation.
-**Scope:** Design correctness, scope discipline, existing-code compatibility, and sufficient Windows acceptance. Read-only review; no implementation, runtime tests, provisioning, remote-host work, or additional agents.
-
-## Initial verdict
-
-Revise before implementation. The reviewer judged the narrowed architecture proportionate and found five mandatory contract corrections. None required restoring the old framework, platform matrix, or eval infrastructure.
-
-| Finding | Evidence in the reviewed source | Spec correction |
-| --- | --- | --- |
-| R1 — Producer attribution and success (P1) | Probe `dispatch` depends on `native_producer`; Bash `emit_bash` selects the first saved pipeline status. The draft required attribution but omitted the input field and aggregate success rule. | Include the field; limit its promise to a direct native invocation or first pipeline stage followed by logging. Keep raw PowerShell status and observed errors distinct. A failed/unknown identified producer cannot yield a successful command result. |
-| R2 — Cached ASR verification (P1) | `narrate`'s unchanged-WAV/manifest branch bypasses all transcription, including under `--verify on`. The controller independently identified and sent this case to the reviewer, who confirmed it. | Explicit `on` verifies reused WAVs as well as new audio. Add the off→on cached-clip case to the focused checks; unchanged text does not establish verified audio. |
-| R3 — Complete asynchronous control (P2) | Probe reads exactly the next consecutive request filename; arbitrary positive IDs can block it. Duplicate rejection prevents resubmitting solely to wait for an existing result. | Specify consecutive IDs, one controller, visible next ID, and prompt gap/duplicate rejection. Add a wait-only `result` command; distinguish acknowledgment from completion and client timeout from cancellation. |
-| R4 — Timeout/shutdown state (P2) | Probe command timeout terminates the session, but the draft left this ambiguous. Probe dispatch publishes close/cancel replies before cleanup. | Command timeout records unknown and tears down the session. Define active-take/pending-command behavior for close/cancel. Successful shutdown completion is published only after resource cleanup, separately from acknowledgment. |
-| R5 — Event timing and automatic capture (P2) | Draft allowed nearly two-second duplicate intervals while requiring only total-duration accuracy; a transient could be omitted without failing that duration check. | Specify monotonic boundaries, capture request/completion timestamps, conservative placement on the output grid, and visible gaps. Require several ordered TUI states in automatic frames before exit; a manual checkpoint cannot substitute. |
-
-The controller checked R1–R4 against the cited implementation paths and R5 against the actual capture/export contract before editing. The revisions add focused checks inside the existing Windows validation work, not new infrastructure.
-
-## Boundaries retained
-
-- Fixed geometry removes runtime resize support. It does not prove wrapping/redraw correctness.
-- Periodic screenshots are a proposed implementation choice. Existing successful single screenshots do not establish sustained capture correctness.
-- A focused actual Windows cadence/wrapping test remains required before further recorder extraction. Its failure requires a bounded design decision, not automatic expansion into the old plan.
-- The three Windows shell runs may share downloaded model caches and fixture setup, but each must produce its own final-code evidence.
-- Existing agents remain stopped. This newly authorized reviewer performs review only.
-
-## Scoped recheck
-
-The same reviewer rechecked R1–R5 and contradictions introduced by their revisions. **All five are resolved sufficiently for implementation; no new concrete blocker or mandatory scope reduction was found.** The revised design is ready to guide implementation when execution is authorized.
-
-The reviewer found no evidence that fixed-geometry periodic screenshots fail as a design. Sustained capture and wrapping correctness remain unproven runtime questions covered by the focused Windows decision gate. Design approval does not establish Windows support or replace that test.
-
-Both review passes were read-only. The controller changed only this review record and the completion spec; implementation did not resume.

+ 0 - 1510
tests/proving-it-works-with-a-movie/probe-windows.py

@@ -1,1510 +0,0 @@
-#!/usr/bin/env python3
-# /// script
-# requires-python = ">=3.12"
-# dependencies = ["websocket-client==1.9.0"]
-# ///
-"""Native Windows terminal feasibility probe; a missing observation fails the gate."""
-import argparse
-import base64
-import codecs
-import ctypes
-import hashlib
-import json
-import os
-import platform
-import re
-import socket
-import struct
-import subprocess
-import sys
-import time
-import traceback
-import urllib.request
-import uuid
-from pathlib import Path
-
-
-def assert_probe(result):
-    assert result["terminal_client_count"] == 1
-    assert result["observed_session_id"] == result["filmed_session_id"]
-    assert result["nonce_before"] == result["nonce_after"]
-    assert result["shell_pid_before"] == result["shell_pid_after"]
-    assert result["variable_after"] == "persisted-λ"
-    assert result["long_command_survived_take_boundary"]
-    assert result["owned_children_remaining"] == []
-    assert result["unrelated_sentinel_alive"]
-    assert result["shutdown_seconds"] < 10
-
-
-def assert_gate(result):
-    assert set(result["shells"]) == {"powershell51", "powershell7", "gitbash"}
-    for shell in result["shells"].values():
-        assert_probe(shell)
-    assert set(result["checks"]) == {"prerequisites"} | {
-        f"{shell}.{check}" for shell in result["shells"]
-        for check in (*REQUIRED, "ordinary_user", "interactive", "bounded_resize", "capture_artifacts")}
-    assert all(check["status"] == "passed" for check in result["checks"].values())
-
-
-def assert_outcomes(records):
-    expected = {
-        "native_success": (True, 0),
-        "cmdlet_failure": (False, None),
-        "native_failure": (False, 7),
-        "cmdlet_success": (True, None),
-        "terminating_error": (False, None),
-        "nonterminating_error": (False, None),
-        "parse_failure": (False, None),
-    }
-    for name, (success, native_exit) in expected.items():
-        record = records[name]
-        assert record["outcome"] == "completed", (name, record)
-        assert record["shell_success"] is success, (name, record)
-        assert record["native_exit_code"] == native_exit, (name, record)
-    for name in ("cmdlet_failure", "terminating_error", "nonterminating_error", "parse_failure"):
-        assert records[name]["shell_error"], (name, records[name])
-    assert records["logging_failure"]["producer_exit_code"] == 7
-    assert records["logging_failure"]["native_exit_code"] == 7
-    assert records["logging_failure"]["producer_success"] is False
-
-
-def file_handoff(operation):
-    # A Windows reader can overlap an atomic replacement. Retry the operation,
-    # never substitute old/empty evidence; persistent access errors still fail.
-    deadline = time.monotonic() + 0.5
-    while True:
-        try:
-            return operation()
-        except PermissionError:
-            if time.monotonic() >= deadline:
-                raise
-            time.sleep(0.01)
-
-
-def read_json(path):
-    return json.loads(file_handoff(lambda: path.read_text(encoding="utf-8")))
-
-
-def write_json(path, value):
-    temporary = path.with_suffix(".tmp")
-    temporary.write_text(json.dumps(value, ensure_ascii=False, indent=2), encoding="utf-8")
-    file_handoff(lambda: temporary.replace(path))
-
-
-def token_info():
-    """Query this process's actual token, without changing it."""
-    from ctypes import wintypes as W
-
-    kernel = ctypes.WinDLL("kernel32", use_last_error=True)
-    security = ctypes.WinDLL("advapi32", use_last_error=True)
-    kernel.GetCurrentProcess.argtypes, kernel.GetCurrentProcess.restype = [], W.HANDLE
-    kernel.CloseHandle.argtypes, kernel.CloseHandle.restype = [W.HANDLE], W.BOOL
-    security.OpenProcessToken.argtypes = [W.HANDLE, W.DWORD, ctypes.POINTER(W.HANDLE)]
-    security.OpenProcessToken.restype = W.BOOL
-    security.GetTokenInformation.argtypes = [W.HANDLE, ctypes.c_int, ctypes.c_void_p, W.DWORD, ctypes.POINTER(W.DWORD)]
-    security.GetTokenInformation.restype = W.BOOL
-    security.GetSidSubAuthorityCount.argtypes = [ctypes.c_void_p]
-    security.GetSidSubAuthorityCount.restype = ctypes.POINTER(W.BYTE)
-    security.GetSidSubAuthority.argtypes = [ctypes.c_void_p, W.DWORD]
-    security.GetSidSubAuthority.restype = ctypes.POINTER(W.DWORD)
-    token = W.HANDLE()
-    if not security.OpenProcessToken(kernel.GetCurrentProcess(), 8, ctypes.byref(token)):
-        raise ctypes.WinError(ctypes.get_last_error())
-    try:
-        def query(kind):
-            length = W.DWORD()
-            security.GetTokenInformation(token, kind, None, 0, ctypes.byref(length))
-            buffer = ctypes.create_string_buffer(length.value)
-            if not security.GetTokenInformation(token, kind, buffer, len(buffer), ctypes.byref(length)):
-                raise ctypes.WinError(ctypes.get_last_error())
-            return buffer
-
-        integrity = query(25)
-        sid = ctypes.cast(integrity, ctypes.POINTER(ctypes.c_void_p))[0]
-        count = security.GetSidSubAuthorityCount(sid)[0]
-        rid = security.GetSidSubAuthority(sid, count - 1)[0]
-        result = {"pid": os.getpid(), "elevated": bool(W.DWORD.from_buffer(query(20)).value),
-                  "elevation_type": W.DWORD.from_buffer(query(18)).value,
-                  "integrity_rid": rid, "integrity_sid": f"S-1-16-{rid}",
-                  "desktop_session": W.DWORD.from_buffer(query(12)).value}
-        result["ordinary_user"] = not result["elevated"] and rid == 8192
-        whoami = Path(os.environ["SystemRoot"]) / "System32" / "whoami.exe"
-        result["whoami_groups"] = subprocess.run([str(whoami), "/all"], capture_output=True, check=True).stdout.decode("utf-8", "replace")
-        return result
-    finally:
-        kernel.CloseHandle(token)
-
-
-def ttyd_output(event: dict) -> bytes:
-    frame = event["params"]["response"]
-    payload = frame["payloadData"]
-    raw = base64.b64decode(payload) if frame["opcode"] == 2 else payload.encode("utf-8")
-    return raw[1:] if raw[:1] == b"0" else b""
-
-
-class CompletionParser:
-    """Parse bounded multiline records; this is not a VT redraw emulator."""
-
-    def __init__(self):
-        self.decoder = codecs.getincrementaldecoder("utf-8")("replace")
-        self.text = ""
-        self.records = []
-
-    def feed(self, data):
-        self.text += self.decoder.decode(data)
-        clean = re.sub(r"\x1b\][^\x07]*(?:\x07|\x1b\\)|\x1b\[[0-?]*[ -/]*[@-~]", "", self.text)
-        for match in re.finditer(r"\[PROBE\|([A-Za-z0-9+/=\s]+)\|END\]", clean):
-            record = json.loads(base64.b64decode(re.sub(r"\s", "", match[1]), validate=True).decode("utf-8"))
-            if record not in self.records:
-                self.records.append(record)
-
-
-class WindowsJob:
-    """Suspended roots enter an unnamed, non-inheritable job before resuming."""
-
-    def __init__(self):
-        from ctypes import wintypes as W
-
-        if sys.platform != "win32" or struct.calcsize("P") != 8:
-            raise RuntimeError("This probe requires native Windows x64 Python")
-        U64, SIZE = ctypes.c_ulonglong, ctypes.c_size_t
-
-        class STARTUPINFOW(ctypes.Structure):
-            _fields_ = [("cb", W.DWORD), ("lpReserved", W.LPWSTR),
-                        ("lpDesktop", W.LPWSTR), ("lpTitle", W.LPWSTR),
-                        ("dwX", W.DWORD), ("dwY", W.DWORD), ("dwXSize", W.DWORD),
-                        ("dwYSize", W.DWORD), ("dwXCountChars", W.DWORD),
-                        ("dwYCountChars", W.DWORD), ("dwFillAttribute", W.DWORD),
-                        ("dwFlags", W.DWORD), ("wShowWindow", W.WORD),
-                        ("cbReserved2", W.WORD), ("lpReserved2", ctypes.POINTER(W.BYTE)),
-                        ("hStdInput", W.HANDLE), ("hStdOutput", W.HANDLE), ("hStdError", W.HANDLE)]
-
-        class PROCESS_INFORMATION(ctypes.Structure):
-            _fields_ = [("hProcess", W.HANDLE), ("hThread", W.HANDLE),
-                        ("dwProcessId", W.DWORD), ("dwThreadId", W.DWORD)]
-
-        class BASIC_LIMIT(ctypes.Structure):
-            _fields_ = [("PerProcessUserTimeLimit", ctypes.c_longlong),
-                        ("PerJobUserTimeLimit", ctypes.c_longlong), ("LimitFlags", W.DWORD),
-                        ("MinimumWorkingSetSize", SIZE), ("MaximumWorkingSetSize", SIZE),
-                        ("ActiveProcessLimit", W.DWORD), ("Affinity", SIZE),
-                        ("PriorityClass", W.DWORD), ("SchedulingClass", W.DWORD)]
-
-        class IO_COUNTERS(ctypes.Structure):
-            _fields_ = [(name, U64) for name in ("ReadOperationCount", "WriteOperationCount",
-                        "OtherOperationCount", "ReadTransferCount", "WriteTransferCount", "OtherTransferCount")]
-
-        class EXTENDED_LIMIT(ctypes.Structure):
-            _fields_ = [("BasicLimitInformation", BASIC_LIMIT), ("IoInfo", IO_COUNTERS),
-                        ("ProcessMemoryLimit", SIZE), ("JobMemoryLimit", SIZE),
-                        ("PeakProcessMemoryUsed", SIZE), ("PeakJobMemoryUsed", SIZE)]
-
-        self.sizes = {"STARTUPINFOW": ctypes.sizeof(STARTUPINFOW),
-                      "PROCESS_INFORMATION": ctypes.sizeof(PROCESS_INFORMATION),
-                      "BASIC_LIMIT": ctypes.sizeof(BASIC_LIMIT),
-                      "IO_COUNTERS": ctypes.sizeof(IO_COUNTERS),
-                      "EXTENDED_LIMIT": ctypes.sizeof(EXTENDED_LIMIT)}
-        assert list(self.sizes.values()) == [104, 24, 64, 48, 144], self.sizes
-        self.SI, self.PI = STARTUPINFOW, PROCESS_INFORMATION
-        self.k = ctypes.WinDLL("kernel32", use_last_error=True)
-        signatures = {
-            "CreateJobObjectW": ([ctypes.c_void_p, W.LPCWSTR], W.HANDLE),
-            "SetInformationJobObject": ([W.HANDLE, ctypes.c_int, ctypes.c_void_p, W.DWORD], W.BOOL),
-            "QueryInformationJobObject": ([W.HANDLE, ctypes.c_int, ctypes.c_void_p, W.DWORD, ctypes.POINTER(W.DWORD)], W.BOOL),
-            "CreateProcessW": ([W.LPCWSTR, W.LPWSTR, ctypes.c_void_p, ctypes.c_void_p, W.BOOL,
-                                W.DWORD, ctypes.c_void_p, W.LPCWSTR, ctypes.POINTER(STARTUPINFOW),
-                                ctypes.POINTER(PROCESS_INFORMATION)], W.BOOL),
-            "AssignProcessToJobObject": ([W.HANDLE, W.HANDLE], W.BOOL),
-            "ResumeThread": ([W.HANDLE], W.DWORD),
-            "TerminateProcess": ([W.HANDLE, W.UINT], W.BOOL),
-            "TerminateJobObject": ([W.HANDLE, W.UINT], W.BOOL),
-            "CloseHandle": ([W.HANDLE], W.BOOL),
-            "WaitForSingleObject": ([W.HANDLE, W.DWORD], W.DWORD),
-            "GetExitCodeProcess": ([W.HANDLE, ctypes.POINTER(W.DWORD)], W.BOOL),
-            "OpenProcess": ([W.DWORD, W.BOOL, W.DWORD], W.HANDLE),
-            "GetProcessTimes": ([W.HANDLE] + [ctypes.POINTER(W.FILETIME)] * 4, W.BOOL),
-            "IsProcessInJob": ([W.HANDLE, W.HANDLE, ctypes.POINTER(W.BOOL)], W.BOOL),
-        }
-        for name, (arguments, returns) in signatures.items():
-            function = getattr(self.k, name)
-            function.argtypes, function.restype = arguments, returns
-        self.handle = self.k.CreateJobObjectW(None, None)
-        if not self.handle:
-            raise ctypes.WinError(ctypes.get_last_error())
-        limits = EXTENDED_LIMIT()
-        limits.BasicLimitInformation.LimitFlags = 0x2000  # KILL_ON_JOB_CLOSE; no breakaway
-        if not self.k.SetInformationJobObject(self.handle, 9, ctypes.byref(limits), ctypes.sizeof(limits)):
-            error = ctypes.get_last_error()
-            self.k.CloseHandle(self.handle)
-            raise ctypes.WinError(error)
-        self.roots = []
-
-    def spawn(self, argv, directory, log, env=None):
-        import msvcrt
-
-        if not Path(argv[0]).is_absolute():
-            raise ValueError("An absolute executable path is required")
-        environment = dict(os.environ) if env is None else env
-        block = ctypes.create_unicode_buffer("\0".join(f"{k}={v}" for k, v in sorted(environment.items())) + "\0\0")
-        si, pi = self.SI(), self.PI()
-        si.cb, si.dwFlags = ctypes.sizeof(si), 0x100  # STARTF_USESTDHANDLES
-        with open(os.devnull, "rb") as stdin, log.open("ab", buffering=0) as output:
-            handles = [msvcrt.get_osfhandle(f.fileno()) for f in (stdin, output)]
-            for handle in handles:
-                os.set_handle_inheritable(handle, True)
-            si.hStdInput, si.hStdOutput, si.hStdError = handles[0], handles[1], handles[1]
-            try:
-                created = self.k.CreateProcessW(argv[0], ctypes.create_unicode_buffer(subprocess.list2cmdline(argv)),
-                    None, None, True, 0x404, block, str(directory), ctypes.byref(si), ctypes.byref(pi))
-            finally:
-                for handle in handles:
-                    os.set_handle_inheritable(handle, False)
-        if not created:
-            raise ctypes.WinError(ctypes.get_last_error())
-        try:
-            if not self.k.AssignProcessToJobObject(self.handle, pi.hProcess):
-                raise ctypes.WinError(ctypes.get_last_error())
-            if self.k.ResumeThread(pi.hThread) == 0xFFFFFFFF:
-                raise ctypes.WinError(ctypes.get_last_error())
-        except BaseException:
-            self.k.TerminateProcess(pi.hProcess, 1)
-            self.k.WaitForSingleObject(pi.hProcess, 5000)
-            self.k.CloseHandle(pi.hProcess)
-            raise
-        finally:
-            self.k.CloseHandle(pi.hThread)
-        self.roots.append((pi.dwProcessId, pi.hProcess))
-        return pi.dwProcessId
-
-    def pids(self):
-        from ctypes import wintypes as W
-
-        class PROCESS_LIST(ctypes.Structure):
-            _fields_ = [("assigned", W.DWORD), ("count", W.DWORD), ("pids", ctypes.c_size_t * 1024)]
-
-        info = PROCESS_LIST()
-        if not self.k.QueryInformationJobObject(self.handle, 3, ctypes.byref(info), ctypes.sizeof(info), None):
-            raise ctypes.WinError(ctypes.get_last_error())
-        return list(info.pids[:info.count])
-
-    def process_time(self, handle):
-        from ctypes import wintypes as W
-
-        times = [W.FILETIME() for _ in range(4)]
-        if not self.k.GetProcessTimes(handle, *(ctypes.byref(value) for value in times)):
-            raise ctypes.WinError(ctypes.get_last_error())
-        return (times[0].dwHighDateTime << 32) | times[0].dwLowDateTime
-
-    def open_process(self, pid, creation=None, terminate=False):
-        handle = self.k.OpenProcess(0x100000 | 0x1000 | int(terminate), False, pid)
-        if not handle:
-            raise ctypes.WinError(ctypes.get_last_error())
-        if creation is not None and self.process_time(handle) != creation:
-            self.k.CloseHandle(handle)
-            raise RuntimeError("Process creation time changed; refusing stale PID")
-        return handle
-
-    def snapshot(self):
-        from ctypes import wintypes as W
-
-        processes = []
-        try:
-            for pid in self.pids():
-                try:
-                    handle = self.open_process(pid)
-                except OSError as error:
-                    if error.winerror == 87:  # Exited between enumeration and open.
-                        continue
-                    raise
-                owned = W.BOOL()
-                if not self.k.IsProcessInJob(handle, self.handle, ctypes.byref(owned)) or not owned.value:
-                    self.k.CloseHandle(handle)
-                    raise RuntimeError("Process is no longer a member of the owned job")
-                processes.append({"pid": pid, "handle": handle, "creation": self.process_time(handle)})
-            return processes
-        except BaseException:
-            for process in processes:
-                self.k.CloseHandle(process["handle"])
-            raise
-
-    def close(self):
-        if not self.handle:
-            return
-        try:
-            if not self.k.TerminateJobObject(self.handle, 1):
-                raise ctypes.WinError(ctypes.get_last_error())
-            deadline = time.monotonic() + 5
-            while self.pids() and time.monotonic() < deadline:
-                time.sleep(0.05)
-            if self.pids():
-                raise TimeoutError("Owned processes remain after job termination")
-        finally:
-            self.k.CloseHandle(self.handle)
-            self.handle = None
-            for _, handle in self.roots:
-                self.k.CloseHandle(handle)
-
-
-def free_port():
-    with socket.socket() as sock:
-        sock.bind(("127.0.0.1", 0))
-        return sock.getsockname()[1]
-
-
-def http_json(url):
-    with urllib.request.urlopen(url, timeout=1) as response:
-        return json.load(response)
-
-
-class CDP:
-    def __init__(self, url, on_event, trace_path):
-        import websocket
-
-        self.ws = websocket.create_connection(url, timeout=5, suppress_origin=True)
-        self.ws.settimeout(0.1)
-        self.on_event, self.counter, self.responses = on_event, 0, {}
-        self.trace = trace_path.open("w", encoding="utf-8")
-
-    def log(self, direction, **fields):
-        self.trace.write(json.dumps({"time": time.time(), "direction": direction, **fields}) + "\n")
-        self.trace.flush()
-
-    def send(self, method, params=None):
-        self.counter += 1
-        self.ws.send(json.dumps({"id": self.counter, "method": method, "params": params or {}}))
-        self.log("sent", id=self.counter, method=method,
-                 frame_session_id=(params or {}).get("sessionId"))
-        return self.counter
-
-    def pump(self):
-        import websocket
-
-        try:
-            raw = self.ws.recv()
-        except websocket.WebSocketTimeoutException:
-            self.log("receive_timeout")
-            return
-        if not raw:
-            raise ConnectionError("Browser CDP connection closed")
-        event = json.loads(raw)
-        self.log("received", id=event.get("id"), method=event.get("method"),
-                 frame_session_id=event.get("params", {}).get("sessionId"))
-        if "id" in event:
-            self.responses[event["id"]] = event
-        else:
-            self.on_event(event)
-
-    def call(self, method, params=None, timeout=5):
-        request = self.send(method, params)
-        deadline = time.monotonic() + timeout
-        while request not in self.responses and time.monotonic() < deadline:
-            self.pump()
-        response = self.responses.pop(request, None)
-        if response is None or "error" in response:
-            raise RuntimeError(f"CDP {method}: {response}")
-        return response.get("result", {})
-
-
-class Terminal:
-    def __init__(self, args, directory):
-        self.directory, self.args = directory, args
-        directory.mkdir(parents=True, exist_ok=False)
-        self.session = uuid.uuid4().hex
-        self.parser = CompletionParser()
-        self.request_id = None
-        self.socket_ids = []
-        self.closed = False
-        self.frames = []
-        self.cdp = None
-        self.raw = (directory / "network.jsonl").open("w", encoding="utf-8")
-        self.output = (directory / "terminal.bin").open("wb")
-        self.job = WindowsJob()
-        self.launches = []
-        self.take = None
-        self.take_frames = []
-        self.terminal_sizes = []
-
-    def event(self, event):
-        method, params = event["method"], event.get("params", {})
-        if method == "Page.screencastFrame":
-            if self.take is not None:
-                frame_path = self.directory / self.take / f"{len(self.take_frames):06d}.jpg"
-                frame_path.write_bytes(base64.b64decode(params["data"]))
-                self.take_frames.append({"path": str(frame_path), "received_at": time.time(),
-                                         "metadata": params["metadata"], "page_id": self.page_id,
-                                         "websocket_id": self.request_id, "session_id": self.session})
-                write_json(self.directory / self.take / "frames.json", self.take_frames)
-            self.cdp.send("Page.screencastFrameAck", {"sessionId": params["sessionId"]})
-            return
-        if method.startswith("Network.webSocket"):
-            self.raw.write(json.dumps(event, ensure_ascii=False) + "\n")
-            self.raw.flush()
-        if method == "Network.webSocketCreated" and params["url"] == self.terminal_url.replace("http:", "ws:") + "ws":
-            self.socket_ids.append(params["requestId"])
-            if self.request_id is None:
-                self.request_id = params["requestId"]
-            else:
-                self.closed = True  # A reconnect is never continuity.
-        if params.get("requestId") != self.request_id or self.request_id is None:
-            return
-        if method == "Network.webSocketClosed":
-            self.closed = True
-        if method == "Network.webSocketFrameSent":
-            frame = params["response"]
-            raw = base64.b64decode(frame["payloadData"]) if frame["opcode"] == 2 else frame["payloadData"].encode("utf-8")
-            payload = raw[1:] if raw[:1] == b"1" else raw
-            if payload[:1] == b"{":
-                size = json.loads(payload)
-                if "columns" in size and "rows" in size:
-                    self.terminal_sizes.append({"columns": size["columns"], "rows": size["rows"],
-                                                "observed_at": time.time(), "request_id": self.request_id})
-        if method == "Network.webSocketFrameReceived":
-            frame = params["response"]
-            raw = base64.b64decode(frame["payloadData"]) if frame["opcode"] == 2 else frame["payloadData"].encode("utf-8")
-            self.frames.append({"opcode": frame["opcode"], "leading_byte": raw[:1].hex(), "bytes": len(raw)})
-            data = ttyd_output(event)
-            self.output.write(data)
-            self.output.flush()
-            self.parser.feed(data)
-
-    def start(self):
-        port, debug_port = free_port(), free_port()
-        self.terminal_url = f"http://127.0.0.1:{port}/"
-        shell_argv = [str(self.args.shell)] + (["--noprofile", "--norc", "-i"] if self.args.shell_kind == "gitbash" else ["-NoLogo", "-NoProfile", "-NoExit"])
-        # ttyd 1.7.7 leaves its ConPTY cwd pointer uninitialized without -w.
-        ttyd_argv = [str(self.args.ttyd), "-i", "127.0.0.1", "-p", str(port), "-W", "-m", "1",
-                     "-w", str(self.directory)] + shell_argv
-        browser_argv = [str(self.args.browser), "--headless=new", "--no-first-run", "--no-default-browser-check",
-                        "--disable-background-networking", "--remote-debugging-address=127.0.0.1",
-                        f"--remote-debugging-port={debug_port}", f"--user-data-dir={self.directory / 'profile'}",
-                        "--window-size=1600,900", "about:blank"]
-        for name, argv in (("ttyd", ttyd_argv), ("browser", browser_argv)):
-            pid = self.job.spawn(argv, self.directory, self.directory / f"{name}.log")
-            self.launches.append({"name": name, "argv": argv, "pid": pid})
-        write_json(self.directory / "launches.json", self.launches)
-        deadline = time.monotonic() + 20
-        last_error = None
-        while time.monotonic() < deadline:
-            try:
-                pages = http_json(f"http://127.0.0.1:{debug_port}/json/list")
-                page = next(p for p in pages if p["type"] == "page")
-                self.cdp = CDP(page["webSocketDebuggerUrl"], self.event, self.directory / "cdp-trace.jsonl")
-                break
-            except (OSError, StopIteration) as error:
-                last_error = error
-                time.sleep(0.1)
-        if self.cdp is None:
-            raise TimeoutError(f"Browser startup: {last_error}")
-        self.page_id = page["id"]
-        self.cdp.call("Network.enable")  # Must precede navigation and terminal socket creation.
-        self.cdp.call("Page.enable")
-        self.cdp.call("Page.navigate", {"url": self.terminal_url})
-        deadline = time.monotonic() + 10
-        while time.monotonic() < deadline:
-            self.cdp.pump()
-            if self.parser.text or self.closed:
-                break
-        self.screenshot("startup.png")
-        if self.closed:
-            raise ConnectionError("Terminal WebSocket closed before readiness")
-        if not self.parser.text:
-            raise TimeoutError("No ttyd output from the filmed terminal within 10 seconds")
-
-    def screenshot(self, name):
-        result = self.cdp.call("Page.captureScreenshot", {"format": "png"})
-        (self.directory / name).write_bytes(base64.b64decode(result["data"]))
-
-    def type(self, text):
-        self.cdp.call("Runtime.evaluate", {"expression": "document.querySelector('.xterm-helper-textarea').focus()"})
-        self.cdp.call("Input.insertText", {"text": text})
-        self.cdp.call("Input.dispatchKeyEvent", {"type": "keyDown", "key": "Enter", "code": "Enter", "windowsVirtualKeyCode": 13, "text": "\r"})
-        self.cdp.call("Input.dispatchKeyEvent", {"type": "keyUp", "key": "Enter", "code": "Enter", "windowsVirtualKeyCode": 13})
-
-    def readiness(self):
-        if self.args.shell_kind == "gitbash":
-            # Native Windows Python is explicit even inside Git Bash.
-            python = str(Path(sys.executable)).replace("\\", "/")
-            code = "import base64,json,os;print('[PROBE|'+base64.b64encode(json.dumps(dict(session=os.environ['PROBE_SESSION'],pid=os.environ['PROBE_SHELL_PID'],shell='gitbash',cwd=os.getcwd())).encode()).decode()+'|END]')"
-            command = f"export PROBE_SESSION={self.session} PROBE_SHELL_PID=$$; '{python}' -c '{code.replace(chr(39), chr(39)+chr(34)+chr(39)+chr(34)+chr(39))}'"
-            # Encode the Python payload to keep the complete framing out of input echo.
-            payload = base64.b64encode(code.encode()).decode()
-            command = f"export PROBE_SESSION={self.session} PROBE_SHELL_PID=$$; '{python}' -c \"import base64;exec(base64.b64decode('{payload}'))\""
-        else:
-            script = "$global:ProbeSession='" + self.session + "'; $r=@{session=$ProbeSession;pid=$PID;shell=$PSVersionTable.PSVersion.ToString();cwd=(Get-Location).Path}; [Console]::WriteLine('[PROBE|'+[Convert]::ToBase64String([Text.Encoding]::UTF8.GetBytes(($r|ConvertTo-Json -Compress)))+'|END]')"
-            payload = base64.b64encode(script.encode("utf-8")).decode()
-            command = ". ([scriptblock]::Create([Text.Encoding]::UTF8.GetString([Convert]::FromBase64String('" + payload + "'))))"
-        write_json(self.directory / "readiness-command.json", {"command": command, "session": self.session})
-        self.type(command)
-        deadline = time.monotonic() + 10
-        while time.monotonic() < deadline:
-            self.cdp.pump()
-            matches = [r for r in self.parser.records if r.get("session") == self.session]
-            if matches:
-                self.screenshot("ready.png")
-                return matches[-1]
-            if self.closed:
-                raise ConnectionError("Terminal connection lost during readiness")
-        self.screenshot("readiness-timeout.png")
-        raise TimeoutError("No nonce/shell/cwd completion record from the filmed terminal")
-
-    def wait_record(self, predicate, timeout=10):
-        deadline = time.monotonic() + timeout
-        while time.monotonic() < deadline:
-            matches = [record for record in self.parser.records if record.get("session") == self.session and predicate(record)]
-            if matches:
-                return matches[-1]
-            if self.closed:
-                raise ConnectionError("Terminal connection lost while waiting for completion")
-            self.cdp.pump()
-        raise TimeoutError("Missing terminal completion record")
-
-    def install_prompt(self):
-        if self.args.shell_kind == "gitbash":
-            python = str(Path(sys.executable)).replace("\\", "/")
-            source = str(Path(__file__).resolve()).replace("\\", "/")
-            script = f'''export PROBE_SESSION={self.session}
-PROBE_PHASE=boot
-probe_prompt() {{
-    local probe_status=$? probe_pipeline=("${{PIPESTATUS[@]}}")
-    if [[ -n "$PROBE_PHASE" ]]; then
-        PROBE_STATUS="$probe_status" PROBE_PIPELINE="${{probe_pipeline[*]}}" PROBE_SHELL_PID=$$ PROBE_CWD="$PWD" PROBE_VALUE="$PROBE_VALUE" PROBE_PHASE="$PROBE_PHASE" PROBE_REQUEST="$PROBE_REQUEST" PROBE_NATIVE="$PROBE_NATIVE" '{python}' '{source}' --emit-bash
-        if [[ "$PROBE_PHASE" == arm ]]; then PROBE_PHASE=run; else PROBE_PHASE=; fi
-    fi
-    return "$probe_status"
-}}
-PROMPT_COMMAND=probe_prompt
-PS1='PROBE $ '
-'''
-            payload = base64.b64encode(script.encode()).decode()
-            self.type(f'''eval "$('{python}' -c "import base64;print(base64.b64decode('{payload}').decode())")"''')
-        else:
-            script = r'''
-$global:ProbeErrorCount=$Error.Count
-$global:ProbePhase='boot'
-function global:prompt {
-    $probeOK=$?
-    $probeNative=$global:LASTEXITCODE
-    if ($global:ProbePhase) {
-        $probeError=$null
-        $parseError=$false
-        if ($Error.Count -gt 0 -and $Error.Count -gt $global:ProbeErrorCount) {
-            $probeError=$Error[0].ToString()
-            $parseError=($Error[0] -is [System.Management.Automation.ParseException] -or $Error[0].Exception -is [System.Management.Automation.ParseException] -or $Error[0].CategoryInfo.Category -eq 'ParserError')
-        }
-        $record=@{session=$global:ProbeSession;request=$global:ProbeRequest;phase=$global:ProbePhase;pid=$PID;cwd=(Get-Location).Path;variable=$global:ProbeValue;outcome='completed';shell_success=($probeOK -and -not $parseError);raw_shell_success=$probeOK;parse_error=$parseError;native_exit_code=$null;shell_error=$probeError;producer_exit_code=$null;producer_success=$null}
-        if ($global:ProbeNative -and -not $parseError) {$record.native_exit_code=$probeNative;$record.producer_exit_code=$probeNative;$record.producer_success=($probeNative -eq 0)}
-        if ($global:ProbePhase -eq 'arm') {$global:ProbePhase='run'} else {$global:ProbePhase=$null}
-        $encoded=[Convert]::ToBase64String([Text.Encoding]::UTF8.GetBytes(($record|ConvertTo-Json -Compress)))
-        [Console]::WriteLine('[PROBE|')
-        for ($offset=0;$offset -lt $encoded.Length;$offset+=60) {[Console]::WriteLine($encoded.Substring($offset,[Math]::Min(60,$encoded.Length-$offset)))}
-        [Console]::WriteLine('|END]')
-    }
-    return 'PROBE PS> '
-}
-'''
-            payload = base64.b64encode(script.encode()).decode()
-            self.type(". ([scriptblock]::Create([Text.Encoding]::UTF8.GetString([Convert]::FromBase64String('" + payload + "'))))")
-        try:
-            return self.wait_record(lambda r: r.get("phase") == "boot")
-        except TimeoutError:
-            if self.args.shell_kind != "gitbash":
-                self.type("$Error | Select-Object -First 3 | Format-List * -Force")
-                deadline = time.monotonic() + 2
-                while time.monotonic() < deadline:
-                    self.cdp.pump()
-                self.screenshot("prompt-error.png")
-            raise
-
-    def arm(self, request, native):
-        if self.args.shell_kind == "gitbash":
-            command = f"PROBE_REQUEST={request}; PROBE_NATIVE={int(bool(native))}; PROBE_PHASE=arm"
-        else:
-            command = f"$global:ProbeRequest={request}; $global:ProbeNative=${str(bool(native)).lower()}; $global:ProbeErrorCount=$Error.Count; $global:ProbePhase='arm'"
-        self.type(command)
-        self.wait_record(lambda r: r.get("phase") == "arm" and str(r.get("request")) == str(request))
-
-    def begin_take(self, name):
-        if self.take is not None:
-            raise RuntimeError("A take is already active")
-        if not re.fullmatch(r"[a-zA-Z0-9_-]+", name):
-            raise ValueError("Invalid take name")
-        (self.directory / name).mkdir(exist_ok=False)
-        self.take, self.take_frames = name, []
-        self.cdp.call("Page.startScreencast", {"format": "jpeg", "quality": 80, "everyNthFrame": 1})
-
-    def end_take(self):
-        if self.take is None:
-            raise RuntimeError("No take is active")
-        self.cdp.call("Page.stopScreencast")
-        self.screenshot(self.take + "-end.png")
-        frames = self.take_frames
-        write_json(self.directory / self.take / "frames.json", frames)
-        self.take = None
-        return frames
-
-    def key(self, key):
-        if key == "CTRL_C":
-            params = {"key": "c", "code": "KeyC", "windowsVirtualKeyCode": 67, "modifiers": 2}
-        elif key == "ENTER":
-            params = {"key": "Enter", "code": "Enter", "windowsVirtualKeyCode": 13, "text": "\r"}
-        else:
-            params = {"key": key, "text": key, "windowsVirtualKeyCode": ord(key.upper())}
-        self.cdp.call("Input.dispatchKeyEvent", {"type": "keyDown", **params})
-        self.cdp.call("Input.dispatchKeyEvent", {"type": "keyUp", **{k: v for k, v in params.items() if k != "text"}})
-
-    def close(self):
-        start = time.monotonic()
-        try:
-            self.job.close()
-        finally:
-            if self.cdp is not None:
-                self.cdp.ws.close()
-                self.cdp.trace.close()
-            self.output.close()
-            self.raw.close()
-        return time.monotonic() - start
-
-
-REQUIRED = ("readiness", "framing", "command_outcomes", "two_takes", "extra_client",
-            "browser_loss", "normal_cleanup", "cancel_cleanup", "crash_cleanup")
-
-
-def emit_bash():
-    env = os.environ
-    status = int(env["PROBE_STATUS"])
-    native = env.get("PROBE_NATIVE") == "1"
-    pipeline = [int(value) for value in env["PROBE_PIPELINE"].split()]
-    record = dict(session=env["PROBE_SESSION"], request=env.get("PROBE_REQUEST"),
-                  phase=env["PROBE_PHASE"], pid=env["PROBE_SHELL_PID"],
-                  pid_kind="msys", cwd=env["PROBE_CWD"], variable=env.get("PROBE_VALUE"),
-                  outcome="completed", shell_success=status == 0,
-                  native_exit_code=pipeline[0] if native else None, shell_error=None,
-                  pipeline_statuses=pipeline, producer_exit_code=pipeline[0] if native else None,
-                  producer_success=pipeline[0] == 0 if native else None)
-    encoded = base64.b64encode(json.dumps(record).encode()).decode()
-    print("[PROBE|\n" + "\n".join(encoded[offset:offset + 60] for offset in range(0, len(encoded), 60)) + "\n|END]", flush=True)
-
-
-def native_command(args, code):
-    # Each beat explicitly identifies the native producer; no stale LASTEXITCODE attribution.
-    python = str(Path(sys.executable))
-    prefix = "& " if args.shell_kind != "gitbash" else ""
-    if args.shell_kind == "gitbash":
-        python = python.replace("\\", "/")
-    return f'''{prefix}'{python}' -c "{code}"'''
-
-
-def outcome_commands(args):
-    native = str(Path(sys.executable))
-    commands = [("native_success", native_command(args, "import sys;sys.exit(0)"), native)]
-    if args.shell_kind != "gitbash":
-        commands += [
-            ("cmdlet_failure", "Get-Item 'Z:\\probe-path-that-does-not-exist'", None),
-            ("native_failure", native_command(args, "import sys;sys.exit(7)"), native),
-            ("cmdlet_success", "Write-Output 'cmdlet success λ'", None),
-            ("terminating_error", "throw 'probe terminating error'", None),
-            ("nonterminating_error", "Write-Error 'probe nonterminating error'", None),
-            ("parse_failure", "Write-Output )", None),
-            ("logging_failure", native_command(args, "import sys;print('producer');sys.exit(7)") + " | Tee-Object -Variable ProbeLog", native),
-            ("expression_wrapper", "(Write-Error 'probe expression wrapper')", None),
-            ("script_exit", f"& '{args.shell}' -NoLogo -NoProfile -Command 'exit 9'", str(args.shell)),
-        ]
-    else:
-        commands += [
-            ("shell_failure", "test -e /probe-path-that-does-not-exist", None),
-            ("native_failure", native_command(args, "import sys;sys.exit(7)"), native),
-            ("shell_success", "printf 'shell success λ\\n'", None),
-            ("parse_failure", "echo )", None),
-            ("logging_failure", native_command(args, "import sys;print('producer');sys.exit(7)") + " | cat", native),
-            ("script_exit", "bash --noprofile --norc -c 'exit 9'", "Git Bash"),
-        ]
-    return commands
-
-
-def run_outcomes(terminal):
-    records = {}
-    terminal.install_prompt()
-    terminal.begin_take("outcomes")
-    for request, (name, command, native) in enumerate(outcome_commands(terminal.args), 1):
-        terminal.arm(request, native)
-        terminal.type(command)
-        record = terminal.wait_record(lambda r: r.get("phase") == "run" and str(r.get("request")) == str(request))
-        records[name] = {**record, "command": command, "native_producer": native}
-        write_json(terminal.directory / "outcomes.json", records)
-        print(json.dumps({"case": name, "record": records[name]}, ensure_ascii=False), flush=True)
-    terminal.end_take()
-    if terminal.args.shell_kind != "gitbash":
-        assert_outcomes(records)
-        assert records["expression_wrapper"]["shell_success"] is (terminal.args.shell_kind == "powershell51")
-    else:
-        assert records["native_success"]["native_exit_code"] == 0
-        assert records["native_failure"]["native_exit_code"] == 7
-        assert records["shell_failure"]["shell_success"] is False
-        assert records["shell_success"]["shell_success"] is True
-        assert records["parse_failure"]["shell_success"] is False
-        assert records["logging_failure"]["pipeline_statuses"] == [7, 0]
-    assert records["script_exit"]["native_exit_code"] == 9
-    return records
-
-
-def heartbeat_worker(directory, role):
-    directory.mkdir(parents=True, exist_ok=True)
-    if role in ("parent", "child"):
-        child_role = "child" if role == "parent" else "grandchild"
-        subprocess.Popen([sys.executable, str(Path(__file__).resolve()), "--worker", child_role,
-                          "--directory", str(directory)], stdin=subprocess.DEVNULL)
-    print(f"HEARTBEAT {role} {os.getpid()}", flush=True)
-    while True:
-        write_json(directory / f"{role}.json", {"pid": os.getpid(), "timestamp": time.time(), "role": role})
-        time.sleep(0.1)
-
-
-def heartbeats(directory):
-    return {path.stem: read_json(path) for path in directory.glob("*.json")}
-
-
-class Supervisor:
-    def __init__(self, args):
-        self.args = args
-        self.terminal = Terminal(args, args.directory)
-        self.control = args.directory / "control"
-        self.control.mkdir()
-        self.pending, self.last_request = None, 0
-        self.results, self.takes = {}, []
-        self.sentinel = None
-        self.report = {"gate": "incomplete", "shell_kind": args.shell_kind,
-                       "environment": {"launch_host": args.launch_host, "recording_host": platform.node(),
-                                       "platform": platform.platform(), "python": sys.executable,
-                                       "python_version": sys.version, "pointer_bits": struct.calcsize("P") * 8,
-                                       "token": token_info()},
-                       "session_id": self.terminal.session, "supervisor_pid": os.getpid(),
-                       "artifacts": str(args.directory), "requests": [], "results": self.results,
-                       "takes": self.takes, "checks": {name: {"status": "unavailable", "reason": "Not exercised in this session"} for name in REQUIRED}}
-
-    def reply(self, request, value, suffix="result"):
-        write_json(self.control / f"{request:06d}.{suffix}.json", value)
-
-    def status(self):
-        terminal = self.terminal
-        processes = terminal.job.snapshot()
-        try:
-            owned = [{key: value for key, value in item.items() if key != "handle"} for item in processes]
-        finally:
-            for item in processes:
-                terminal.job.k.CloseHandle(item["handle"])
-        supervisor_handle = terminal.job.open_process(os.getpid())
-        try:
-            creation = terminal.job.process_time(supervisor_handle)
-        finally:
-            terminal.job.k.CloseHandle(supervisor_handle)
-        sentinel_handle = terminal.job.open_process(self.sentinel.pid)
-        try:
-            sentinel = {"pid": self.sentinel.pid, "creation": terminal.job.process_time(sentinel_handle)}
-        finally:
-            terminal.job.k.CloseHandle(sentinel_handle)
-        status = {"session": terminal.session, "control": str(self.control), "supervisor_pid": os.getpid(),
-                  "supervisor_creation": creation, "owned": owned, "sentinel": sentinel,
-                  "token": self.report["environment"]["token"],
-                  "pending": self.pending, "last_request": self.last_request,
-                  "terminal_sizes": terminal.terminal_sizes, "take": terminal.take,
-                  "heartbeats": heartbeats(self.args.directory / "heartbeats")}
-        write_json(self.control / "status.json", status)
-        return status
-
-    def finish_pending(self):
-        if self.pending is None:
-            return
-        request = self.pending
-        records = [r for r in self.terminal.parser.records if r.get("session") == self.terminal.session
-                   and r.get("phase") == "run" and str(r.get("request")) == str(request["id"])]
-        if records:
-            result = {**records[-1], "command": request["command"], "native_producer": request.get("native_producer"),
-                      "finished_at": time.time(), "page_id": self.terminal.page_id,
-                      "websocket_id": self.terminal.request_id}
-            self.results[str(request["id"])] = result
-            self.reply(request["id"], result)
-            self.pending = None
-        elif time.time() > request["deadline"]:
-            self.reply(request["id"], {"outcome": "unknown", "reason": "Completion deadline expired"})
-            raise TimeoutError("An in-flight command has no completion record")
-
-    def dispatch(self, request):
-        terminal = self.terminal
-        operation, number = request["operation"], request["id"]
-        self.report["requests"].append({**request, "dispatched_at": time.time(), "active_take": terminal.take})
-        if operation == "run":
-            if self.pending is not None:
-                raise RuntimeError("Another command still owns terminal input")
-            terminal.arm(number, request.get("native_producer"))
-            self.pending = {**request, "deadline": time.time() + request.get("timeout", 90)}
-            terminal.type(request["command"])
-            self.reply(number, {"accepted": True, "in_flight": True, "time": time.time()}, "ack")
-            return
-        self.reply(number, {"accepted": True, "time": time.time()}, "ack")
-        value = {"operation": operation, "time": time.time()}
-        if operation == "begin-take":
-            terminal.begin_take(request["name"])
-            self.takes.append({"name": request["name"], "begin": time.time(), "begin_request": number,
-                               "pending_at_begin": self.pending, "session": terminal.session,
-                               "page_id": terminal.page_id, "websocket_id": terminal.request_id})
-        elif operation == "end-take":
-            frames = terminal.end_take()
-            self.takes[-1].update(end=time.time(), end_request=number, frames=len(frames), pending_at_end=self.pending)
-            value["frames"] = len(frames)
-        elif operation == "key":
-            terminal.key(request["key"])
-        elif operation == "resize":
-            terminal.cdp.call("Emulation.setDeviceMetricsOverride", {"width": request["width"], "height": request["height"], "deviceScaleFactor": 1, "mobile": False})
-            deadline = time.monotonic() + 2
-            while time.monotonic() < deadline:
-                terminal.cdp.pump()
-            value["terminal_sizes"] = terminal.terminal_sizes
-            assert terminal.terminal_sizes[-1]["columns"] >= 80, "Probe framing is only tested at >=80 columns"
-        elif operation == "inspect":
-            value.update(self.status())
-        elif operation == "snapshot":
-            name = request["name"]
-            if not re.fullmatch(r"[a-zA-Z0-9_-]+", name):
-                raise ValueError("Invalid snapshot name")
-            assert request["expected_text"] in terminal.parser.text
-            terminal.screenshot(name + ".png")
-            value.update(screenshot=name + ".png", pending=self.pending,
-                         observed_text=request["expected_text"])
-        elif operation == "extra-client":
-            import websocket
-
-            candidate = None
-            try:
-                candidate = websocket.create_connection(terminal.terminal_url.replace("http:", "ws:") + "ws",
-                    subprotocols=["tty"], timeout=2, origin=terminal.terminal_url.rstrip("/"))
-                value.update(accepted=True, candidate_continuity=False, reason="Unexpected second WebSocket accepted; no continuity claimed")
-            except (websocket.WebSocketBadStatusException, websocket.WebSocketConnectionClosedException) as error:
-                value.update(accepted=False, candidate_continuity=False, http_status=getattr(error, "status_code", None), error=str(error))
-            finally:
-                if candidate:
-                    candidate.close()
-            self.report["extra_client"] = value
-            if value.get("accepted") is False:
-                terminal.arm(number, None)
-                terminal.type("printf 'original still alive\\n'" if self.args.shell_kind == "gitbash" else "Write-Output 'original still alive'")
-                value["original_round_trip"] = terminal.wait_record(lambda r: r.get("phase") == "run" and str(r.get("request")) == str(number))
-            self.report["checks"]["extra_client"] = {"status": "passed" if value.get("accepted") is False else "failed"}
-        elif operation == "browser-loss":
-            terminal.screenshot("before-browser-loss.png")
-            browser = next(entry for entry in terminal.launches if entry["name"] == "browser")
-            handle = next(handle for pid, handle in terminal.job.roots if pid == browser["pid"])
-            if not terminal.job.k.TerminateProcess(handle, 1):
-                raise ctypes.WinError(ctypes.get_last_error())
-            self.report["browser_loss_requested"] = True
-        elif operation in ("close", "cancel"):
-            self.report["close_operation"] = operation
-        else:
-            raise ValueError(f"Unknown operation: {operation}")
-        self.reply(number, value)
-
-    def run(self):
-        terminal = self.terminal
-        failure = None
-        try:
-            terminal.start()
-            before = terminal.readiness()
-            terminal.install_prompt()
-            self.report["readiness"] = before
-            self.report["checks"]["readiness"] = {"status": "passed", "screenshot": str(self.args.directory / "ready.png")}
-            self.report["checks"]["framing"] = {"status": "passed" if any(f["leading_byte"] == "30" for f in terminal.frames) else "failed"}
-            self.sentinel = subprocess.Popen([sys._base_executable, str(Path(__file__).resolve()), "--worker", "sentinel",
-                                             "--directory", str(self.args.directory / "sentinel")],
-                                            stdin=subprocess.DEVNULL, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
-            ready = self.status()
-            write_json(self.control / "ready.json", ready)
-            print(json.dumps({"ready": ready, "directory": str(self.args.directory)}), flush=True)
-            deadline = time.monotonic() + 1200
-            while not self.report.get("close_operation"):
-                if time.monotonic() > deadline:
-                    raise TimeoutError("Supervisor maximum lifetime expired")
-                terminal.cdp.pump()
-                if terminal.closed:
-                    raise ConnectionError("Recorded terminal disconnected")
-                self.finish_pending()
-                path = self.control / f"{self.last_request + 1:06d}.request.json"
-                if path.exists():
-                    request = json.loads(path.read_text(encoding="utf-8"))
-                    if request["id"] != self.last_request + 1:
-                        raise ValueError("Request id does not match its ordered filename")
-                    self.last_request += 1
-                    self.dispatch(request)
-                    self.status()
-        except BaseException as error:
-            failure = repr(error)
-            self.report["failure"] = {"error": failure, "traceback": traceback.format_exc()}
-            if self.report["checks"]["readiness"]["status"] == "unavailable":
-                self.report["checks"]["readiness"] = {"status": "failed", "error": failure}
-            if self.report.get("browser_loss_requested"):
-                self.report["checks"]["browser_loss"] = {"status": "passed", "continuity": False, "outcome": "interrupted", "observed_error": failure}
-        finally:
-            cleanup_ok = self.cleanup(failure)
-        return 1 if not cleanup_ok or (failure and not self.report.get("browser_loss_requested")) else 0
-
-    def cleanup(self, failure):
-        terminal, job = self.terminal, self.terminal.job
-        started = time.monotonic()
-        mode = "cancel_cleanup" if self.report.get("close_operation") == "cancel" else "normal_cleanup"
-        self.report["checks"][mode] = {"status": "failed"}
-        diagnostics = self.report["cleanup_diagnostics"] = []
-        handles = before = after = stable = None
-        remaining, sentinel_alive = None, False
-
-        def attempt(stage, operation):
-            try:
-                return operation()
-            except BaseException as error:
-                diagnostic = {"stage": stage, "error": repr(error), "traceback": traceback.format_exc()}
-                diagnostics.append(diagnostic)
-                # Preserve diagnostics even when the evidence directory is unwritable.
-                try:
-                    print(json.dumps({"cleanup_error": diagnostic}), file=sys.stderr, flush=True)
-                except (OSError, ValueError):
-                    pass
-                return None
-
-        try:
-            handles = attempt("snapshot", job.snapshot)
-            before = attempt("heartbeats_before", lambda: heartbeats(self.args.directory / "heartbeats"))
-            if self.pending is not None:
-                result = {"outcome": "interrupted" if failure or self.report.get("close_operation") else "unknown",
-                          "shell_success": None, "native_exit_code": None, "shell_error": failure,
-                          "command": self.pending["command"]}
-                self.results[str(self.pending["id"])] = result
-                attempt("pending_reply", lambda: self.reply(self.pending["id"], result))
-
-            def graceful_shutdown():
-                if terminal.cdp and not failure:
-                    if terminal.take:
-                        terminal.end_take()
-                    terminal.key("CTRL_C")
-                    deadline = time.monotonic() + 1.5
-                    while time.monotonic() < deadline:
-                        terminal.cdp.pump()
-                    terminal.type("exit")
-                    deadline = time.monotonic() + 1.5
-                    while time.monotonic() < deadline and not terminal.closed:
-                        terminal.cdp.pump()
-
-            attempt("graceful_shutdown", graceful_shutdown)
-        finally:
-            try:
-                try:
-                    attempt("terminal_close", terminal.close)
-                finally:
-                    # Terminal.close may fail before reaching the job. Job.close
-                    # releases its kill-on-close handle even if termination fails.
-                    if job.handle:
-                        attempt("job_close", job.close)
-                    for name, resource in (("terminal_output", terminal.output), ("network_log", terminal.raw)):
-                        attempt(name + "_close", resource.close)
-                    if terminal.cdp:
-                        attempt("cdp_close", terminal.cdp.ws.close)
-                        attempt("cdp_trace_close", terminal.cdp.trace.close)
-                if handles is not None:
-                    def wait_handles():
-                        self.report["initial_unsignaled_handles"] = [p["pid"] for p in handles if job.k.WaitForSingleObject(p["handle"], 0) != 0]
-                        deadline = started + 8
-                        while time.monotonic() < deadline and any(job.k.WaitForSingleObject(p["handle"], 0) != 0 for p in handles):
-                            time.sleep(0.05)
-                        return [p["pid"] for p in handles if job.k.WaitForSingleObject(p["handle"], 0) != 0]
-
-                    remaining = attempt("owned_handle_wait", wait_handles)
-                sentinel_alive = attempt("sentinel_status", lambda: self.sentinel is not None and self.sentinel.poll() is None)
-                after = attempt("heartbeats_after", lambda: heartbeats(self.args.directory / "heartbeats"))
-                time.sleep(0.4)
-                stable = attempt("heartbeats_verify", lambda: heartbeats(self.args.directory / "heartbeats"))
-            finally:
-                try:
-                    for item in handles or []:
-                        def release_handle(item=item):
-                            if not job.k.CloseHandle(item["handle"]):
-                                raise ctypes.WinError(ctypes.get_last_error())
-                        attempt("owned_handle_close", release_handle)
-                finally:
-                    # Sentinel cleanup is independent of every job/terminal step.
-                    try:
-                        if self.sentinel:
-                            attempt("sentinel_terminate", self.sentinel.terminate)
-                    finally:
-                        if self.sentinel:
-                            attempt("sentinel_wait", lambda: self.sentinel.wait(timeout=1))
-
-        stopped = after is not None and stable is not None and after == stable
-        self.report.update(owned_children_remaining=remaining, unrelated_sentinel_alive=sentinel_alive,
-                           shutdown_seconds=time.monotonic() - started)
-        self.report["cleanup"] = {"before": before, "after": after, "heartbeats_stopped": stopped,
-                                  "owned_handles": None if handles is None else
-                                  [{k: v for k, v in item.items() if k != "handle"} for item in handles]}
-        passed = (not diagnostics and remaining == [] and sentinel_alive is True and stopped
-                  and before is not None and set(before) == {"parent", "child", "grandchild"}
-                  and self.report["shutdown_seconds"] < 10)
-        self.report["checks"][mode] = {"status": "passed" if passed else "failed"}
-        self.report.update(terminal_client_count=len(terminal.socket_ids), observed_session_id=terminal.session,
-                           filmed_session_id=self.takes[0]["session"] if self.takes else None,
-                           socket_ids=terminal.socket_ids, observed_frames=terminal.frames,
-                           terminal_sizes=terminal.terminal_sizes, framing_chunk_columns=60,
-                           framing_scope="Observed dimensions only; arbitrary resizing/redraw is unverified",
-                           job_structure_sizes=job.sizes, launches=terminal.launches)
-        attempt("report_write", lambda: write_json(self.args.directory / "probe.json", self.report))
-        if diagnostics:
-            self.report["checks"][mode] = {"status": "failed"}
-        prior_errors = len(diagnostics)
-        attempt("summary_write", lambda: print(json.dumps({"finished": str(self.args.directory),
-                "checks": self.report["checks"], "failure": failure, "cleanup_diagnostics": diagnostics}), flush=True))
-        if len(diagnostics) != prior_errors:
-            self.report["checks"][mode] = {"status": "failed"}
-            attempt("report_write", lambda: write_json(self.args.directory / "probe.json", self.report))
-        return passed and not diagnostics
-
-
-def submit_request(args):
-    request = json.loads(args.submit)
-    number = request["id"]
-    if not isinstance(number, int) or number < 1:
-        raise ValueError("Request ids must be positive integers")
-    path = args.directory / "control" / f"{number:06d}.request.json"
-    if path.exists():
-        raise FileExistsError("Request id already submitted; refusing replay")
-    write_json(path, request)
-    deadline = time.monotonic() + args.wait
-    reply = path.with_name(f"{number:06d}.{'result' if args.wait_result else 'ack'}.json")
-    while not reply.exists() and time.monotonic() < deadline:
-        time.sleep(0.05)
-    if not reply.exists():
-        raise TimeoutError(f"No acknowledgment/result for request {number}")
-    print(reply.read_text(encoding="utf-8"), flush=True)
-
-
-def wait_json(path, timeout=60):
-    deadline = time.monotonic() + timeout
-    while time.monotonic() < deadline:
-        if path.exists():
-            return read_json(path)
-        time.sleep(0.05)
-    raise TimeoutError(f"Missing artifact: {path}")
-
-
-def drive_phase(args):
-    """Run a bounded group of requests; invoke each phase in a separate harness call."""
-    control = args.directory / "control"
-    state = wait_json(control / "ready.json")
-    submitted = list(control.glob("*.request.json"))
-    number = max((int(path.name.split(".")[0]) for path in submitted), default=0)
-
-    def send(operation, wait=True, **fields):
-        nonlocal number
-        number += 1
-        request = {"id": number, "operation": operation, **fields}
-        submit_request(argparse.Namespace(directory=args.directory, submit=json.dumps(request), wait=10, wait_result=wait))
-        return number
-
-    def native_fixture(role):
-        prefix = "& " if args.shell_kind != "gitbash" else ""
-        paths = [str(Path(sys.executable)), str(Path(__file__).resolve()), str(args.directory / "heartbeats")]
-        if args.shell_kind == "gitbash":
-            paths = [path.replace("\\", "/") for path in paths]
-        return f"{prefix}'{paths[0]}' '{paths[1]}' --worker {role} --directory '{paths[2]}'"
-
-    if args.phase == "first":
-        send("begin-take", name="take-one")
-        command = "PROBE_VALUE='persisted-λ'" if args.shell_kind == "gitbash" else "$global:ProbeValue='persisted-λ'"
-        send("run", command=command)
-        delayed = send("run", wait=False, command=native_command(args, "import time;print('LONG_STARTED',flush=True);time.sleep(45);print('LONG_FINISHED',flush=True)"), native_producer=sys.executable)
-        send("end-take")
-        write_json(control / "first-phase.json", {"delayed_request": delayed, "returned_at": time.time(), "session": state["session"]})
-    elif args.phase == "second":
-        first = wait_json(control / "first-phase.json")
-        send("begin-take", name="take-two")
-        inspect_id = send("inspect")
-        inspection = wait_json(control / f"{inspect_id:06d}.result.json")
-        assert inspection["pending"] and inspection["pending"]["id"] == first["delayed_request"], "Delayed command completed before the second take began"
-        send("resize", width=1200, height=800)
-        delayed = wait_json(control / f"{first['delayed_request']:06d}.result.json", 70)
-        assert delayed["outcome"] == "completed" and delayed["shell_success"] is True
-        command = "printf 'after λ\\n'" if args.shell_kind == "gitbash" else "Write-Output 'after λ'"
-        after_id = send("run", command=command)
-        after = wait_json(control / f"{after_id:06d}.result.json")
-        assert after["variable"] == "persisted-λ"
-        send("end-take")
-        write_json(control / "second-phase.json", {"after": after, "delayed": delayed, "inspection": inspection,
-                                                  "returned_at": time.time()})
-    if args.phase in ("second", "interactive"):
-        send("begin-take", name="interactive")
-        prefix = "& " if args.shell_kind != "gitbash" else ""
-        python, source = str(Path(sys.executable)), str(Path(__file__).resolve())
-        if args.shell_kind == "gitbash":
-            python, source = python.replace("\\", "/"), source.replace("\\", "/")
-        tui = send("run", wait=False, command=f"{prefix}'{python}' '{source}' --tui", native_producer=sys.executable)
-        # Application readiness is a native fixture artifact, not echoed command text.
-        wait_json(args.directory / "tui-ready.json")
-        # Give the recorded application a visible hold; the supervisor continues
-        # draining CDP/screencast events while this separate driver waits.
-        time.sleep(1)
-        send("snapshot", name="tui-live", expected_text="Native Windows TUI λ")
-        send("key", key="q")
-        outcome = wait_json(control / f"{tui:06d}.result.json")
-        assert outcome["shell_success"] is True
-        send("end-take")
-        send("extra-client")
-        write_json(control / "interactive-phase.json", {"interactive": outcome, "returned_at": time.time()})
-    elif args.phase == "prepare-cleanup":
-        send("begin-take", name="ownership")
-        send("run", wait=False, command=native_fixture("parent"), native_producer=sys.executable, timeout=600)
-        deadline = time.monotonic() + 10
-        while time.monotonic() < deadline:
-            found = heartbeats(args.directory / "heartbeats")
-            if set(found) == {"parent", "child", "grandchild"}:
-                break
-            time.sleep(0.1)
-        assert set(found) == {"parent", "child", "grandchild"}, found
-        send("inspect")
-    elif args.phase in ("close", "cancel", "browser-loss"):
-        send(args.phase)
-        result = wait_json(args.directory / "probe.json", 15)
-        print(json.dumps({"checks": result["checks"], "cleanup": result.get("cleanup")}), flush=True)
-
-
-def tui_fixture():
-    import msvcrt
-
-    print("\x1b[2J\x1b[HNative Windows TUI λ\n[q] Finish this screen", flush=True)
-    write_json(Path.cwd() / "tui-ready.json", {"pid": os.getpid(), "ready_at": time.time()})
-    while True:
-        key = msvcrt.getwch()
-        if key == "q":
-            print("\nTUI_EXIT:q", flush=True)
-            return
-
-
-def crash_check(args):
-    """An independent harness call retains native handles before killing the supervisor."""
-    state = wait_json(args.directory / "control" / "status.json")
-    observer = WindowsJob()  # Empty observer job; no process is assigned to it.
-    handles, sentinel, supervisor = [], None, None
-    report = {"status": "failed", "state_before": state, "launch_host": args.launch_host,
-              "recording_host": platform.node(), "prior_children_already_exited": []}
-    try:
-        supervisor = observer.open_process(state["supervisor_pid"], state["supervisor_creation"], terminate=True)
-        sentinel = observer.open_process(state["sentinel"]["pid"], state["sentinel"]["creation"], terminate=True)
-        for process in state["owned"]:
-            try:
-                handle = observer.open_process(process["pid"], process["creation"])
-            except OSError as error:
-                if error.winerror == 87:
-                    report["prior_children_already_exited"].append(process)
-                    continue
-                raise
-            handles.append({**process, "handle": handle})
-        before = heartbeats(args.directory / "heartbeats")
-        assert set(before) == {"parent", "child", "grandchild"}, before
-        assert {record["pid"] for record in before.values()} <= {record["pid"] for record in handles}
-        started = time.monotonic()
-        if not observer.k.TerminateProcess(supervisor, 99):
-            raise ctypes.WinError(ctypes.get_last_error())
-        deadline = time.monotonic() + 5
-        while time.monotonic() < deadline and any(observer.k.WaitForSingleObject(p["handle"], 0) != 0 for p in handles):
-            time.sleep(0.05)
-        remaining = [p["pid"] for p in handles if observer.k.WaitForSingleObject(p["handle"], 0) != 0]
-        report["shutdown_seconds"] = time.monotonic() - started
-        report["owned_children_remaining"] = remaining
-        report["unrelated_sentinel_alive"] = observer.k.WaitForSingleObject(sentinel, 0) == 258
-        after = heartbeats(args.directory / "heartbeats")
-        sentinel_before = heartbeats(args.directory / "sentinel")
-        time.sleep(0.5)
-        report.update(heartbeats_before=before, heartbeats_after=after,
-                      heartbeats_stopped=after == heartbeats(args.directory / "heartbeats"),
-                      sentinel_heartbeat_advanced=heartbeats(args.directory / "sentinel") != sentinel_before)
-        assert remaining == []
-        assert report["heartbeats_stopped"] and report["unrelated_sentinel_alive"] and report["sentinel_heartbeat_advanced"]
-        assert report["shutdown_seconds"] < 10
-        report["status"] = "passed"
-    except BaseException as error:
-        report.update(error=repr(error), traceback=traceback.format_exc())
-        raise
-    finally:
-        write_json(args.directory / "crash.json", report)
-        if sentinel:
-            observer.k.TerminateProcess(sentinel, 0)
-            observer.k.WaitForSingleObject(sentinel, 5000)
-            observer.k.CloseHandle(sentinel)
-        if supervisor:
-            observer.k.CloseHandle(supervisor)
-        for process in handles:
-            observer.k.CloseHandle(process["handle"])
-        observer.close()
-    print(json.dumps(report), flush=True)
-
-
-def startup_probe(args):
-    result = {"gate": "incomplete", "environment": {"launch_host": args.launch_host,
-              "recording_host": platform.node(), "platform": platform.platform(),
-              "python": sys.executable, "python_version": sys.version,
-              "pointer_bits": struct.calcsize("P") * 8, "token": token_info()},
-              "shell_kind": args.shell_kind, "shell_executable": str(args.shell),
-              "checks": {name: {"status": "unavailable", "reason": "Not reached"} for name in REQUIRED},
-              "artifacts": str(args.directory)}
-    terminal = None
-    try:
-        terminal = Terminal(args, args.directory)
-        terminal.start()
-        result["readiness"] = terminal.readiness()
-        result["checks"]["readiness"] = {"status": "passed", "screenshot": str(args.directory / "ready.png")}
-        assert any(frame["leading_byte"] == "30" for frame in terminal.frames)
-        result["checks"]["framing"] = {"status": "passed"}
-        if args.outcomes:
-            result["outcomes"] = run_outcomes(terminal)
-            result["checks"]["command_outcomes"] = {"status": "passed"}
-    except Exception as error:
-        check = "command_outcomes" if args.outcomes and result["checks"]["readiness"]["status"] == "passed" else "readiness"
-        result["checks"][check] = {"status": "failed", "error": repr(error), "traceback": traceback.format_exc()}
-    finally:
-        if terminal:
-            result["socket_ids"] = terminal.socket_ids
-            result["observed_frames"] = terminal.frames
-            result["terminal_sizes"] = terminal.terminal_sizes
-            result["framing_chunk_columns"] = 60
-            result["framing_scope"] = "Bounded multiline records at the observed dimensions; arbitrary resize/redraw decoding is unverified"
-            result["job_structure_sizes"] = terminal.job.sizes
-            result["launches"] = terminal.launches
-            result["terminal_text"] = terminal.parser.text
-            try:
-                result["forced_cleanup_seconds"] = terminal.close()
-            except Exception as error:
-                result["cleanup_error"] = repr(error)
-        args.directory.mkdir(parents=True, exist_ok=True)
-        write_json(args.directory / "probe.json", result)
-        print(json.dumps({"directory": str(args.directory), "shell": args.shell_kind, "checks": result["checks"]}, ensure_ascii=False), flush=True)
-    return 1  # Startup alone cannot open the full feasibility gate.
-
-
-def aggregate(root):
-    """Validate retained runtime evidence, including separate-call control artifacts."""
-    result = {"gate": "incomplete", "shells": {}, "checks": {}, "artifacts": str(root),
-              "scope": "Native Windows x64 at observed dimensions with an explicit same-page snapshot checkpoint",
-              "required_production_work": ["Resize/redraw-aware completion decoding",
-                                           "General capture freshness fix and acceptance tests in Tasks 8–9; stalled takes must fail"]}
-
-    def check(name, operation):
-        try:
-            detail = operation()
-            result["checks"][name] = {"status": "passed", "evidence": detail}
-        except FileNotFoundError as error:
-            result["checks"][name] = {"status": "unavailable", "error": str(error)}
-        except Exception as error:
-            result["checks"][name] = {"status": "failed", "error": repr(error), "traceback": traceback.format_exc()}
-
-    def prerequisites():
-        data = read_json(root / "prerequisites.json")
-        expected = {"uv-version", "python-identity", "ffmpeg-version", "ffmpeg-filters", "ffmpeg-encoders",
-                    "ffmpeg-devices", "ffprobe-version", "ttyd-version", "powershell51-version",
-                    "powershell7-version", "gitbash-version", "gitbash-native-tools", "feature-libx264",
-                    "feature- aac ", "feature-subtitles", "feature---enable-libass"}
-        assert expected <= set(data["commands"])
-        assert all(record["status"] == "passed" for record in data["commands"].values())
-        assert all(record["pe_machine"] == "0x8664" for record in data["binaries"].values())
-        return "prerequisites.json"
-
-    check("prerequisites", prerequisites)
-    for kind in ("powershell51", "powershell7", "gitbash"):
-        normal_path = root / f"standard-normal-{kind}"
-
-        def report(mode):
-            return read_json(root / f"standard-{mode}-{kind}" / "probe.json")
-
-        def recorded(mode, name):
-            data = report(mode)
-            assert data["checks"][name]["status"] == "passed", data["checks"][name]
-            if mode != "loss":
-                assert not data.get("failure"), data.get("failure")
-            return f"standard-{mode}-{kind}/probe.json"
-
-        for name, mode in (("readiness", "normal"), ("framing", "normal"), ("command_outcomes", "outcomes"),
-                           ("extra_client", "normal"), ("normal_cleanup", "normal"), ("cancel_cleanup", "cancel")):
-            check(f"{kind}.{name}", lambda name=name, mode=mode: recorded(mode, name))
-
-        def continuity():
-            normal = report("normal")
-            first = read_json(normal_path / "control" / "first-phase.json")
-            second = read_json(normal_path / "control" / "second-phase.json")
-            before, after = normal["readiness"], second["after"]
-            take1, take2 = normal["takes"][:2]
-            delayed = first["delayed_request"]
-            assert take1["pending_at_end"]["id"] == take2["pending_at_begin"]["id"] == second["inspection"]["pending"]["id"] == delayed
-            assert take1["end"] <= first["returned_at"] < take2["begin"] < second["delayed"]["finished_at"]
-            assert second["delayed"]["outcome"] == "completed" and second["delayed"]["shell_success"] is True
-            assert take1["page_id"] == take2["page_id"] == after["page_id"]
-            assert take1["websocket_id"] == take2["websocket_id"] == after["websocket_id"]
-            original = normal["extra_client"]["original_round_trip"]
-            assert normal["extra_client"]["accepted"] is False and normal["extra_client"]["candidate_continuity"] is False
-            assert original["session"] == after["session"] and str(original["pid"]) == str(after["pid"])
-            shell = {key: normal[key] for key in ("terminal_client_count", "owned_children_remaining", "unrelated_sentinel_alive", "shutdown_seconds")}
-            shell.update(observed_session_id=after["session"], filmed_session_id=take2["session"],
-                         nonce_before=before["session"], nonce_after=after["session"],
-                         shell_pid_before=str(before["pid"]), shell_pid_after=str(after["pid"]),
-                         variable_after=after["variable"], long_command_survived_take_boundary=True,
-                         environment=normal["environment"], terminal_sizes=normal["terminal_sizes"],
-                         artifacts=normal["artifacts"])
-            assert_probe(shell)
-            result["shells"][kind] = shell
-            return {"delayed_request": delayed, "first_call_returned": first["returned_at"], "take_two_began": take2["begin"],
-                    "command_finished": second["delayed"]["finished_at"], "page_id": after["page_id"], "websocket_id": after["websocket_id"]}
-
-        check(f"{kind}.two_takes", continuity)
-
-        def ordinary_user():
-            tokens = [report(mode)["environment"]["token"] for mode in ("normal", "outcomes", "cancel", "loss")]
-            tokens.append(read_json(root / f"standard-crash-{kind}" / "control" / "ready.json")["token"])
-            assert all(not token["elevated"] and token["integrity_rid"] == 8192 for token in tokens)
-            return [{key: token[key] for key in ("pid", "elevated", "integrity_sid", "elevation_type", "desktop_session")} for token in tokens]
-
-        check(f"{kind}.ordinary_user", ordinary_user)
-
-        def browser_loss():
-            data = report("loss")
-            recorded("loss", "browser_loss")
-            recorded("loss", "normal_cleanup")
-            assert data["checks"]["browser_loss"]["continuity"] is False
-            assert data["failure"] and data["results"]
-            assert all(record["outcome"] == "interrupted" and record["shell_success"] is None for record in data["results"].values())
-            return data["checks"]["browser_loss"]
-
-        check(f"{kind}.browser_loss", browser_loss)
-
-        def crash():
-            data = read_json(root / f"standard-crash-{kind}" / "crash.json")
-            assert data["status"] == "passed" and data["owned_children_remaining"] == []
-            assert data["heartbeats_stopped"] and data["unrelated_sentinel_alive"] and data["sentinel_heartbeat_advanced"]
-            assert data["shutdown_seconds"] < 10
-            return {key: data[key] for key in ("shutdown_seconds", "owned_children_remaining", "heartbeats_stopped", "unrelated_sentinel_alive")}
-
-        check(f"{kind}.crash_cleanup", crash)
-
-        def interactive():
-            directory = root / f"tui-verified-{kind}"
-            phase = read_json(directory / "control" / "interactive-phase.json")
-            assert phase["interactive"]["outcome"] == "completed" and phase["interactive"]["shell_success"] is True
-            assert read_json(directory / "tui-ready.json")["pid"] > 0
-            data = read_json(directory / "probe.json")
-            assert data["environment"]["token"]["ordinary_user"] and not data.get("failure")
-            assert data["checks"]["normal_cleanup"]["status"] == "passed"
-            frames = read_json(directory / "interactive" / "frames.json")
-            assert frames and all(frame["session_id"] == data["readiness"]["session"] for frame in frames)
-            snapshot = next(request for request in data["requests"] if request["operation"] == "snapshot")
-            observed = read_json(directory / "control" / f"{snapshot['id']:06d}.result.json")
-            assert observed["pending"]["id"] == int(phase["interactive"]["request"])
-            assert (directory / observed["screenshot"]).stat().st_size > 0
-            assert observed["time"] < phase["interactive"]["finished_at"]
-            inspected = read_json(root / "active-tui-pixel-inspection.json")[kind]
-            assert inspected["status"] == "passed" and inspected["pending_request"] == 2
-            pixel_path = directory / "interactive" / f"{inspected['frame_index']:06d}.jpg"
-            assert hashlib.sha256(pixel_path.read_bytes()).hexdigest() == inspected["sha256"]
-            frame = frames[inspected["frame_index"]]
-            assert inspected["snapshot_at"] < frame["received_at"] < inspected["q_at"]
-            assert frame["metadata"]["timestamp"] < inspected["q_at"]
-            assert inspected["all_hold_acks_replied"] and inspected["hold_receive_timeouts"] > 0
-            text = (directory / "terminal.bin").read_bytes().decode("utf-8")
-            assert "Native Windows TUI λ" in text and "TUI_EXIT:q" in text
-            return {"directory": str(directory), "frames": len(frames), "native_TUI_and_non_ASCII": True}
-
-        check(f"{kind}.interactive", interactive)
-
-        def resize():
-            data = report("normal")
-            columns = [size["columns"] for size in data["terminal_sizes"]]
-            assert columns == [217, 167] and data["framing_chunk_columns"] == 60
-            return {"observed_columns": columns, "base64_chunk_columns": 60, "scope": data["framing_scope"]}
-
-        check(f"{kind}.bounded_resize", resize)
-
-        def captures():
-            data = report("normal")
-            counts = {}
-            for take in data["takes"]:
-                frames = read_json(normal_path / take["name"] / "frames.json")
-                assert frames
-                if "end_request" in take:
-                    assert take["frames"] == len(frames)
-                else:
-                    # The ownership take is stopped by cleanup, not an end-take
-                    # request. Its persisted frame manifest is the count source.
-                    assert take["name"] == "ownership" and data["close_operation"] == "close"
-                for index, frame in enumerate(frames):
-                    assert frame["session_id"] == data["readiness"]["session"]
-                    assert frame["page_id"] == take["page_id"] and frame["websocket_id"] == take["websocket_id"]
-                    assert (normal_path / take["name"] / f"{index:06d}.jpg").stat().st_size > 0
-                counts[take["name"]] = len(frames)
-            assert (normal_path / "ready.png").stat().st_size > 0
-            return counts
-
-        check(f"{kind}.capture_artifacts", captures)
-    try:
-        assert_gate(result)
-        result["gate"] = "passed"
-    except (AssertionError, KeyError):
-        result["gate"] = "failed" if any(c["status"] == "failed" for c in result["checks"].values()) else "incomplete"
-    write_json(root / "probe.json", result)
-    print(json.dumps(result, ensure_ascii=False, indent=2))
-    return 0 if result["gate"] == "passed" else 1
-
-
-def main():
-    parser = argparse.ArgumentParser(description=__doc__)
-    parser.add_argument("--assert-result", type=Path)
-    parser.add_argument("--aggregate", type=Path)
-    parser.add_argument("--startup", action="store_true")
-    parser.add_argument("--outcomes", action="store_true")
-    parser.add_argument("--emit-bash", action="store_true")
-    parser.add_argument("--serve", action="store_true")
-    parser.add_argument("--worker", choices=["parent", "child", "grandchild", "sentinel"])
-    parser.add_argument("--tui", action="store_true")
-    parser.add_argument("--crash-check", action="store_true")
-    parser.add_argument("--phase", choices=["first", "second", "interactive", "prepare-cleanup", "close", "cancel", "browser-loss"])
-    parser.add_argument("--submit")
-    parser.add_argument("--wait", type=float, default=10)
-    parser.add_argument("--wait-result", action="store_true")
-    parser.add_argument("--ttyd", type=Path)
-    parser.add_argument("--browser", type=Path)
-    parser.add_argument("--shell", type=Path)
-    parser.add_argument("--shell-kind", choices=["powershell51", "powershell7", "gitbash"])
-    parser.add_argument("--directory", type=Path)
-    parser.add_argument("--launch-host")
-    args = parser.parse_args()
-    if args.aggregate:
-        return aggregate(args.aggregate)
-    if args.emit_bash:
-        emit_bash()
-        return 0
-    if args.worker:
-        heartbeat_worker(args.directory, args.worker)
-        return 0
-    if args.tui:
-        tui_fixture()
-        return 0
-    if args.crash_check:
-        crash_check(args)
-        return 0
-    if args.phase:
-        drive_phase(args)
-        return 0
-    if args.submit:
-        submit_request(args)
-        return 0
-    if args.startup or args.outcomes or args.serve:
-        for name in ("ttyd", "browser", "shell", "directory"):
-            path = getattr(args, name)
-            if path is None or not path.is_absolute():
-                parser.error(f"--{name} requires an explicit absolute path")
-        if not args.shell_kind or not args.launch_host:
-            parser.error("--shell-kind and --launch-host are required")
-        return Supervisor(args).run() if args.serve else startup_probe(args)
-    assert_gate(json.loads(args.assert_result.read_text(encoding="utf-8"))
-                if args.assert_result else {})
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 153
tests/proving-it-works-with-a-movie/test-probe-cleanup.py

@@ -1,153 +0,0 @@
-#!/usr/bin/env python3
-"""Focused native Windows cleanup failures, verified with real process handles."""
-import argparse
-import contextlib
-import ctypes
-import hashlib
-import importlib.util
-import json
-from pathlib import Path
-import subprocess
-import sys
-import time
-import traceback
-from unittest.mock import patch
-
-
-def exercise(probe, root, case):
-    directory = root / case
-    supervisor = probe.Supervisor(argparse.Namespace(directory=directory, shell_kind="powershell51",
-                                                     launch_host="workerbee"))
-    terminal, job = supervisor.terminal, supervisor.terminal.job
-    real_close = terminal.close
-    handles = []
-    observed = {"case": case, "passed": False, "token": supervisor.report["environment"]["token"]}
-
-    def fail(label):
-        def injected(*args, **kwargs):
-            raise PermissionError(f"injected persistent {label} failure")
-        return injected
-
-    try:
-        job.spawn([sys.executable, str(Path(probe.__file__).resolve()), "--worker", "parent",
-                   "--directory", str(directory / "heartbeats")], directory, directory / "worker.log")
-        supervisor.sentinel = subprocess.Popen([sys.executable, str(Path(probe.__file__).resolve()),
-                                               "--worker", "sentinel", "--directory", str(directory / "sentinel")],
-                                              stdin=subprocess.DEVNULL, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
-        deadline = time.monotonic() + 10
-        while time.monotonic() < deadline:
-            before = probe.heartbeats(directory / "heartbeats")
-            if set(before) == {"parent", "child", "grandchild"} and probe.heartbeats(directory / "sentinel"):
-                break
-            time.sleep(0.05)
-        assert set(before) == {"parent", "child", "grandchild"}, before
-        handles = job.snapshot()  # Independent verifier retains handles before injecting faults.
-        assert {entry["pid"] for entry in before.values()} <= {entry["pid"] for entry in handles}
-        sentinel_handle = job.open_process(supervisor.sentinel.pid)
-        try:
-            member = ctypes.c_long()
-            assert job.k.IsProcessInJob(sentinel_handle, job.handle, ctypes.byref(member))
-            assert member.value == 0, "Sentinel must be outside the owned job"
-        finally:
-            job.k.CloseHandle(sentinel_handle)
-        supervisor.pending = {"id": 1, "command": "native heartbeat parent fixture"}
-        supervisor.report["close_operation"] = "close"
-        observed["before"] = before
-        with contextlib.ExitStack() as faults:
-            if case == "snapshot":
-                faults.enter_context(patch.object(job, "snapshot", fail("snapshot")))
-            elif case == "heartbeat":
-                faults.enter_context(patch.object(probe, "heartbeats", fail("heartbeat access")))
-            elif case == "pending_reply":
-                faults.enter_context(patch.object(supervisor, "reply", fail("pending reply")))
-            elif case == "terminal_job":
-                faults.enter_context(patch.object(terminal, "close", fail("terminal close")))
-
-                def termination_failure(*args):
-                    ctypes.set_last_error(5)  # Access denied; real CloseHandle must still kill the job.
-                    return 0
-
-                faults.enter_context(patch.object(job.k, "TerminateJobObject", termination_failure))
-            elif case == "report_write":
-                real_write = probe.write_json
-
-                def report_failure(path, value):
-                    if path.name == "probe.json":
-                        raise PermissionError("injected persistent report write failure")
-                    return real_write(path, value)
-
-                faults.enter_context(patch.object(probe, "write_json", report_failure))
-            started = time.monotonic()
-            try:
-                observed["cleanup_return"] = supervisor.cleanup(None)
-            except BaseException as error:
-                observed.update(escaped_error=repr(error), escaped_traceback=traceback.format_exc())
-            observed["cleanup_seconds"] = time.monotonic() - started
-        observed["owned_alive_after_cleanup"] = [p["pid"] for p in handles if job.k.WaitForSingleObject(p["handle"], 0) != 0]
-        observed["sentinel_alive_after_cleanup"] = supervisor.sentinel.poll() is None
-        after = probe.heartbeats(directory / "heartbeats")
-        time.sleep(0.3)
-        observed["heartbeats_stopped"] = after == probe.heartbeats(directory / "heartbeats")
-        observed["probe_report_exists"] = (directory / "probe.json").exists()
-        observed["report"] = supervisor.report
-        assert observed["owned_alive_after_cleanup"] == [], observed
-        assert not observed["sentinel_alive_after_cleanup"], observed
-        assert observed["heartbeats_stopped"], observed
-        assert observed["cleanup_seconds"] < 10, observed
-        assert "escaped_error" not in observed, observed
-        assert observed["cleanup_return"] is (case == "normal"), observed
-        expected_status = "passed" if case == "normal" else "failed"
-        assert supervisor.report["checks"]["normal_cleanup"]["status"] == expected_status, observed
-        if case != "normal":
-            assert supervisor.report["cleanup_diagnostics"], observed
-            expected = {"snapshot": {"snapshot"}, "heartbeat": {"heartbeats_before", "heartbeats_after"},
-                        "pending_reply": {"pending_reply"}, "terminal_job": {"terminal_close", "job_close"},
-                        "report_write": {"report_write"}}[case]
-            assert expected <= {d["stage"] for d in supervisor.report["cleanup_diagnostics"]}, observed
-        assert observed["probe_report_exists"] is (case != "report_write"), observed
-        if case != "report_write":
-            assert probe.read_json(directory / "probe.json")["checks"]["normal_cleanup"]["status"] == expected_status
-        observed["passed"] = True
-    except BaseException as error:
-        observed.update(assertion_error=repr(error), assertion_traceback=traceback.format_exc())
-    finally:
-        # Emergency verifier teardown is outside the code under test. The above
-        # observations are captured before it, so RED cannot borrow this cleanup.
-        try:
-            real_close()
-        finally:
-            if supervisor.sentinel and supervisor.sentinel.poll() is None:
-                supervisor.sentinel.terminate()
-                supervisor.sentinel.wait(timeout=5)
-            for process in handles:
-                job.k.WaitForSingleObject(process["handle"], 5000)
-                job.k.CloseHandle(process["handle"])
-        probe.write_json(directory / "test-observation.json", observed)
-    return observed
-
-
-def main():
-    parser = argparse.ArgumentParser(description=__doc__)
-    parser.add_argument("--directory", type=Path, required=True)
-    parser.add_argument("--probe", type=Path, default=Path(__file__).with_name("probe-windows.py"))
-    args = parser.parse_args()
-    if sys.platform != "win32":
-        parser.error("Native Windows is required; this test cannot be skipped green")
-    args.directory.mkdir(parents=True, exist_ok=False)
-    spec = importlib.util.spec_from_file_location("windows_probe", args.probe)
-    probe = importlib.util.module_from_spec(spec)
-    spec.loader.exec_module(probe)
-    cases = [exercise(probe, args.directory, case) for case in
-             ("snapshot", "heartbeat", "pending_reply", "terminal_job", "report_write", "normal")]
-    report = {"probe_sha256": hashlib.sha256(args.probe.read_bytes()).hexdigest(),
-              "test_sha256": hashlib.sha256(Path(__file__).read_bytes()).hexdigest(),
-              "cases": cases, "passed": sum(case["passed"] for case in cases), "total": len(cases)}
-    probe.write_json(args.directory / "cleanup-tests.json", report)
-    print(json.dumps({"passed": report["passed"], "total": report["total"],
-                      "cases": [{key: c.get(key) for key in ("case", "passed", "owned_alive_after_cleanup",
-                                                             "sentinel_alive_after_cleanup", "escaped_error")} for c in cases]}), flush=True)
-    return 0 if report["passed"] == report["total"] else 1
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 28 - 5
tests/proving-it-works-with-a-movie/test_terminal.py

@@ -183,12 +183,35 @@ class NativeTerminalTests(unittest.TestCase):
         self.run_command('echo recovered')
         self.request('close');self.assertEqual(self.server.wait(10),0)
 
+    def outcome_commands(self):
+        """Shell outcomes the recorder must attribute correctly: (name, command, native_producer)."""
+        exe=os.environ['MOVIE_TEST_SHELL_EXE'];python=sys.executable
+        commands=[('native_success',self.native('import sys;sys.exit(0)'),python)]
+        if self.shell!='gitbash':
+            commands+=[
+                ('cmdlet_failure',"Get-Item 'Z:\\probe-path-that-does-not-exist'",None),
+                ('native_failure',self.native('import sys;sys.exit(7)'),python),
+                ('cmdlet_success',"Write-Output 'cmdlet success λ'",None),
+                ('terminating_error',"throw 'probe terminating error'",None),
+                ('nonterminating_error',"Write-Error 'probe nonterminating error'",None),
+                ('parse_failure','Write-Output )',None),
+                ('logging_failure',self.native("import sys;print('producer');sys.exit(7)")+' | Tee-Object -Variable ProbeLog',python),
+                ('expression_wrapper',"(Write-Error 'probe expression wrapper')",None),
+                ('script_exit',f"& '{exe}' -NoLogo -NoProfile -Command 'exit 9'",exe),
+            ]
+        else:
+            commands+=[
+                ('shell_failure','test -e /probe-path-that-does-not-exist',None),
+                ('native_failure',self.native('import sys;sys.exit(7)'),python),
+                ('shell_success',"printf 'shell success λ\\n'",None),
+                ('parse_failure','echo )',None),
+                ('logging_failure',self.native("import sys;print('producer');sys.exit(7)")+' | cat',python),
+                ('script_exit',"bash --noprofile --norc -c 'exit 9'",'Git Bash'),
+            ]
+        return commands
+
     def test_native_shell_outcomes(self):
-        spec=importlib.util.spec_from_file_location('historical_probe',Path(__file__).with_name('probe-windows.py'))
-        probe=importlib.util.module_from_spec(spec);spec.loader.exec_module(probe)
-        import argparse
-        args=argparse.Namespace(shell_kind=self.shell,shell=os.environ['MOVIE_TEST_SHELL_EXE'])
-        for name,command,native in probe.outcome_commands(args):
+        for name,command,native in self.outcome_commands():
             success=name in ('native_success','cmdlet_success','shell_success')
             result=self.run_command(command,native_producer=native,expected=0 if success else 1)
             self.assertEqual(result['outcome'],'completed')