Przeglądaj źródła

tests: update SDD assertions and prompts to current skill behavior

The SDD skill tests still asserted on pre-rename skill text, so correct
model answers failed the suite:

- "Read at beginning" asserted "Step 1|beginning|start|Load Plan";
  the current skill has no numbered steps or Load Plan phase (setup
  covers it: "Read the plan once" during Setup). A live run today
  failed this assertion when the model correctly answered "during
  setup". Pattern now accepts setup/before-dispatch paraphrases while
  still requiring an at-the-start answer.
- "Provides text directly" asserted the removed provide-full-task-text
  behavior; SDD now routes task requirements through brief files
  (scripts/task-brief). The test now asks brief-file-vs-whole-plan and
  asserts the brief-based flow.
- The integration test's prompt and summary told the agent to
  "provide full task text to subagents (don't make them read files)",
  contradicting the skill it verifies; reworded to the task-brief flow.
  Also gave its direct `timeout 1800 claude -p` the same </dev/null
  stdin guard as run_claude.

Live-LLM tests; verified with bash -n on every touched file.

Part of #2130; defects documented in PR #2071 by @ericyen97903-lab.
Jesse Vincent 1 miesiąc temu
rodzic
commit
c7b456dd0e

+ 5 - 5
tests/claude-code/test-subagent-driven-development-integration.sh

@@ -23,7 +23,7 @@ echo "========================================"
 echo ""
 echo "This test executes a real plan using the skill and verifies:"
 echo "  1. Plan is read once (not per task)"
-echo "  2. Full task text provided to subagents"
+echo "  2. Task requirements routed to subagents via brief files"
 echo "  3. Subagents perform self-review"
 echo "  4. Spec compliance review before code quality"
 echo "  5. Review loops when issues found"
@@ -136,7 +136,7 @@ I want you to execute the implementation plan at docs/superpowers/plans/implemen
 
 IMPORTANT: Follow the skill exactly. I will be verifying that you:
 1. Read the plan once at the beginning
-2. Provide full task text to subagents (don't make them read files)
+2. Route each task's requirements to subagents via a task brief file (don't make them read the whole plan)
 3. Ensure subagents do self-review before reporting
 4. Run spec compliance review before code quality review
 5. Use review loops when issues are found
@@ -150,7 +150,7 @@ PROMPT="Execute the implementation plan at docs/superpowers/plans/implementation
 
 IMPORTANT: Follow the skill exactly. I will be verifying that you:
 1. Read the plan once at the beginning
-2. Provide full task text to subagents (don't make them read files)
+2. Route each task's requirements to subagents via a task brief file (don't make them read the whole plan)
 3. Ensure subagents do self-review before reporting
 4. Run spec compliance review before code quality review
 5. Use review loops when issues are found
@@ -164,7 +164,7 @@ PLUGIN_DIR=$(cd "$SCRIPT_DIR/../.." && pwd)
 # other concurrent claude sessions.
 echo "Running Claude (plugin-dir: $PLUGIN_DIR, cwd: $TEST_PROJECT)..."
 echo "================================================================================"
-cd "$TEST_PROJECT" && timeout 1800 claude -p "$PROMPT" --plugin-dir "$PLUGIN_DIR" --allowed-tools=all --permission-mode bypassPermissions 2>&1 | tee "$OUTPUT_FILE" || {
+cd "$TEST_PROJECT" && timeout 1800 claude -p "$PROMPT" --plugin-dir "$PLUGIN_DIR" --allowed-tools=all --permission-mode bypassPermissions < /dev/null 2>&1 | tee "$OUTPUT_FILE" || {
     echo ""
     echo "================================================================================"
     echo "EXECUTION FAILED (exit code: $?)"
@@ -316,7 +316,7 @@ if [ $FAILED -eq 0 ]; then
     echo ""
     echo "The subagent-driven-development skill correctly:"
     echo "  ✓ Reads plan once at start"
-    echo "  ✓ Provides full task text to subagents"
+    echo "  ✓ Routes task requirements via brief files"
     echo "  ✓ Enforces self-review"
     echo "  ✓ Runs spec compliance before code quality"
     echo "  ✓ Spec reviewer verifies independently"

+ 6 - 6
tests/claude-code/test-subagent-driven-development.sh

@@ -4,7 +4,7 @@
 #
 # No drill coverage: this test asks the agent to *describe* SDD (string-
 # matches its verbal explanation against expected keywords like
-# "self-review", "skeptical", "worktree", "Step 1", "loop"). Drill scenarios
+# "self-review", "skeptical", "worktree", "setup", "loop"). Drill scenarios
 # test behavior (real subagent dispatch, plan-following, review loops),
 # not description-recall. Kept by design.
 set -euo pipefail
@@ -83,7 +83,7 @@ else
     exit 1
 fi
 
-if assert_contains "$output" "Step 1\|beginning\|start\|Load Plan" "Read at beginning"; then
+if assert_contains "$output" "beginning\|start\|setup\|before.*dispatch\|before.*task" "Read at beginning"; then
     : # pass
 else
     exit 1
@@ -133,16 +133,16 @@ echo ""
 echo "Test 7: Task context provision..."
 
 output=$(run_claude "In subagent-driven-development, how does the controller provide task information to the implementer subagent? Answer using exactly this structure:
-Controller provides: <directly or by file>
-Implementer must read plan file: <yes or no>" "$CLAUDE_PROMPT_TIMEOUT")
+Controller provides: <brief file or whole plan file>
+Implementer must read whole plan file: <yes or no>" "$CLAUDE_PROMPT_TIMEOUT")
 
-if assert_contains "$output" "provide.*directly\|full.*text\|paste\|include.*prompt" "Provides text directly"; then
+if assert_contains "$output" "task-brief\|brief file\|brief.*path\|Controller provides:.*brief" "Provides task brief file"; then
     : # pass
 else
     exit 1
 fi
 
-if assert_contains "$output" "Implementer must read plan file:.*no" "Doesn't make subagent read file"; then
+if assert_contains "$output" "Implementer must read whole plan file:.*no" "Doesn't make subagent read whole plan"; then
     : # pass
 else
     exit 1