mirror of
https://github.com/garrytan/gstack.git
synced 2026-05-08 13:39:45 +08:00
* feat: gstack-gbrain-mcp-verify helper for remote MCP probe
Probes a remote gbrain MCP endpoint with bearer auth. POSTs initialize,
classifies failures into NETWORK / AUTH / MALFORMED with one-line
remediation hints, and runs a tools/list capability probe to detect
sources_add MCP support (forward-compat for when gbrain ships URL ingest).
Token consumed from GBRAIN_MCP_TOKEN env, never argv. Required to set
both 'application/json' AND 'text/event-stream' in Accept; that gotcha
costs 10 minutes of debugging when missed (regression-tested).
Live-verified against wintermute (gbrain v0.27.1).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: gstack-artifacts-init + gstack-artifacts-url helpers
artifacts-init replaces brain-init with provider choice (gh / glab /
manual), per-user gstack-artifacts-$USER repo, HTTPS-canonical storage in
~/.gstack-artifacts-remote.txt, and a "send this to your brain admin"
hookup printout. Always prints the command, never auto-executes — gbrain
v0.26.x has no admin-scope MCP probe (codex Finding #3).
artifacts-url centralizes HTTPS↔SSH/host/owner-repo conversion so callers
don't each string-mangle (codex Finding #10). The remote-conflict check in
artifacts-init compares at the canonical level so re-running with HTTPS
input doesn't trip on a stored SSH URL for the same logical repo.
The "URL form not supported" branch prints a two-line clone-then-path
form for gbrain v0.26.x; the supported branch is a one-liner with --url
ready for when gbrain ships URL ingest.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: extend gstack-gbrain-detect with mcp_mode + artifacts_remote
Adds two new fields to detect's JSON output:
- gbrain_mcp_mode: local-stdio | remote-http | none
Resolved via 3-tier fallback (codex Finding D3): claude mcp get --json
→ claude mcp list text-grep → ~/.claude.json jq read. If Anthropic moves
the file format, the first two tiers absorb it.
- gstack_artifacts_remote: HTTPS URL from ~/.gstack-artifacts-remote.txt
Falls back to ~/.gstack-brain-remote.txt during the v1.27.0.0 migration
window so detect doesn't return empty between upgrade and migration.
Existing detect tests still pass (15/15). New 19 tests cover every fallback
tier independently, plus a schema regression for /sync-gbrain compat.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: setup-gbrain Path 4 (remote MCP) + artifacts rename
Path 4 lets users paste an HTTPS MCP URL + bearer token and registers it
as an HTTP-transport MCP without needing a local gbrain CLI install. The
flow:
- Step 2 gains a fourth option (Remote gbrain MCP)
- Step 4 adds Path 4 sub-flow: collect URL, secret-read bearer, verify
via gstack-gbrain-mcp-verify (NETWORK / AUTH / MALFORMED classifier)
- Step 5 (local doctor), Step 7.5 (transcript ingest), Step 5a's stdio
branch all skip on Path 4
- Step 5a adds an HTTP+bearer registration form: claude mcp add
--transport http --header "Authorization: Bearer ..."
- Step 7 renamed "session memory sync" → "artifacts sync" and now calls
gstack-artifacts-init (which always prints the brain-admin hookup
command — no auto-execute, codex Finding #3)
- Step 8 CLAUDE.md block branches: remote-http includes URL + server
version (never the token); local-stdio keeps engine + config-file
- Step 9 smoke test on Path 4 prints the curl-equivalent for
post-restart verification (MCP tools aren't visible mid-session)
- Step 10 verdict block has separate templates per mode
Idempotency: re-running with gbrain_mcp_mode=remote-http already in
detect output skips Step 2 entirely and goes to verification.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor: rename gbrain_sync_mode → artifacts_sync_mode (v1.27.0.0 prep)
Hard rename, no dual-read alias (codex Finding D4). The on-disk migration
script (Phase C, separate commit) renames the config key in users'
~/.gstack/config.yaml and any CLAUDE.md blocks.
Touched call sites:
- bin/gstack-config defaults + validation + list/defaults output
- bin/gstack-gbrain-detect (gstack_brain_sync_mode field still emitted
with the same name for downstream-tool compat; reads new key)
- bin/gstack-brain-sync, bin/gstack-brain-enqueue, bin/gstack-brain-uninstall
- bin/gstack-timeline-log (comment ref)
- scripts/resolvers/preamble/generate-brain-sync-block.ts: renames key,
branches on gbrain_mcp_mode=remote-http to emit "ARTIFACTS_SYNC:
remote-mode (managed by brain server <host>)" instead of the local
mode/queue/last_push line (codex Finding #11)
- bin/gstack-brain-restore + bin/gstack-gbrain-source-wireup: read
~/.gstack-artifacts-remote.txt with ~/.gstack-brain-remote.txt fallback
during the migration window
- bin/gstack-artifacts-init: tolerant of unrecognized URL forms (local
paths, file://, self-hosted gitea) so test infrastructure and unusual
remotes work without canonicalization
- test/brain-sync.test.ts: gstack-brain-init → gstack-artifacts-init
- test/skill-e2e-brain-privacy-gate.test.ts: artifacts_sync_mode keys
- test/gen-skill-docs.test.ts: budget 35K → 36.5K for the new MCP-mode
probe in the preamble resolver
- health/SKILL.md.tmpl, sync-gbrain/SKILL.md.tmpl: comment + verdict line
Hard delete:
- bin/gstack-brain-init (replaced by bin/gstack-artifacts-init in v1.27.0.0)
- test/gstack-brain-init-gh-mock.test.ts (replaced by gstack-artifacts-init.test.ts)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: regenerate SKILL.md files after artifacts-sync rename
Mechanical regen via \`bun run gen:skill-docs --host all\`. All */SKILL.md
files reflect the renamed config key (gbrain_sync_mode →
artifacts_sync_mode), the renamed remote-helper file
(~/.gstack-artifacts-remote.txt with brain fallback), the renamed init
script (gstack-artifacts-init), and the new ARTIFACTS_SYNC: remote-mode
status line that fires when a remote-http MCP is registered.
Golden fixtures (test/fixtures/golden/*-ship-SKILL.md) refreshed to match
the regenerated default-ship output.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: v1.27.0.0 migration — gstack-brain → gstack-artifacts rename
Journaled, interruption-safe migration. Six steps, each writes to
~/.gstack/.migrations/v1.27.0.0.journal on success; re-entry resumes
from the next un-done step. On final success, journal is replaced by
~/.gstack/.migrations/v1.27.0.0.done.
Steps:
1. gh_repo_renamed gh/glab repo rename gstack-brain-$USER →
gstack-artifacts-$USER (idempotent: detects
already-renamed and skips)
2. remote_txt_renamed mv ~/.gstack-brain-remote.txt → artifacts file,
rewriting URL path to match the new repo name
3. config_key_renamed sed -i in ~/.gstack/config.yaml flips
gbrain_sync_mode → artifacts_sync_mode
4. claude_md_block sed flips "- Memory sync:" → "- Artifacts sync:"
in cwd CLAUDE.md and ~/.gstack/CLAUDE.md
5. sources_swapped gbrain sources add NEW (verify) → remove OLD
(codex Finding #6: add-before-remove ordering,
no downtime window). On remote-MCP mode, prints
commands for the brain admin instead of executing.
6. done touchfile + delete journal
User opt-out: any "n" or "skip-for-now" answer at the initial prompt
writes a marker file that prevents re-prompting; user can re-invoke
via /setup-gbrain --rerun-migration.
11 unit tests cover: nothing-to-migrate, GitHub happy path, idempotent
re-run, journal-resume mid-flight, remote-MCP print-only path,
add-before-remove ordering verification, add-fail → old source stays
registered, CLAUDE.md field rewrite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test: regression suite + E2E for v1.27.0.0 rename
Three new regression tests guard the rename's blast radius (per codex
Findings #1, #8, #9, #12):
- test/no-stale-gstack-brain-refs.test.ts: greps bin/, scripts/, *.tmpl,
test/ for forbidden identifiers (gstack-brain-init, gbrain_sync_mode);
fails CI if any non-allowlisted file references them.
- test/post-rename-doc-regen.test.ts: confirms gen-skill-docs output has
no stale references in any */SKILL.md (the cross-product blind spot).
- test/setup-gbrain-path4-structure.test.ts: structural lint over the
Path 4 prose contract — STOP gates after verify failure, never-write-
token rules, mode-aware CLAUDE.md block, bearer always via env-var.
Two new gate-tier E2E tests (deterministic stub HTTP server, fixed inputs):
- test/skill-e2e-setup-gbrain-remote.test.ts: Path 4 happy path. Stubs
an HTTP MCP server, drives the skill via Agent SDK with a stubbed
bearer, asserts claude.json gets the http MCP entry, CLAUDE.md gets
the remote-http block, the secret token NEVER leaks to CLAUDE.md.
- test/skill-e2e-setup-gbrain-bad-token.test.ts: stub server returns 401;
asserts the AUTH classifier hint surfaces, no MCP registration occurs,
CLAUDE.md is unchanged. Regression guard for the "verify failed → STOP"
rule.
touchfiles.ts: setup-gbrain-remote and setup-gbrain-bad-token added at
gate-tier so CI catches Path 4 regressions on every PR.
Plus a few comment refs flipped: bin/gstack-jsonl-merge, bin/gstack-timeline-log
(legacy gstack-brain-init mentions in headers).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* release: v1.27.0.0 — /setup-gbrain Path 4 + brain → artifacts rename
Bumps VERSION 1.26.4.0 → 1.27.0.0 (MINOR per CLAUDE.md scale-aware bump
guidance: ~1500 line net change including a new path in /setup-gbrain,
two new bin helpers, a journaled migration, 59 new tests, and a config
key rename across the codebase).
CHANGELOG entry covers: Path 4 (Remote MCP) end-to-end, the brain →
artifacts rename, the journaled migration, the verify-helper error
classifier, the artifacts-init multi-host provider choice. Includes
the canonical Garry-voice headline + numbers table + audience close
per the release-summary format.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test: demote setup-gbrain Path 4 E2E to periodic-tier
The Agent SDK E2E tests for Path 4 (skill-e2e-setup-gbrain-remote and
skill-e2e-setup-gbrain-bad-token) are inherently non-deterministic —
the model interprets "follow Path 4 only" prompts flexibly and can
skip Step 8 (CLAUDE.md write) or shortcut past the verify helper, which
makes the gate-tier assertions flaky.
The deterministic gate coverage for Path 4 is in
test/setup-gbrain-path4-structure.test.ts: a fast structural lint that
catches AUQ-pacing regressions and prose contract drift in <200ms with
zero token spend. That test is the right tool for catching the failure
mode the gate-tier was meant to guard against.
The Agent SDK E2E tests stay available on-demand for periodic-tier runs
(EVALS=1 EVALS_TIER=periodic bun test test/skill-e2e-setup-gbrain-*.test.ts).
Also tightened the verify-error assertion to the literal field shape
("error_class": "AUTH") instead of a substring match that false-matches
the parent claude session's "needs-auth" MCP discovery markers.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: sync package.json version to 1.27.0.0
VERSION was bumped to 1.27.0.0 in f6ec11eb but package.json was not
updated in the same commit. The gen-skill-docs.test.ts assertion
"package.json version matches VERSION file" caught the drift.
This is the DRIFT_STALE_PKG case the /ship Step 12 idempotency check
is designed for; the fix is the documented sync-only repair (no
re-bump, package.json synced to existing VERSION).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
228 lines
9.5 KiB
TypeScript
228 lines
9.5 KiB
TypeScript
/**
|
|
* Privacy-gate E2E (periodic tier, paid).
|
|
*
|
|
* The gbrain-sync preamble block instructs the model to fire a one-time
|
|
* AskUserQuestion when:
|
|
* - `BRAIN_SYNC: off` in the preamble echo (sync mode not on)
|
|
* - config `artifacts_sync_mode_prompted` is "false"
|
|
* - gbrain is detected on the host (binary on PATH or `gbrain doctor`
|
|
* --fast --json succeeds)
|
|
*
|
|
* This test stages all three conditions (via env + a fake `gbrain` binary
|
|
* on PATH), runs a cheap gstack skill through the Agent SDK, intercepts
|
|
* every tool use via canUseTool, and asserts: one of the AskUserQuestions
|
|
* fired by the preamble is the privacy gate with its distinctive prose
|
|
* and three options (full / artifacts-only / decline).
|
|
*
|
|
* Cost: ~$0.30-$0.50 per run. Periodic tier (EVALS=1 EVALS_TIER=periodic).
|
|
*
|
|
* See scripts/resolvers/preamble/generate-brain-sync-block.ts for the
|
|
* prose contract this test locks in.
|
|
*/
|
|
|
|
import { describe, test, expect } from 'bun:test';
|
|
import * as fs from 'fs';
|
|
import * as os from 'os';
|
|
import * as path from 'path';
|
|
import { runAgentSdkTest, passThroughNonAskUserQuestion, resolveClaudeBinary } from './helpers/agent-sdk-runner';
|
|
|
|
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
|
|
const describeE2E = shouldRun ? describe : describe.skip;
|
|
|
|
describeE2E('gbrain-sync privacy gate fires once via preamble', () => {
|
|
test('gstack skill preamble fires the 3-option AskUserQuestion when gbrain is detected', async () => {
|
|
// Stage a fresh GSTACK_HOME with artifacts_sync_mode_prompted=false.
|
|
const gstackHome = fs.mkdtempSync(path.join(os.tmpdir(), 'privacy-gate-gstack-'));
|
|
const fakeBinDir = fs.mkdtempSync(path.join(os.tmpdir(), 'privacy-gate-bin-'));
|
|
|
|
// Seed the config so the gate's condition passes.
|
|
fs.writeFileSync(
|
|
path.join(gstackHome, 'config.yaml'),
|
|
'artifacts_sync_mode: off\nartifacts_sync_mode_prompted: false\n',
|
|
{ mode: 0o600 }
|
|
);
|
|
|
|
// Fake `gbrain` binary that makes the host-detection probe succeed.
|
|
// The preamble checks `gbrain doctor --fast --json` OR `which gbrain`.
|
|
// Either branch counts as "gbrain detected."
|
|
fs.writeFileSync(
|
|
path.join(fakeBinDir, 'gbrain'),
|
|
'#!/bin/bash\n' +
|
|
'case "$1" in\n' +
|
|
' doctor) echo \'{"status":"ok","schema_version":2}\' ; exit 0 ;;\n' +
|
|
' --version) echo "0.18.2" ; exit 0 ;;\n' +
|
|
' *) exit 0 ;;\n' +
|
|
'esac\n',
|
|
{ mode: 0o755 }
|
|
);
|
|
|
|
const askUserQuestions: Array<{ input: Record<string, unknown> }> = [];
|
|
const binary = resolveClaudeBinary();
|
|
|
|
// Ambient env mutations — restored in finally so other tests in the file
|
|
// don't inherit them.
|
|
const origGstackHome = process.env.GSTACK_HOME;
|
|
const origPath = process.env.PATH;
|
|
process.env.GSTACK_HOME = gstackHome;
|
|
process.env.PATH = `${fakeBinDir}:${process.env.PATH ?? '/usr/bin:/bin:/opt/homebrew/bin'}`;
|
|
|
|
try {
|
|
// Pick a small skill with the preamble and load it via Read to force
|
|
// the model to execute every preamble directive. A narrow "run /learn"
|
|
// prompt often gets reduced to a direct action, skipping the preamble
|
|
// gates. Mirror the plan-mode-no-op test pattern: ask the model to
|
|
// follow the skill's instructions in full.
|
|
const learnSkill = path.resolve(
|
|
import.meta.dir,
|
|
'..',
|
|
'learn',
|
|
'SKILL.md'
|
|
);
|
|
await runAgentSdkTest({
|
|
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
|
userPrompt:
|
|
`Read the skill file at ${learnSkill} and follow its instructions from the top, including every preamble directive. Execute every bash block. If any AskUserQuestion fires, present it.`,
|
|
workingDirectory: gstackHome,
|
|
maxTurns: 10,
|
|
allowedTools: ['Read', 'Grep', 'Glob', 'Bash'],
|
|
// NOTE: do NOT pass `env:` here. When the Agent SDK gets an explicit
|
|
// env object, its auth pipeline doesn't pick up ANTHROPIC_API_KEY the
|
|
// same way as when env is undefined (SDK-internal detail, verified
|
|
// against the plan-mode-no-op test which passes no env and auths
|
|
// cleanly). Instead, mutate process.env before the call so the SDK
|
|
// inherits our overrides ambiently.
|
|
...(binary ? { pathToClaudeCodeExecutable: binary } : {}),
|
|
canUseTool: async (toolName, input) => {
|
|
if (toolName === 'AskUserQuestion') {
|
|
askUserQuestions.push({ input });
|
|
// Auto-answer "Decline — keep everything local" (option C)
|
|
// so the skill can continue without actually turning on sync.
|
|
const q = (input.questions as Array<{
|
|
question: string;
|
|
options: Array<{ label: string }>;
|
|
}>)[0];
|
|
const decline =
|
|
q.options.find((o) => /decline|keep everything local|no thanks/i.test(o.label)) ??
|
|
q.options[q.options.length - 1]!;
|
|
return {
|
|
behavior: 'allow',
|
|
updatedInput: {
|
|
questions: input.questions,
|
|
answers: { [q.question]: decline.label },
|
|
},
|
|
};
|
|
}
|
|
return passThroughNonAskUserQuestion(toolName, input);
|
|
},
|
|
});
|
|
|
|
// Assertion 1: the privacy gate fired.
|
|
const privacyQuestions = askUserQuestions.filter((aq) => {
|
|
const qs = aq.input.questions as Array<{ question: string }>;
|
|
return qs.some(
|
|
(q) =>
|
|
/publish.*session memory|private github repo|gbrain indexes/i.test(q.question)
|
|
);
|
|
});
|
|
expect(privacyQuestions.length).toBeGreaterThanOrEqual(1);
|
|
|
|
// Assertion 2: the question has the three expected options.
|
|
const gate = privacyQuestions[0]!.input.questions as Array<{
|
|
question: string;
|
|
options: Array<{ label: string }>;
|
|
}>;
|
|
const labels = gate[0]!.options.map((o) => o.label.toLowerCase()).join(' | ');
|
|
// Full / artifacts-only / decline are the three canonical options.
|
|
expect(labels).toMatch(/everything|allowlisted|full/);
|
|
expect(labels).toMatch(/artifact/);
|
|
expect(labels).toMatch(/decline|local|no thanks/);
|
|
|
|
// Assertion 3: the gate should NOT fire twice in one run.
|
|
// (The preamble is supposed to be idempotent within a session.)
|
|
expect(privacyQuestions.length).toBe(1);
|
|
} finally {
|
|
// Restore ambient env before other tests.
|
|
if (origGstackHome === undefined) delete process.env.GSTACK_HOME;
|
|
else process.env.GSTACK_HOME = origGstackHome;
|
|
if (origPath === undefined) delete process.env.PATH;
|
|
else process.env.PATH = origPath;
|
|
fs.rmSync(gstackHome, { recursive: true, force: true });
|
|
fs.rmSync(fakeBinDir, { recursive: true, force: true });
|
|
}
|
|
}, 180_000);
|
|
|
|
test('privacy gate does NOT fire when artifacts_sync_mode_prompted is already true', async () => {
|
|
// Same staging, but prompted=true this time. Gate should be silent.
|
|
const gstackHome = fs.mkdtempSync(path.join(os.tmpdir(), 'privacy-gate-off-'));
|
|
const fakeBinDir = fs.mkdtempSync(path.join(os.tmpdir(), 'privacy-gate-off-bin-'));
|
|
|
|
fs.writeFileSync(
|
|
path.join(gstackHome, 'config.yaml'),
|
|
'artifacts_sync_mode: off\nartifacts_sync_mode_prompted: true\n',
|
|
{ mode: 0o600 }
|
|
);
|
|
|
|
fs.writeFileSync(
|
|
path.join(fakeBinDir, 'gbrain'),
|
|
'#!/bin/bash\necho \'{"status":"ok"}\'\nexit 0\n',
|
|
{ mode: 0o755 }
|
|
);
|
|
|
|
const askUserQuestions: Array<{ input: Record<string, unknown> }> = [];
|
|
const binary = resolveClaudeBinary();
|
|
|
|
// Ambient env mutations (see note on the first test).
|
|
const origGstackHome = process.env.GSTACK_HOME;
|
|
const origPath = process.env.PATH;
|
|
process.env.GSTACK_HOME = gstackHome;
|
|
process.env.PATH = `${fakeBinDir}:${process.env.PATH ?? '/usr/bin:/bin:/opt/homebrew/bin'}`;
|
|
|
|
try {
|
|
await runAgentSdkTest({
|
|
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
|
userPrompt:
|
|
'Run /learn with no arguments. Just report the learnings count.',
|
|
workingDirectory: gstackHome,
|
|
maxTurns: 4,
|
|
allowedTools: ['Read', 'Grep', 'Glob', 'Bash'],
|
|
...(binary ? { pathToClaudeCodeExecutable: binary } : {}),
|
|
canUseTool: async (toolName, input) => {
|
|
if (toolName === 'AskUserQuestion') {
|
|
askUserQuestions.push({ input });
|
|
// Pass through whatever the model asks; don't prefer anything.
|
|
const q = (input.questions as Array<{
|
|
question: string;
|
|
options: Array<{ label: string }>;
|
|
}>)[0];
|
|
return {
|
|
behavior: 'allow',
|
|
updatedInput: {
|
|
questions: input.questions,
|
|
answers: { [q.question]: q.options[0]!.label },
|
|
},
|
|
};
|
|
}
|
|
return passThroughNonAskUserQuestion(toolName, input);
|
|
},
|
|
});
|
|
|
|
// No AskUserQuestion should have matched the privacy gate's prose.
|
|
const privacyQuestions = askUserQuestions.filter((aq) => {
|
|
const qs = aq.input.questions as Array<{ question: string }>;
|
|
return qs.some(
|
|
(q) =>
|
|
/publish.*session memory|private github repo|gbrain indexes/i.test(q.question)
|
|
);
|
|
});
|
|
expect(privacyQuestions.length).toBe(0);
|
|
} finally {
|
|
if (origGstackHome === undefined) delete process.env.GSTACK_HOME;
|
|
else process.env.GSTACK_HOME = origGstackHome;
|
|
if (origPath === undefined) delete process.env.PATH;
|
|
else process.env.PATH = origPath;
|
|
fs.rmSync(gstackHome, { recursive: true, force: true });
|
|
fs.rmSync(fakeBinDir, { recursive: true, force: true });
|
|
}
|
|
}, 180_000);
|
|
});
|