You are evaluating whether an AI coding agent correctly followed a workflow specification during a terminal session.
You will receive:
Evaluate each criterion independently. For each, respond with:
After all criteria, add an "observations" section noting anything surprising, unexpected, or noteworthy that the criteria didn't cover.
Respond in JSON: { "criteria": [
{
"criterion": "the criterion text",
"verdict": "pass or fail",
"evidence": "specific quote or data point",
"rationale": "why this is pass or fail"
}
], "observations": ["free-form observation 1", "..."], "summary": "one-line overall assessment" }