By the end of this chapter you can close the repository's quality loop with four more patterns — Review, Testing, CI-Doctor, and Refactoring — while keeping humans firmly on the merge decision. This is the second half of the Continuous-X library and the close of Part II.
This chapter targets gh aw v0.88.7 (Public Preview), inspected for this update on 2026-09-15. The Repo Assistant graduates from tending the issue tracker to helping tend the code.
A repository's quality has a loop: code is proposed (a PR), reviewed, tested, merged, and — when something slips — fixed. Traditional CI automates the deterministic checks in that loop: does it compile, do the tests pass, does the linter approve. But the judgement steps — is this a good change? is this test worth adding? why did CI actually break? — still wait on a human.
Continuous-X patterns fill exactly those judgement gaps. Where Chapter 9 kept the inbox honest, these keep the codebase honest — each one a mini-product owning one link in the quality loop.
Our policy: humans keep the merge
For this pattern library, the agent proposes; a human decides. The review agent comments rather than approving; the test-improver proposes a draft PR. This is a deliberate human-in-the-loop policy, not a product prohibition: gh-aw also offers opt-in merge capabilities, including an experimental merge safe output (v0.88.7 PR outputs). We do not enable those capabilities in these recipes.
The Chapter 6 boundary constrains which requested effects can be applied. It does not make an allowed comment accurate or an allowed code change correct. Use COMMENT-only reviews and draft PRs alongside repository review rules and human judgment, rather than treating either setting as a complete safety guarantee.
Choose the work, then constrain the effects
Admission decides which events merit agent work; enforcement bounds what the resulting proposal can change. Filtering redundant PR events can reduce repeated reviews, but it is not a concurrency lock or a spending cap. Likewise, asking for “tests only” describes intent; enforcing changed-file scope requires a separate control. These distinctions connect the clock from Chapter 4 to the mediated writes from Chapter 6.
Four patterns, each triggered by a different moment in the quality loop — and, under our policy, each writing through a safe output without delegating the merge.
Pattern
Trigger
Writes via
Review
pull_request
submit-pull-request-review (COMMENT only)
Testing
schedule
create-pull-request (draft)
CI-Doctor
workflow_run (CI failed)
add-comment / create-issue
Refactoring
schedule or command
create-pull-request (draft)
Review: comment, rather than approve
The Review pattern reads a PR diff and leaves inline feedback. Keep allowed-events: [COMMENT] explicit under safe-outputs.submit-pull-request-review: the handler rejects APPROVE and REQUEST_CHANGES decisions for this output. Omitting the allowlist permits all three decisions; COMMENT-only is a recommended configuration, not the implicit default (v0.88.7 review outputs). This enforces a narrow review contract, not a promise that other required checks can never block a merge.
Review admission: top of the stack by default
A stack is a chain of PRs in which each successive PR targets the previous one's branch. In v0.88.7, on.pull_request.max-stack defaults to 1: only the top-most PR is admitted by this filter; lower layers are skipped. The same default applies to pull_request_review. Ordinary, non-stacked PRs are unaffected by this setting. The unchanged review recipe below therefore has different admission behavior for stacks without gaining a new frontmatter line (tagged trigger reference).
Leave the default when you want to avoid repeated reviews of related diffs. If every stack layer needs its own review, explicitly set max-stack: -1 under on.pull_request to disable stack filtering. That chooses broader review coverage at the possible cost of more agent runs, AI-credit spend, and comment noise; it is not a measured savings comparison. As with the admission principle, keep concurrency and budgets separate. Existing role and fork guards still apply; disabling the stack filter does not bypass them.
CI-Doctor: react to the failure
CI-Doctor applies the same admission idea to the workflow_run trigger from Chapter 4. A named CI workflow completes; gh-aw's conclusion filter admits agent work only for failure. The compiler turns conclusion into a guarded if: condition — it is not a native Actions event filter, nor does it mean no Actions jobs can start for other conclusions (v0.88.7 conclusion filtering).
Trigger excerpt from examples/ch10/workflow-run-conclusion.md — a compile-only fixture for CI-failure admission, not a complete CI-Doctor recipe
on:
workflow_run:
workflows: [CI]
types: [completed]
branches: [main]
conclusion: [failure]
The compiler also adds repository-ID and fork checks for workflow_run; the explicit branch filter further narrows admission. These controls do not make CI logs trusted instructions. A complete doctor would still need reviewed log-reading tools and a bounded diagnosis output such as add-comment or create-issue.
Testing & Refactoring: propose a diff
Both patterns ask for focused work and propose it through create-pull-request with draft: true. Testing asks for new tests against current behavior; Refactoring asks for a small, behavior-preserving cleanup. The draft setting is an enforced creation policy that the agent cannot override, but it neither proves the diff correct nor governs every later merge action (tagged PR-creation reference).
A tests-only or docs-only prompt is not mechanical file-scope enforcement. The retained test-improver has no safe-outputs.create-pull-request.allowed-files allowlist, so its prompt does not prevent production files from appearing in a proposal. For enforced changed-file scope, configure that exclusive allowlist for your repository's test paths and retain an appropriate protected-files policy. Those are separate apply-time controls, not restrictions on every local edit or proof of behavior preservation. Default protected-file handling is not a test-directory allowlist (allowed files; protected-file policies).
Quality patterns pay off when they act as a tireless first pass — catching the obvious before a human spends attention, never replacing the human's final say. The line to hold: automate the noticing and the drafting; reserve the deciding.
For these patterns, the agent may…
Humans keep by policy…
comment on a PR, flag risks
approve / request changes / merge
open a draft test or refactor PR
review and merge that PR
diagnose a CI failure, file an issue
decide the fix and ship it
When not to
Don't accidentally make model judgment a merge gate. A REQUEST_CHANGES review can leave a persistent blocking state. Keep allowed-events: [COMMENT] for this non-gating review policy.
Don't assume every PR in a stack gets reviewed. Choose the stack admission policy before rollout; a skipped lower layer is not necessarily a broken review agent.
Don't trust “tests only” as enforcement. Inspect the complete diff. A PR that changes production code to make a new test pass violates this recipe's intent; configure and verify file-scope controls when that boundary must be mechanical.
Don't delegate the merge in this library. Retain human review and repository protections; do not enable merge outputs or another auto-merge automation for these recipes. A draft is one constraint, not the entire policy.
Don't run a refactoring agent on a repo without good tests. Tests provide regression evidence, not proof that all behavior is preserved. Ship Testing before Refactoring.
Two complementary quality agents: one reacts to admitted PR events, the other looks for a useful test addition each day. Their source recipes are retained unchanged and passed the v0.88.7 strict-compilation preflight. Below are the complete Markdown workflows, including their prompts — not abbreviated frontmatter presented as full recipes.
examples/ch10/continuous-review.md — complete COMMENT-only review workflow; v0.88.7 strict preflight PASS, live run not performed
---
on:
pull_request:
types: [opened, synchronize]
permissions:
contents: read
pull-requests: read
engine: copilot
network:
allowed:
- defaults
- github
tools:
github:
toolsets: [pull_requests]
safe-outputs:
create-pull-request-review-comment:
max: 10
submit-pull-request-review:
allowed-events: [COMMENT]
max: 1
---
# Continuous Review
You are the Repo Assistant's **review** agent. A pull request was opened or
updated. Give it a focused, constructive review.
1. Read the diff and the surrounding code for context.
2. Leave inline review comments on specific lines where you see real problems:
likely bugs, missing edge cases, unclear names, or missing tests. Be specific
and kind. Skip style nits a linter already covers.
3. Submit a single **COMMENT** review summarizing what you found. Never approve or
request changes — a human decides the merge.
If the PR looks good, submit a short COMMENT review saying so rather than
inventing problems.
examples/ch10/daily-test-improver.md — complete draft-PR workflow with a tests-only prompt, not a file allowlist; v0.88.7 strict preflight PASS with a schedule-context warning, live run not performed
---
on:
schedule: daily
workflow_dispatch:
permissions:
contents: read
engine: copilot
network:
allowed:
- defaults
- github
- node
tools:
github:
toolsets: [repos]
bash: ["npm ci", "npm test", "npx jest", "npx vitest run"]
edit:
safe-outputs:
create-pull-request:
title-prefix: "[tests] "
labels: [tests, automated]
draft: true
---
# Daily Test Improver
You are the Repo Assistant's **test-improver** agent, running once a day.
1. Run the existing test suite and inspect coverage to find one under-tested area
of the code that matters (core logic, a bug-prone module, an untested branch).
2. Write **new tests only** — do not change production code. Make them pass
against the current behavior.
3. Open one **draft** pull request titled `[tests] ...` adding those tests, with a
body explaining what you covered and why it matters.
Keep the change small and focused: one area, a handful of solid tests. If the
suite is already well covered, report no action instead of padding it.
The review agent's repository permissions are read-only; its two declared safe outputs permit up to ten inline comments and one COMMENT review. The test-improver gets the supported Bash allowlist and edit for local work, then proposes a draft through a separate write-capable safe-output job. Read these as bounded output contracts, not evidence that the model followed every prompt instruction.
What compilation verified
The 2026-09-15 preflight emitted locks for both unchanged recipes using v0.88.7 and --strict. Review had no compiler warnings; the test-improver warned that fuzzy scheduling lacked repository context in the isolated fixture. Both also emitted an informational Copilot billing tip. The separate CI-trigger probe passed target strict compilation with --validate; that evidence covers the diagnostic fixture, not a live diagnosis.
Reproduce strict compilation in a scratch copy with the v0.88.7 compiler selected — commands, not a runtime transcript
gh aw version
gh aw compile --strict examples/ch10/continuous-review.md
gh aw compile --strict examples/ch10/daily-test-improver.md
gh aw compile --strict examples/ch10/workflow-run-conclusion.md
Check that the version command reports v0.88.7 before compiling; see Chapter 2 for compiler setup. Compilation writes generated artifacts and may resolve dependencies, but it does not invoke the agent or require its credential. No test suite, model, runtime container, or security scanner was exercised for these checks; they produced no review, PR, coverage improvement, or measured cost.
Quality automation fills the judgement gaps CI can't: is this change good, is this test worth adding, why did CI break.
Four patterns — Review (PR, comment-only), Testing (scheduled draft PR), CI-Doctor (workflow_run on failure), Refactoring (scheduled draft PR).
Humans keep the merge is our policy. Explicit allowed-events: [COMMENT] and draft: true enforce narrower output constraints; repository review rules and human judgment complete the policy.
Review defaults to top-of-stack admission in v0.88.7; ordinary non-stacked PRs are unaffected. Choosing every layer means accepting potentially more work, not gaining a budget guarantee.
Tests-only intent needs file controls when enforcement matters. Keep Copilot's supported Bash restriction rather than swapping to an engine that cannot enforce it.
Automate the noticing and drafting; reserve the deciding. Ship Testing before Refactoring, and distinguish a compile PASS from evidence that tests improved or a PR succeeded.
What's next. You now have a shelf of patterns — and you're about to notice how much they repeat. Part III scales from one repo to an org. Chapter 11: Reuse & Memory factors the shared parts into imported components and gives the Repo Assistant memory that persists across runs.