By the end of this chapter you can inspect, debug, and audit what your workflows do — with gh aw logs, gh aw audit, run summaries, and OpenTelemetry — so you can trust a fleet you can't watch by hand.
This chapter's fixed target is gh aw v0.88.7. Building on the shared workflows in Chapter 11, we trace a run that went wrong from an overview table down to the failing step. The revised source and matching embedded workflow passed strict v0.88.7 compilation; the debugging scenario is illustrative, not a captured live-run transcript.
Everything in Part III assumes a fleet running unattended — agents triaging, reviewing, and opening PRs across many repos while you sleep. That only works if you can answer, after the fact: what did it do, why, what did it cost, and what did it touch?Observability is the precondition for trust. You can't govern — can't budget, can't secure, can't improve — what you can't see.
The compile model from Chapter 3 gives you a reviewable job plan; runtime records let you compare that plan with what happened. Prompts, recorded outputs, patches, and logs provide evidence for post-hoc analysis (v0.88.7 artifact reference). These are defined artifacts, available when produced and retained — not a complete snapshot of the runner's filesystem. Debugging means reading the record and recognizing its limits.
Inspectable orchestration is not deterministic judgment. Both the main agent and the default threat detector perform separate, probabilistic AI inference (tagged threat-detection reference). The compiled job ordering, permissions, and gates are inspectable, but neither AI judgment is deterministic or proof of safety. Inspect both the agent's reported outcome and the detector's verdict alongside the available artifacts. Detection evidence is also bounded: the retained verdict and summary do not include the deliberately omitted raw detector log.
Three CLI commands and one export give you a practical starting point: find anomalies, investigate the evidence, and bring run telemetry into your existing operations tools. The CLI recipes below were checked against v0.88.7 help and tagged sources, not exercised on live runs.
gh aw logs — the overview and the artifacts
It downloads and analyzes workflow logs and artifacts, then gives you an overview with duration, token usage, and cost information. In the inspected v0.88.7 help, the default download is still just the compact usage artifact, with a default limit of ten matching runs per workflow. Widen the evidence with --artifacts (tagged CLI reference; gh aw logs --help):
Fetch runs and choose how much to download; replace owner/repo with your repository
gh aw logs # overview: duration, tokens, cost
gh aw logs repo-assistant # just one workflow's runs
gh aw logs owner/repo/repo-assistant # a workflow in a remote repository
gh aw logs --artifacts all # all available artifact sets
gh aw logs --artifacts agent,firewall # only what you need
gh aw logs --help # current download and pruning controls
Bound the sample before downloading. v0.88.7 accepts multiple workflow targets, including cross-repository paths, and combines their results. --count applies per workflow: two workflows with --count 20 can return up to twenty matching runs each, not twenty in total. Date filters such as --start-date narrow the window but do not remove that count limit. Use --exclude-staged to omit runs that used staged safe outputs; it replaces the older --no-staged spelling.
v0.88.7 recipes: compare named workflows, or inspect a bounded sample from the last week
gh aw logs repo-assistant repo-assistant-observable --count 20
gh aw logs repo-assistant --start-date -1w --count 20 --exclude-staged
Choose comparable evidence when workflows use different sandbox runtimes or record different measurements. These selectors filter existing run records; they do not change or instrument the workflow:
Target CLI filters for a more focused investigation
Investigate…
Command recipe
runs using the Docker agent runtime
gh aw logs repo-assistant --runtime docker --count 10
runs with recorded evaluation results
gh aw logs repo-assistant --evals --count 10
runs with recorded deterministic grader results
gh aw logs repo-assistant --graders --count 10
The OTLP example in this chapter does not configure evaluations or graders. Those filters are useful for a fleet that already records such results; a metric or grader result is supporting evidence, not a substitute for checking the outcome or reviewing the work (tagged audit reference).
Bound the download footprint too. v0.88.7 resolves remote workflow names automatically and improves cache-size diagnostics and budget-driven pruning during concurrent downloads (v0.88.7 release notes). After choosing the runs and artifact sets, choose a storage policy or an API reserve. Preserve required incident evidence under your retention policy before enabling cleanup:
Download-control recipes — the storage recipe can delete cached run evidence
gh aw logs repo-assistant --count 20 --max-storage 10240 --prune-older-runs
gh aw logs repo-assistant --count 20 --max-github-api-rate-limit -2000 --timeout 30
--max-storage is a log-cache budget in MB; zero means unlimited. The first recipe sets 10,240 MB. Pruning first removes nonessential data from completed runs, preserving summaries and metadata where possible.
--prune-older-runs allows removal of the oldest completed runs if that selective cleanup still cannot meet the budget. Do not mistake selective preservation for guaranteed retention: this additional mode can remove the remaining run record.
--max-github-api-rate-limit -2000 reserves 2,000 requests from the GitHub core API allowance. A positive value instead sets a maximum used core-request count before waiting for reset. --timeout 30 sets a thirty-minute download timeout. These are client download controls, not AI-credit or token budgets.
Check gh aw logs --help for the full option definitions. Local cache cleanup does not extend GitHub's artifact retention, and a cached report cannot restore evidence that was never collected or is no longer available.
Depending on what the run produced and what you downloaded, the record can include agent-stdio.log, safe_output.jsonl (recorded agent output), aw-{branch}.patch (changes), workflow-logs/, and summary.json. Useful artifact sets include activation, agent, detection, firewall, github-api, mcp, usage.
All artifacts does not mean all files. v0.88.7 restricts agent artifact packaging to known files to reduce accidental data exposure. Claude debug logs also moved outside the agent data directory to avoid interfering with output packaging (release notes). Don't assume an arbitrary diagnostic file will be in the uploaded agent output, or interpret an absent file as proof that nothing happened.
Some raw diagnostics are deliberately omitted. The default external threat-detection path uploads detection_result.json and step-summary.md, not detection.log, because the raw log can contain sensitive content derived from the agent transcript (tagged detection artifact reference). Redaction and packaging limits reduce exposure; they do not guarantee that every sensitive value is removed. Control access to downloaded records and review them before sharing.
gh aw audit — the focused report
Where logs is broad, audit is deep. It downloads artifacts and logs, detects errors, analyzes MCP tool usage, and generates a focused report (tagged audit reference). Unlike the lightweight logs default, audit requests all available artifact sets for the selected run by default.
Use a full run URL so the repository context travels with the run ID. For a bare numeric ID, the inspected target help requires --repo owner/repo. Replace the sample ID, repository, and URL placeholders below with your own:
Investigate one run, or diff two
gh aw audit 1234567890 --repo owner/repo # bare ID with explicit repository context
gh aw audit <run-url> # detailed Markdown report for one run
gh aw audit <run-url>/job/<id> # a job URL — extracts the first failing step
gh aw audit <baseline-url> <comparison-url> # compare two runs (first = baseline)
Given a job URL without a step anchor, it extracts the first failing step's output. Its Firewall Analysis section connects to Chapter 7: it reports domains and allow/deny decisions found in the available firewall evidence. Use that report to form a diagnosis, then confirm it against the relevant logs.
For repeat investigations, gh aw logs --audit generates or reuses each cached run's audit.json from downloaded data (cached-audit change shipped before the target). Choose enough artifact sets for the question you are asking:
Create cached reports from selected evidence for up to five runs in a remote repository
gh aw logs owner/repo/repo-assistant --count 5 --artifacts agent,firewall --audit
Report generation uses the downloaded evidence; it does not make the overall logs invocation offline or API-free. A cached report can still have missing data. Widen the artifact selection or investigate the original run when the report cannot answer your question.
Run summaries and OpenTelemetry
Read the Markdown step summaries in the Actions UI alongside the workflow status from gh aw status. In v0.88.7, an agent calling report_incomplete makes the workflow's conclusion step fail rather than silently succeed; optional incomplete-work issue reporting can still proceed (incomplete-work failure change). That is an outcome signal, not necessarily an engine crash. A noop (“no work was needed”) and incomplete work (“the agent could not finish”) are different outcomes.
For centralized, cross-run visibility, observability.otlp configures trace export to an OpenTelemetry Protocol (OTLP) compatible backend (tagged frontmatter reference). With a working, reviewed runtime configuration, agent runs can appear in the same tracing tool as the rest of your systems.
That export is also a data-flow decision. v0.88.7 hardens OTLP handling against scheme-only authorization headers — for example, Bearer without a credential (release notes). This guard does not validate your collector or approve sending telemetry to it. The example below supplies Authorization directly from OTLP_TOKEN; its value must match the collector's required header format.
You can't read every run of a busy fleet — nor should you. The skill is knowing which runs earn a look. Let the cheap signals (the overview table, the safe-outputs boundary, the threat-detection gate) carry the routine cases, and spend attention where the signal says something's off.
Inspect closely when…
Trust the guardrails when…
a run failed, timed out, or reported incomplete work
it succeeded and produced expected safe outputs
tokens/cost spiked vs. the norm
cost is in the usual band
the firewall logged unexpected domains
egress stayed within the allowlist
you're rolling out a new or changed workflow
a stable workflow is running unchanged
When not to
Don't skip observability because “it's working.” A silent fleet is not a healthy fleet — it's an unmonitored one. Glance at gh aw logs regularly even when nothing's on fire.
Don't debug from the model's chat alone. Compare its narration with the patch, safe-output JSON, and firewall evidence. Read the record, not just the story.
Don't read a filtered sample as fleet-wide proof. Count, date, runtime, and result filters change what you see. No matching runs — or no recorded evaluation result — does not mean there were no failures.
Don't confuse missing evidence with a clean run. Check which artifacts were produced and selected, and whether redaction, retention, or local pruning limits what you can inspect. Keep required evidence before cleaning up a cache.
Don't treat observability as a substitute for the guardrails. Seeing a bad action after the fact is no help if it already shipped. Logs and audit complement safe outputs and review gates; they don't replace them.
Imagine the Repo Assistant's nightly run failed. Here's a three-step path from “something's wrong” to a supported diagnosis. The IDs, metrics, and failure below are illustrative; no workflow was run for this chapter update.
1. Get the overview. Start broad to find the bad run and its ID:
An illustrative overview surfaces the anomaly — not a measured CLI transcript
gh aw logs repo-assistant --start-date -1w --count 20 --exclude-staged
# RUN ID WORKFLOW STATUS DURATION TOKENS COST
# 1234567890 repo-assistant failure 4m12s 182,400 …
# 1234567889 repo-assistant success 0m48s 12,100 …
In this scenario, the failed run also burned roughly 15× the tokens of a healthy one — two signals pointing at the same run.
2. Audit that run. Open its focused report, then use a job URL if you need the first failing step's output. Substitute your repository and real run ID:
Audit a run URL with explicit repository context — placeholder URL, not an executed command
gh aw audit https://github.com/OWNER/REPO/actions/runs/1234567890
# Examine errors, MCP tool usage, safe outputs, and Firewall Analysis.
# Read the reported outcome too: incomplete work now fails the workflow.
Say the report points to the agent looping on a tool call to a domain the firewall denied. Confirm the repeated calls and timeout in the raw evidence; in this scenario, those retries explain both the failure and the token blow-up.
3. Confirm and fix. Inspect the raw files downloaded by the audit. If you also need that evidence across the workflow's recent runs, request their available artifact sets explicitly with the logs command below. A denial is not permission to widen the firewall: first establish whether the host is a legitimate dependency. If it is, review a narrow change to network.allowed (Chapter 7); otherwise fix the prompt or tool path and keep the denial. Recompile with strict mode:
Inspect the evidence, review the cause, then recompile
gh aw logs repo-assistant --start-date -1w --count 20 --exclude-staged --artifacts all
# Same filters as the overview, with wider artifact selection.
# Illustrative diagnosis: denied egress, retried to timeout.
# Review whether the dependency is legitimate before changing network.allowed.
gh aw compile --strict .github/workflows/repo-assistant.md
examples/ch12/repo-assistant-observable.md — complete workflow; strict v0.88.7 compilation PASS; runtime NOT RUN
---
on:
issues:
types: [opened]
workflow_dispatch:
permissions:
contents: read
issues: read
engine: copilot
network:
allowed:
- defaults
- github
safe-outputs:
add-comment:
max: 1
observability:
otlp:
endpoint: ${{ secrets.OTLP_ENDPOINT }}
headers:
Authorization: ${{ secrets.OTLP_TOKEN }}
---
# Repo Assistant — observable triage
You are the **Repo Assistant**. Triage the new issue with a single, concise
comment summarizing it and any missing information.
This example is about **operating** the workflow, not the triage itself. It
exports distributed traces to an OpenTelemetry (OTLP) backend via the
`observability:` block when the runtime is configured. Traces, token usage,
timing, and the records from `gh aw logs` and `gh aw audit` provide
complementary evidence about observable activity and reported outcomes,
bounded by collection, redaction, and retention.
Historical compilation evidence, not verification of this revision. The preflight recorded actual gh aw version v0.88.7. The then-current standalone source passed strict compilation with exit code 0 and a nonempty emitted lock. It also emitted one safe-update warning, including SECURITY REVIEW REQUIRED and the following secret references. The exact result is retained in content/research/updates/v0.88.7/preflight-verification.json.
Selected lines from the historical preflight compiler warning — secret names, not secret values
New restricted secret(s):
- OTLP_ENDPOINT
- OTLP_TOKEN
Those preflight fixtures had no approval manifest. That explains the warning; it does not waive the review. Before deployment, review why these credentials are needed, who controls the telemetry destination, what data will leave the workflow, and who can access or retain it. Keep credentials scoped to that intended use. Never add --approve merely to silence the warning, or weaken strict mode to avoid it.
Revised example: strict compilation PASS. The revised standalone source and matching embedded copy each passed with exit code 0 and a nonempty lock whose metadata confirms compiler_version: v0.88.7 and strict: true. Each compilation emitted one safe-update warning naming OTLP_ENDPOINT and OTLP_TOKEN; no approval was granted. Fresh final source and embedded evidence is recorded in content/research/updates/v0.88.7/verification.json and embedded-verification.json; the earlier pilot-revision report remains historical. This is compile-time technical evidence, not editorial acceptance or deployment approval; runtime remains NOT RUN.
Keep verification gates separate. The historical result above is from compile --strict. The separate --validate gate adds checks whose results depend on repository features, dependency resolution, and available tooling (target validation implementation). A PASS against a reference repository does not certify your deployment repository. Docker-backed checks and optional scanners were unavailable in the assessment environment; no PASS for those checks is claimed here. Retain each gate's exit status and diagnostics rather than collapsing everything into one “verified” label.
Runtime: NOT RUN. No engine or OTLP secret values were supplied in the historical preflight. A live deployment needs the Copilot credentials discussed in Chapter 5 and reviewed values for OTLP_ENDPOINT and OTLP_TOKEN. A compilation PASS does not validate endpoint reachability or authentication, and it does not approve secret exposure. Confirm trace delivery only in a separately authorized live test; combine those traces with Actions summaries and logs/audit rather than treating any one source as a complete record.
You can now see what your fleet does, and debug it when it misbehaves:
Observability is the precondition for trust — you can't govern what you can't see. Know which evidence is produced, packaged, and retained.
gh aw logs gives the overview + artifacts (duration, tokens, cost; --artifacts to download more). gh aw audit gives a focused report on tool/firewall use and, with a job URL, extracts the first failing step's output.
Bound the investigation. Counts apply per workflow, dates and other filters select a sample, and download budgets control the local cache and GitHub API usage — not model spend. Cached audit reports reuse evidence; they are not permanent, complete history.
Run step summaries, gh aw status, and OpenTelemetry (observability.otlp) round out the picture.
Inspect the runs that signal trouble (including incomplete outcomes, cost spikes, and denied egress). Investigate a denied dependency before widening access; never let observability replace the guardrails.
Compile PASS is not deployment approval. The revised OTLP source and matching embedded copy passed strict v0.88.7 compilation with a restricted-secret review warning; runtime remains NOT RUN.
What's next. Seeing cost is the first step; controlling it is the next. In Chapter 13: Governance & FinOps, we cap and meter agentic spend with max-ai-credits and set the org policy that keeps a fleet affordable and compliant.