Separate agent, detector, admission, and compute costs; apply scoped budgets and policy; and use forecasts without mistaking them for a complete bill cap.
By the end of this chapter you can cap, meter, and gate agentic work by choosing scoped budgets, model and capability policies, and dependency controls for a growing fleet.
This chapter targets gh aw v0.88.7, inspected on ; the separate APM integration targets APM 0.28.0, inspected on . Build on the evidence loop in Chapter 12 and governed reuse in Chapter 11. Compilation evidence here is not a live billing or policy-enforcement test.
Even ordinary CI has variable compute costs. Agentic work adds another variable: how much model inference each run uses to read, reason, and retry. A looping agent or an over-eager schedule can spend resources without producing useful work. FinOps therefore starts with both a budget and an outcome to measure, not just a cheaper model.
At one repo, this is a cost knob. Across an org, it becomes a policy surface: which workflows may run, which capabilities and dependencies they may use, what they may spend, and which model they select. Central governance makes those choices reviewable, but a central setting only governs the execution paths that actually consume it.
Keep three distinctions in mind. Admission decides whether to start work; budget enforcement limits a defined resource path after it starts. Defaults fill gaps; policy gates restrict supported behavior. Forecasts estimate future usage; neither a forecast nor a main-agent cap is a ceiling on the entire bill. Main inference, threat-detection inference, Actions compute, and execution time need separate accounting (Cost Management).
AI Credits (AIC) provide a common inference-cost metric, calculated from model pricing data. They are best-effort estimates, not a substitute for the provider's billing dashboard; Actions compute is billed separately (AIC reference).
Omitted-setting behavior in paired minimal compiler probes, inspected 2026-09-15; these are fallbacks, not measured usage
Control
v0.81.6
v0.88.7
Main-agent inference budget
1000 AIC
1000 AIC
Daily workflow admission threshold
5000 AIC
5000 AIC
Independent detector inference budget
400 AIC
400 AIC
Agentic step timeout
20 minutes
20 minutes, with runtime-variable fallback
Generated agent job timeout
No explicit value emitted in the probe
60 minutes
Generated detection job timeout
No explicit value emitted in the probe
10 minutes
max-ai-credits bounds the main agent's AWF-proxied inference, with steering messages at 80%, 90%, 95%, and 99% of its budget; integer values and K/M suffixes are supported (Frontmatter, inspected 2026-09-15). Detection has its own safe-outputs.threat-detection.max-ai-credits setting and fallback: lowering the main budget does not lower that separate allowance (Detection Budget).
Frontmatter excerpt from examples/ch13/repo-assistant-budgeted.md; the complete workflow is below
The daily setting implements admission, not reservation. Before admitting an applicable run, it looks back over the same workflow's previous 24 hours. If recorded usage already meets or exceeds the threshold, activation warns, attempts an issue report, and skips the agent. Two activations can read the same history and both pass; the check does not atomically reserve their future spend. Its history and artifact lookups also consume GitHub API requests (daily guardrail).
The documented bypass paths matter: the check is skipped for workflow_call, repository_dispatch, and workflow_dispatch carrying internal aw_context metadata. Do not multiply a daily threshold by a calendar period and call the result a guaranteed bill ceiling. Trigger filtering and Actions concurrency need their own design; see Chapter 4.
timeout-minutes limits the agentic execution step. It is not the generated agent job's timeout or the detection job's timeout: those cover every step of their respective jobs. The target's separate controls are jobs.agent.timeout-minutes and jobs.detection.timeout-minutes (timeout defaults and precedence). Nor is on.stop-after a runtime timer: it supplies an admission deadline, with stop-time preservation and explicit refresh covered in Chapter 4.
Token efficiency remains useful: narrow the task, context, and tool results before increasing the budget. Read-only permissions primarily constrain authority; they do not by themselves make inference cheaper. Security and FinOps reinforce one another, but solve different problems.
Org defaults, runtime gates, and repository strictness
To implement the default-versus-policy distinction, first ask when a setting is read. gh aw env manages GH_AW_DEFAULT_* Actions variables through a separate YAML file with default_-prefixed keys. It does not inject every variable into every local compiler or rewrite every deployed lock (Configuration Governance).
Three different resolution paths at v0.88.7
Surface
Enforcement or resolution point
What a central edit can change
Compiler-process defaults
GH_AW_DEFAULT_MAX_TURNS, some token guardrails, and GH_AW_DEFAULT_DETECTION_MODEL are read from the compiler's environment.
Supply them to the compile process, then regenerate and deploy locks. An Actions variable alone does not populate a developer's shell.
Emitted runtime defaults
Budget, model-fallback, and timeout paths can contain runtime expressions — for example, ${{ vars.GH_AW_DEFAULT_MAX_AI_CREDITS || '1000' }} when the main budget is omitted from both frontmatter and imports.
Later runs resolve visible variables where that expression was emitted. Explicit budget values in frontmatter or imports take precedence over budget defaults.
Runtime capability gates
Supported GH_AW_POLICY_* variables are checked by runtime components.
For example, GH_AW_POLICY_ALLOW_CREATE_PULL_REQUEST=false prevents the safe-outputs server from starting when PR creation is configured. It is not a numeric default.
For runtime defaults, applicable repository variables take precedence over organization variables, then enterprise variables, then the built-in fallback. Variable visibility and local overrides still matter. A supported runtime gate can affect later runs without recompilation when the deployed workflow includes that gate; it does not govern arbitrary manually written Actions steps. These paths are documented separately in Enterprise Environment Controls and Runtime Policy Variables.
New repository compile policy: put the following setting in .github/workflows/aw.json. Unlike an overridable default, it forces effective strict compilation for workflows processed by that compiler. The setting accepts true, not false (repository schema; strict-policy change).
Repository JSON configuration excerpt, not workflow frontmatter; fixture: examples/ch13/strict-policy/aw.json, paired with the strict-policy probe
{
"strict": true
}
A harmless workflow containing strict: false can still compile successfully: the emitted metadata says "strict": true. A source opt-out does not defeat this policy; the write-permission variant of the shared probe failed strict validation. Regenerate, review, and deploy locks to apply the policy. It does not repair old locks or automatically protect a manual workflow that bypasses the compiler. Keep the surrounding review and required-check controls from Chapter 7.
Model policy and uncertain forecasts
Model choice implements a preference; model permission policy implements a gate. At this target, engine.model takes precedence over top-level model; it has not been removed. Copilot's omitted-model fallback changed from claude-sonnet-4.6 to auto. That selector can vary the concrete model used, so an unchanged task is not necessarily a like-for-like cost comparison (fallback change; override precedence; Chapter 5).
Use models.allowed and models.blocked to express model access policy — the final key is blocked, not disallowed. The complete probe below permits the claude-* family while excluding claude-opus-*. This is not a credit allocation or proof that the provider will serve a permitted model (workflow schema). For dynamic selectors such as auto, the firewall accounts against the resolved concrete model rather than assigning a fictitious fixed price to “auto” (dynamic-model accounting).
CLI examples checked against v0.88.7 help on 2026-09-15 — sample window and count are choices, not measured results
gh aw models --json --refresh-observed=false
gh aw forecast repo-assistant-budgeted --period week --days 7 --sample 50 --json
models reports catalog data, aliases, and local observations; the flag above avoids its default observed-model log refresh. forecast uses historical completed runs to project usage and requires repository access and usable history. Its experimental label was removed in v0.84.0, but promotion is not an accuracy guarantee (target CLI reference).
No forecast or billing rate was measured for this chapter. Record the observation date, workflow/model configuration, sample coverage, and pricing source with any forecast you produce. Small samples, changed triggers, or auto routing can invalidate an extrapolation. Reconcile inference estimates with provider charges and Actions compute instead of presenting a catalog estimate as an invoice.
Governing the agent supply chain (APM)
Budgets and capability gates do not decide which instructions an agent should trust. Skills and prompts are an input supply chain: approved provenance and repeatable installation reduce risk, but do not prove that content is safe. APM is independently versioned, not another name for native gh-aw skills, plugins, or imports. Use the separately pinned bridge in Chapter 11; gh aw compile does not install its package graph or test APM policy.
The following corrections are ERRATA to the earlier governance promises, not new gh-aw security guarantees. For APM 0.28.0, inspected 2026-09-16:
Policy is preview and discovery is bounded. Configured extends chains merge tighten-only; do not assume automatic inheritance at every enterprise/org/repo level. A repo-local policy can explicitly use extends: org. Discovery derives the owner from the remote and tries .github-private before .github and other candidates. Fetch failures can warn and proceed rather than block; review cache and no-cache failure behavior before relying on enforcement (Policy Reference; discovery implementation).
Bounded constraints and locks are different.dependencies.require_pinned_constraint: true accepts bounded version constraints, including suitable ranges and version tags, not just exact SHAs. apm.lock.yaml separately records resolved commits and deployment integrity data. Confirm that your chosen install path consumes that lock; apm install --frozen rejects missing or out-of-sync locks, while apm audit checks installed integrity. A committed lock is not a security approval or proof that the bridge's isolated inline-package install replays it (constraint contract; lock specification).
Context preparation is not an agent sandbox.isolated: true is an apm-action input: it ignores the host apm.yml and clears known primitive directories under .github before preparing inline dependencies. The shared bridge uses it for packing. It is not a per-skill gh-aw import option, nor proof that repository instructions cannot influence the model (apm-action v1.10.0 contract).
Scanning has a defined threat scope. APM detects specified hidden Unicode, including bidi overrides and zero-width characters. Critical findings block deployment; some findings are only warnings. It explicitly does not detect homoglyph substitution or ordinary visible prompt injection. Keep content review and the Chapter 7 runtime defenses (APM security model).
APM apm-policy.yml excerpt, not gh-aw frontmatter or aw.json; illustrative package path, source examples/ch13/apm-policy.excerpt.yml
A nonempty allowlist already rejects unmatched sources. Do not add deny: ["*"] to mean “everything else”: * matches one path segment. Replacing it with ** would deny approved paths too, because matching denies take precedence. The fields above were checked against the versioned policy reference, and pure-source matcher diagnostics confirmed the unmatched-source rejection and the two wildcard cases (APM matcher). This is not a full policy installation test.
PROXY_REGISTRY_ONLY=1 restricts direct VCS fallback and lock replay; it does not make the whole workflow air-gapped. Policy-file fetching still uses GitHub APIs, and other workflow components have their own network paths (proxy coverage). No APM install or runtime policy-enforcement test was performed for this chapter.
Budgets trade off resource bounds against task completion: too tight and useful runs get cut off; too loose and a bad run consumes more. Tune from observed work, without turning a historical estimate into a guarantee.
Choose the control that matches the problem
For…
Use…
a frequent triage job
a narrow task, low main-agent budget, daily admission threshold, and deliberate trigger/concurrency settings
an occasional deep task
a reviewed larger allocation, separate detector/time limits, and evidence of useful completion
an org baseline
defaults at the right resolution phase, with visible exceptions and rollout checks
a capability or model you want to restrict
a supported runtime capability gate or models.allowed/models.blocked, not merely a preferred default
no effective workflow strictness opt-out
repository aw.json strict policy plus a protected compile/review/deploy path
approved skills and prompts
version-scoped APM policy, reviewed constraints and apm.lock.yaml, and evidence that the install consumed them
When not to
Don't disable budgets to “unblock” a workflow.max-ai-credits: -1 disables enforcement and steering. Diagnose looping, excessive context, or an under-sized allocation first; do not weaken strict mode to make a recipe compile.
Don't confuse noise control with a shared wallet. Cooldown and daily history checks do not reserve spend. Cooldown's history lookup can fail open; parallel and internally dispatched runs need separate review (Triggers).
Don't set defaults so tight that every repo overrides them. Inspect the exceptions and effective values. Conversely, do not describe an overridable default or a most-specific-wins runtime variable as an immutable organization policy.
Don't rely on budgets or package scans for complete security. Keep safe outputs, the firewall, permission boundaries, and human review. Approved content can still contain a bad instruction.
Don't equate compilation with billing authorization. Centralized Copilot CLI billing requires an explicit permissions.copilot-requests: write, organization policy allowing it, and the updated deployed lock. The compiler does not add that permission automatically. The unchanged permission configuration in this chapter's samples uses the COPILOT_GITHUB_TOKEN PAT path for a live run, not an interactive CLI OAuth token (Billing; Authentication).
The running example keeps its original 200/2000 settings. Those are illustrative task allocations, not prices or measured forecasts. Its credit, step-time, and calendar controls address different scopes; it does not claim to satisfy an unseen organization's policy.
examples/ch13/repo-assistant-budgeted.md — complete workflow, frontmatter and prompt kept in sync with the source; live run requires Copilot credentials
---
on:
issues:
types: [opened]
schedule: daily
workflow_dispatch:
stop-after: "+30d"
permissions:
contents: read
issues: read
engine: copilot
network:
allowed:
- defaults
- github
max-ai-credits: 200
max-daily-ai-credits: 2000
timeout-minutes: 10
safe-outputs:
add-comment:
max: 1
add-labels:
allowed: [bug, enhancement, question, documentation]
max: 1
---
# Repo Assistant — budgeted triage
You are the **Repo Assistant**. Triage the new issue with one concise comment and
at most one label. Keep it efficient: read only what you need, and don't spend
effort re-deriving context you already have.
On a scheduled or manual run without an issue target, report no work rather than
inventing a target.
This example demonstrates **scoped budgets and policy**. The main agent has a
`max-ai-credits: 200` budget; the rolling historical admission threshold is
`max-daily-ai-credits: 2000`. Neither is a total bill cap or an atomic reservation.
`timeout-minutes: 10` bounds the agentic step, not every job, and
`stop-after: "+30d"` supplies an admission deadline.
Threat detection and Actions compute are separate. Explicit frontmatter budgets
are not replaced by organization defaults. Compile-time policy and supported
runtime capability gates have different enforcement points.
The main agent gets 200 AIC, while the unchanged independent detector fallback remains 400 AIC. The 2000 AIC threshold is a rolling historical admission check, and the 10-minute timeout covers only the agentic step; the target's generated job timeouts remain separate. stop-after bounds admission after its deadline, not the total duration of every already-started run.
Live-run prerequisites: supply the Copilot PAT secret, enable Issues, and ensure the allowlisted labels exist. This recipe does not opt into label creation. The book repository had Issues disabled during inspection; compilation in a reference context is not evidence that issue outputs will deploy there (safe-output contract). The prompt asks scheduled or manual runs without an issue target to report no work rather than inventing one.
Compile in an isolated copy of the examples tree; these commands assume gh aw resolves to v0.88.7, not an older personal extension
gh aw version
gh aw compile examples/ch13/repo-assistant-budgeted.md --strict --no-check-update
If you use the isolated compiler from Chapter 2, invoke that executable instead. Compilation emits a lock without invoking the engine; it does not test credentials or demonstrate a successful triage run.
Check a model policy independently of the triage task
examples/ch13/models-policy.md — complete compile-only diagnostic adopted from the successful v0.88.7 strict probe, not a budgeted-triage variant
---
on:
workflow_dispatch:
model: auto
models:
allowed: ["claude-*"]
blocked: ["claude-opus-*"]
safe-outputs:
noop:
---
Report no work. This compile-only probe demonstrates model access policy, not a price cap.
This uses the default Copilot engine and leaves budgets at their fallback paths. Compile success establishes the accepted configuration shape, not availability of the model family, successful auto routing, or a price. Keep the budget controls when applying model policy to a real task.
Check effective repository strictness, not rejection of a word
examples/ch13/strict-policy/opt-out.md — complete compile-only diagnostic; use with the separate aw.json fixture above, not as a recommended opt-out
---
on:
workflow_dispatch:
strict: false
safe-outputs:
noop:
---
Report no work. This compile-only probe tests repository-level strict-mode enforcement.
In a scratch Git repository, place both fixtures under .github/workflows/, then run gh aw compile opt-out --no-check-update with the target compiler. The shared probe compiled without a CLI --strict override and emitted effective "strict": true metadata under v0.88.7. Checking only for a successful compile would miss the point: inspect both the compiler version and effective strictness. The final book verification gate uses CLI --strict for ordinary workflows, but stages adjacent aw.json and omits that flag for strict-policy/ fixtures. The final run also recorded a separate no-policy control; positive fixtures still require exact-version, effective-strict metadata and nonempty locks.
Diagnostic boundary: neither probe was run live. Declaring only safe-outputs.noop still emitted an automatic create_issue fallback in these probes; they are not recipes for disabling all writes. Provider credentials and repository features remain live-run prerequisites. The final authored source and embedded copies passed v0.88.7 compilation with exit code 0 and nonempty strict target locks; content/research/updates/v0.88.7/verification.json and embedded-verification.json record these technical results. Editorial review remains required before publication.
Roll out the central settings deliberately
After reviewing effective values and variable visibility, an administrator can use this sequence. It illustrates the defaults and runtime-gate paths; no organization settings were changed for this chapter.
Administrative CLI sequence — not executed; replace the organization placeholder and review deletions before applying
gh aw env get org-defaults.yml --scope org --org MY_ORG
# Edit the exported file; omitted/null keys are deletions.
gh aw env update org-defaults.yml --scope org --org MY_ORG --dry-run
# Apply only after review:
gh aw env update org-defaults.yml --scope org --org MY_ORG
gh variable set GH_AW_POLICY_ALLOW_CREATE_PULL_REQUEST --org MY_ORG --body "false"
You can now govern a fleet without mistaking one guardrail for the whole system:
Budget main inference, detection, compute, and time separately. A daily historical threshold is not an atomic reservation or a complete bill cap.
Distinguish compiler-process defaults, emitted runtime variables, and capability gates. A central edit only affects consumers of that setting.
Repository aw.json can enforce effective strict compilation, but locks must be regenerated and deployed through a governed path.
models.allowed/models.blocked restrict model access; auto can vary model choice. Forecasts remain estimates even after promotion out of experimental status.
APM policy, bounded constraints, consumed locks, and content scans support reviewed dependency distribution, not guaranteed instruction safety or a whole-workflow air gap.
What's next.Chapter 14: Fleets & Adoption takes the Repo Assistant from one repo to a multi-repo fleet, with deliberate consumer updates and a staged adoption playbook.