opencode plugin: LLM permission judge — asks get decided by a model with the agent's own prompt as context, every request logged
  • TypeScript 100%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
emidab b78a2813bf feat: workspace context for judge and worktree-aware guardrails
Feed the judge explicit workspace facts in every decision prompt: the
requesting session's workdir (from the shared session.get read) and the
project's git worktree directories (via ctx.worktree.list, structured API,
no subprocess or permission recursion). A new WORKSPACE CONTEXT section
gets a dedicated ~10% share of max_context_chars; missing data renders as
"(not available)", empty lists as "(none)".

Worktree lists are cached per projectID with a 60s TTL (successful lists
only) and deduplicated in flight, so concurrent judged asks of the same
project share one fetch instead of racing N concurrent git calls.

Guardrails updated to match: worktrees listed under WORKSPACE CONTEXT
count as inside the workspace (writes there are fine), and reads (never
writes) of other git repositories, package caches, and remote sources are
typically approvable — only writes are location-restricted.

110 tests, tsc clean.
2026-10-07 20:21:47 +00:00
.gitignore feat: add opencode ai-approve plugin 2026-10-07 10:13:09 +00:00
index.test.ts feat: workspace context for judge and worktree-aware guardrails 2026-10-07 20:21:47 +00:00
index.ts feat: workspace context for judge and worktree-aware guardrails 2026-10-07 20:21:47 +00:00
package.json feat: add opencode ai-approve plugin 2026-10-07 10:13:09 +00:00
README.md feat: workspace context for judge and worktree-aware guardrails 2026-10-07 20:21:47 +00:00
tsconfig.json feat: add opencode ai-approve plugin 2026-10-07 10:13:09 +00:00

opencode-ai-approve

opencode plugin that runs an LLM judge over permission asks: when opencode would stop and prompt you to approve a tool call, the judge model decides instead — allowing safe work to continue unattended, denying dangerous work outright, and falling back to a real ask whenever it cannot decide.

Mains and subagents stall on every permission prompt. This plugin keeps agents working on the long tail of safe asks (reads, builds, tests, lint) while still gating the dangerous ones — and it never risks a silent "yes" on something it failed to evaluate.

Works with opencode V2 (2.x) via the permission.evaluate hook: the hook fires for every configured allow AND ask decision, and its effect field is mutable — rewriting it changes the outcome. Explicit configured denies never reach the hook.

Runtime toggle

While opencode is running, the /ai-approval command flips the judge on or off in any session: bare /ai-approval toggles, /ai-approval on|off sets it explicitly, /ai-approval status reports it without flipping (on|off|status is case-insensitive; every invocation reports the resulting state as a synthetic session message). The state is per-runner and in-memory — it resets on restart to the enabled option, which sets the startup state (true by default, so judging starts on). While off, permission asks pass through exactly as without the plugin — nothing is judged and the JSONL log records nothing.

What it does

For every permission evaluation that arrives with effect ask (and is not excluded, and matches the optional rules allowlist), the plugin:

  1. Builds a judge prompt — instructions + hard guardrails, then the payload: the tool action and resources, the requesting agent's name, description, mode and configured permission rules, the agent's own system prompt under a clearly labelled REQUESTING AGENT'S SYSTEM PROMPT / AUTHORITY CONTEXT: section (or an explicit (not available) marker when agent info is unavailable), the tool input (event.metadata) if present, a WORKSPACE CONTEXT: section, and a truncated tail of recent session context. WORKSPACE CONTEXT: gives the judge the requesting session's working directory (from session.get) and the project's git worktree directories (from worktree.list, so the judge can verify whether a path is inside the workspace or one of its worktrees) — every missing fact (no session location, no projectID, a failing list, not a git project) is stated explicitly as (not available), never silently dropped. The worktree list is fetched per projectID within the pre-judge timeout race and cached in memory for 60 seconds; a failed or unavailable list is never cached and never breaks judging. The prompt embeds hard guardrails: deny writes outside the workspace (the project's own git worktrees listed in WORKSPACE CONTEXT count as inside the workspace — writes there are legitimate), deny deletions, privilege escalation (sudo, package-manager script phases), secret/env exfiltration, and calls to private or IP-literal hosts; allow safe reads, builds, tests and lint — including reads (never writes) outside the workspace such as other git repositories, package caches (node_modules, ~/.cache), or remote URL fetches. The ALLOW/DENY format reminder is always appended as a guaranteed final block — never relying on a custom prompt to state it — and the whole prompt respects max_context_chars — a hard cap, except that the guaranteed format reminder may push an oversized custom-template prompt over it (the instruction text is charged to the budget; the agent's system prompt gets a dedicated share of up to 40% of the cap — not leftovers — tool input and session tail up to 30% each, workspace context ~10%; the reminder and the agent's system prompt survive truncation at the cap).
  2. Calls the judge model session-less via ctx.generate.text (no session, no tools — so the judge cannot recurse into the permission system), bounded by timeout_ms; the timeout races the judge call so the permission flow stays bounded, and in-flight cancellation is not delivered through the current V2 promise adapter (@opencode/plugin's promise adapter drops the abort signal).
  3. Applies the verdict — the verdict token is only trusted in a small leading window: it must occur within the first ~32 characters of the first non-empty line, case-sensitively, with backtick-quoted code spans stripped out (so a judged command echoing echo ALLOW can never flip the decision); anything else is unparseable — including a reply whose preamble contains a negator ("not", "no", "cannot", "never", "shouldn't", "nope", "non-", ...), which flips the token's meaning — and falls back to asking. Judged as allow the ask becomes an auto-allow, deny becomes a deny, with the model's one-line reason attached to the permission message.
  4. Logs the decision — one JSONL line per judged request (including failures), appended without ever blocking or failing the permission flow. The record is fully self-contained: it embeds the verbatim judge reply and the complete assembled judge prompt, so the JSONL alone is the whole audit trail. (JSONL lines are therefore long — that is accepted by design.)

The three locked decisions

  1. On any decision failure — generate error, timeout, unparseable verdict, missing model — the event stays (or returns to) ask with a message explaining the failure. Never silently allow, never deny on failure.
  2. Only asks are judged. Events whose configured effect is allow pass through untouched (configured deny never reaches the hook at all).
  3. Scope — everything that would ask is judged, unless an optional rules allowlist narrows it. If rules is set, only matching action/resource pairs are judged; everything else passes through as a normal ask.

Install

Add the plugin to your opencode config (~/.config/opencode/opencode.jsonc, or the project's opencode.json):

{
  "plugins": [
    {
      "package": "/path/to/opencode-ai-approve",
      "options": {
        // REQUIRED: the judge model as "providerID/modelID".
        // Without it every judged ask falls back to a normal ask.
        "model": "anthropic/claude-sonnet-4-5",
        "timeout_ms": 30000,
        "max_context_chars": 8000,
        // "rules": [{ "action": "bash", "resource": "git *" }],
        "exclude": ["sudo *"],
        "redact": true,
        // "enabled": true, — startup state of the judge; /ai-approval toggles at runtime (per-runner).
        // Custom judging prompt: full template takes precedence over prompt_file.
        // "{data}" inserts the payload; without it the payload is appended.
        // "prompt": "You are the permission gate for $ORG. {data} Apply company policy.",
        // "prompt_file": "~/prompts/judge.md",
        // Custom verdict tokens: first token = allowing verdict, second = denying.
        // "format": ["APPROVED", "BLOCKED"]
      }
    }
  ]
}

The judge model should be fast and cheap — it runs on every non-excluded permission ask.

Options

Option Default Meaning
model — (required) Judge model as providerID/modelID. Absent or invalid → every judged ask falls back to ask.
enabled true Startup state of the judge (Boolean-coerced; invalid values fall back to the default). /ai-approval flips it at runtime — per-runner, in-memory, reset on restart to this value.
timeout_ms 30000 Hard bound for the judge request (race; abort is not delivered in flight).
max_context_chars 8000 Hard cap on the judge prompt length — except a guaranteed payload section may push an oversized custom-template prompt over it.
log_path ~/.local/share/opencode/ai-approve.jsonl Decision log (~/ is expanded). The directory is created lazily.
rules — (judge all asks) Allowlist of { action, resource }; only matching asks are judged. Both fields support */? wildcards.
exclude [] Resource wildcard patterns that are never judged — they pass through as normal asks.
redact true Strip obvious secrets (KEY=… assignments, Authorization: headers) from logged resources, from the resources and tool input sent to the judge, and from the tool input in the log. Fidelity tradeoff: the judge itself sees these redacted forms — when redact: true the whole judge payload (resources, tool input, key-redacted TOOL INPUT metadata) is the log's view, not the raw command — and value redaction can swallow content embedded inside a quoted secret value (PASSWORD="$(rm -rf /)" → PASSWORD=[redacted]). Redaction is idempotent: re-redacting an already-redacted string changes nothing.
prompt built-in Full custom judging prompt template. If it contains {data}, the payload (action, resources, agent authority context, tool input, workspace context, session tail) is inserted there — every {data} occurrence is replaced by the payload, and the context budget accounts for one copy (the cap safety net trims the rest). Otherwise the payload is appended after the template. A template too large for max_context_chars warns at setup, and a minimal ACTION/resources section is always guaranteed regardless. The ALLOW/DENY format reminder is appended separately and unconditionally — a comment-prompt cannot remove it. Takes precedence over prompt_file.
prompt_file built-in Path to a custom judging prompt template in a file (~/ is expanded). Read once at plugin setup and cached; a read failure (or empty file) warns and falls back to the built-in template. Same {data} behaviour as prompt.
format ["ALLOW", "DENY"] Custom verdict tokens as [allowToken, denyToken], e.g. ["APPROVED", "BLOCKED"]. The appended format reminder quotes the tokens, and the verdict parser accepts them (instead of ALLOW/DENY) — still only within the trusted leading window. Tokens are uppercased before use — two tokens that collide after casing (e.g. ["ALLOW", "allow"]) are invalid and fall back to ALLOW/DENY with a warning.

Wildcard semantics (same family as permission rules): * matches any run of characters, ? exactly one, case-sensitively; a pattern without wildcards also matches any resource it is a case-sensitive prefix of. Empty/whitespace patterns in exclude are ignored with a warning; rules with empty action/resource are dropped silently, and a warning fires only when the rules array yields zero usable entries; an invalid format option falls back to ALLOW/DENY with a warning.

Log

One JSONL line per judged request, including failures:

{"ts":"2026-10-07T05:00:00.000Z","sessionID":"ses_…","agent":"build","action":"bash","resources":["git push"],"decision":"allow","reason":"safe push to a feature branch","metadata":{"command":"git push origin main","API_TOKEN":"[redacted]"},"model":"anthropic/claude-sonnet-4-5","latency_ms":1841,"source":{"type":"tool","messageID":"msg_…","id":"tool_…"},"prompt":"You are the permission judge for an opencode session. …\n\nACTION: bash\nRESOURCES: git push\n…","reply":"ALLOW\nsafe push to a feature branch","plugin":"ai-approve","v":1}

decision is allow, deny, or fallback-ask (the judge failed and the event stayed an ask). Two fields make the record fully self-contained, so the JSONL alone is the complete audit trail:

  • reply — the judge's verbatim reply, trimmed. Always present; "" when no reply was received (error, timeout, missing model).
  • prompt — the full assembled judge prompt, exactly as sent to the judge: instructions incl. the format reminder, action/resources, the agent's authority context, redacted tool input, the workspace context, and the session tail. With redact: true (the default) this embedded prompt is already redacted and is stable under re-redaction, but redaction is lossy — secret-value redaction can swallow content embedded inside quoted secret values, so the raw command may not be recoverable from the log.

Because the prompt is embedded verbatim, JSONL lines are long — that is accepted by design; the record is judged-prompt-complete rather than terse. Log writes are chained through a promise queue, never block the decision, and never fail the permission flow — write errors are swallowed with a console.warn.

Failure policy

  • generate error / timeout / unparseable verdict / missing model → effect = "ask", message explains the failure, logged as fallback-ask.
  • Excluded and rule-skipped asks are not judged and not logged — they behave exactly as without the plugin.
  • Each stage of the hook path (the pre-judge reads — agent info, workspace context, session context — and the judge call) is individually bounded by timeout_ms, so no hung read or hung generate can stall the permission flow (worst case ~2× timeout_ms before the ask proceeds).
  • The verdict token is trusted only in the leading window of the judge reply — within the first ~32 characters of the first non-empty line, case-sensitively; a judged resource echoing ALLOW cannot flip the decision, and a negator in the preamble ("I can not ALLOW...") makes the verdict unparseable.
  • A hung model cannot prevent the permission ask from proceeding; the timeout races the judge call so the flow stays bounded (in-flight cancellation is not delivered through the current V2 promise adapter).

Development

bun install
bun test
bun node_modules/typescript/bin/tsc --noEmit --project tsconfig.json

Note: under this proot environment bun install may drop package.json files during extraction — if typechecking complains about missing modules, copy a known-good node_modules (e.g. from the opencode-session-restart repo) over this one.