Spec: Prompt exception requests
Status: APPROVED 2026-10-02 (user: “approved”). Tasks 1 to 5 built; task 6 open. Builds on:
docs/spec-hitl-breakglass.md(signed approval grants and the approval queue, shipped), the browser extension (examples/browser-extension),/guardrails/input.
0. The two cases, and why they need different answers
A blocked prompt is one of two things:
- The guardrail was wrong (a false positive). Nobody should have to wait for a human. A second look by a larger model can fix it in seconds.
- The guardrail was right, but the user is authorised (sharing a margin with a partner under NDA). That is a business decision. A model must not make it; a person must.
So the flow has tiers. This spec builds tier 3 first, because it covers both cases and reuses the most existing code, then tier 2.
| Tier | Decides | Wait | For |
|---|---|---|---|
| 1. Give a reason | the user | none | exists on the device agent (“justify”); unchanged here |
| 2. Second opinion | a larger model | seconds | blocks that came from a model’s judgement (task 6, opt-in) |
| 3. Exception request | an admin | minutes to hours | everything appealable (tasks 2 to 5) |
| 4. Not appealable | nobody at request time | n/a | guardrails the tenant marks as hard rules |
1. Problem & outcome
Today: the extension shows “Blocked by Shield: custom_policy_input” and the user is stuck. They cannot tell what was wrong, cannot ask for a review, and the admin never learns the policy misfired. The usual result is the user moving to a personal device.
Outcome
- The block banner says which policy blocked the prompt and why, and has a Request exception button.
- The user adds a reason and submits. The request appears in the portal with the prompt, the reason, the user, the device and the policy that blocked it. Admins are notified by webhook.
- An admin approves or denies. Approval issues a signed, single-use grant bound to that exact prompt, user and destination.
- The extension shows “Approved. Send again.” The resend carries the grant, and Shield lets that prompt through once.
- Every request and decision is audited, and each policy shows how often its blocks were overturned.
Non-goals (v1)
- File attachments. Prompts only.
- The device agent’s local decisions. Its prompts never leave the laptop, and it already has “justify”. Exception requests are for the path where the extension or an API client screens against Shield.
- Standing exemptions (“this user is exempt”). Every approval is for one prompt, once.
- “Approve and change the policy” in one click. An approver who thinks the policy is wrong marks the request as a false positive; editing the policy stays a separate, deliberate act.
- Tool calls. Those already have approvals (
docs/hitl-approvals.md).
2. Plane & latency contract
- Data plane: create and poll a request (
/v1/shield/exceptions), and honour a grant on/guardrails/input. - Admin + data plane: the review queue (
/v1/tenant/me/exceptions), like the other tenant APIs. - Guard path:
/guardrails/inputis touched, minimally.- A request without the grant header pays one header lookup. Nothing else changes.
- A request with the header pays the grant check only when the pipeline’s verdict is a block: an Ed25519 signature check, then about six store operations (settings, two revocation checks, the request, the single-use marker, the request’s new status). The pipeline always runs in full first; a grant never skips screening.
- Budget: under 1 ms added without a grant. With one, a handful of store round trips, once per approved request.
- Creating a request re-screens the prompt once (§4.1). That is a pipeline run on a request the user is waiting on anyway, not on the inline guard call, and it is rate-limited.
3. Data model
| Redis key | Value | TTL |
|---|---|---|
prompt_exc:{tenant}:{request_id} |
the request (below) | request_ttl_s + 7 days |
prompt_exc_idx:{tenant} |
sorted set of request_id by created_at |
trimmed with the requests |
prompt_exc_user:{tenant}:{user_hash} |
this user’s open request ids, for the per-user limit and for finding a duplicate | 7 days |
prompt_exc_fp:{tenant} |
one hash of counters per guardrail and policy: requested, approved, false_positive |
none |
prompt_exc_settings:{tenant} |
the tenant’s settings | none |
prompt_exc_block:{tenant}:{user_hash}:{destination_hash}:{sha256} |
{at, blocked_by}: what blocked this prompt, written by /guardrails/input in the background |
1 hour |
Request record:
request_id, tenant_id, status (pending | approved | denied | expired | used),
created_at, expires_at, decided_at,
user_id, device_id, destination,
prompt_sha256, prompt (up to 4000 characters), prompt_len,
reason (up to 500 characters),
blocked_by: [{guardrail, policy, message}], # from the server's own re-screen
decision: {approver, method, reason, false_positive},
grant_id
- The prompt text is stored. A reviewer cannot judge a prompt they cannot read. It is kept until the request’s TTL ends, then gone. The extension tells the user before they submit that reviewers will see the prompt.
- Tenant scoping: keys are prefixed by the tenant resolved from the API key. A request id from another tenant is a 404.
- Tenant settings have their own key and their own routes
(
GETandPUT /v1/tenant/me/exceptions/settings), not the agentic control-plane config: aPUTthere replaces every section it is not given, so a settings change here could have reset a tenant’s tool approval rules.
{"enabled": false, "non_appealable": [], "request_ttl_s": 86400,
"grant_ttl_s": 900, "max_pending_per_user": 3, "auto_review": false}
A request’s status is decided by the deadline stored in the request, never by a Redis expiry.
4. API / interface
4.1 Request an exception (data plane)
POST /v1/shield/exceptions, tenant key, X-Agent-Key (user) and
X-Device-Id as the extension already sends.
{"prompt": "...", "destination": "chatgpt.com", "reason": "Partner is under NDA"}
What blocked the prompt is always Shield’s finding, never the caller’s. When
/guardrails/input blocks a prompt it records, in the background, what blocked
it for that tenant, user, destination and prompt hash, for an hour
(prompt_exc_block:*). A request uses that record. Only without one does Shield
screen the prompt again. This matters for policies judged by a model: they can
block a prompt and pass the same prompt a moment later, and a request that
relied on a second screen was refused as “no longer blocked” (found in the
first production test, 2026-10-02).
| Result | Status |
|---|---|
| Created | 201 {request_id, status: "pending", expires_at, blocked_by} |
| Same user, same prompt, already pending | 200 with the existing request |
| The prompt is not blocked now | 409 not_blocked (the policy changed; just resend) |
| Blocked by a non-appealable guardrail | 403 not_appealable, naming it |
| Too many pending requests for this user | 429 |
| Feature off for the tenant, or no signing key | 404 / 503 |
4.2 Poll (data plane)
GET /v1/shield/exceptions/{request_id} returns status and, once approved,
grant (the signed token) and approver. Only the same tenant key with the
same X-Agent-Key may read it.
4.3 Review (both planes)
GET /v1/tenant/me/exceptions?status=pendingPOST /v1/tenant/me/exceptions/{id}/approve{reason, false_positive}POST /v1/tenant/me/exceptions/{id}/deny{reason}
Approval records the decision. The grant is minted by the data plane when the
requester polls an approved request, so its short lifetime (grant_ttl_s)
starts when they are there to use it, not when the reviewer clicked. It uses
the existing core.approvals.mint_grant: tool = "prompt_exception",
resource = "prompt:<sha256>@<destination>", agent_id = <user_id>, the
approver’s identity, and a hash of the guardrails it may waive. Each poll mints
a fresh grant; the request can still be redeemed only once, because redeeming
claims a per-request marker atomically. Writes go through the registry write
gate; the approver must not be the requester; the first decision stands (a
second gets 409).
4.4 Using the grant (guard path)
The resend adds X-Shield-Exception-Grant: <token> to /guardrails/input.
After the pipeline has run, and only if its verdict is a block:
- Verify the signature, audience, expiry and tenant.
- Check the grant’s resource equals the SHA-256 of this exact message and
this destination, and its
agent_idequals this caller’s user id. - Check every guardrail that failed now is in the request’s
blocked_byand none is non-appealable. - Claim the request’s single-use marker.
Nothing is spent until every check has passed, so a resend that does not match
leaves the approval usable. If all four pass, the response is safe: true, action: "pass", with the
failed results kept and marked exception_granted, an exception object
naming the request and approver, and the request becomes used. If any fails,
the block stands and the response carries exception_error with the reason
(grant_mismatch, grant_expired, grant_used, grant_invalid,
new_violation, not_appealable, exceptions_disabled,
store_unavailable).
The prompt hash is over the text after Unicode NFC normalisation and trimming outer whitespace, so the user must resend the same prompt.
4.5 Notification
A webhook event exception_requested (and exception_decided) through the
existing webhook delivery, with the request id, user, policy and a link to the
portal. Never the prompt text.
4.6 Extension
- Banner: the policy’s name and message instead of the guardrail’s internal name.
- Request exception opens a small form: a reason, and the notice that reviewers will see the prompt.
- A pending request is remembered per tab and polled every 30 seconds while the tab is open. The banner then shows approved, denied (with the reviewer’s reason) or expired.
- On approval, the next send of the same prompt attaches the grant.
5. Security & backward compatibility
- Opt-in.
exceptions.enableddefaults to false. With it off, no route accepts requests and the grant header is ignored. No existing default changes. - The grant cannot be forged or reused. It is the same signed, single-use, short-lived token tool approvals use, with its own resource namespace. A grant for one prompt does nothing for another, for another user, or for another destination.
- A grant never widens. If the same prompt now also fails a guardrail it did not fail when requested, the block stands.
- Hard rules stay hard. Guardrails in
non_appealablecannot be requested or waived. - The reviewer model (tier 2) cannot be argued with. It sees the policy text and the prompt. The user’s reason is never given to it.
- Abuse. Per-user pending limit, request expiry, requester cannot approve, and every request and decision is in the admin audit log.
- What a malicious caller can do: with a tenant key, create requests up to the rate limit for prompts that really are blocked. Nothing is released without an approver.
6. Packaging & deploy
- New modules
core/prompt_exceptions.py(store, hashing, grant check),api/routes_exceptions.py(ask and poll, data plane only, because asking runs the screening pipeline) andapi/routes_exception_review.py(settings and review, both planes).admin_app.pymounts the review module, so it andcore/prompt_exceptions.pyare inDockerfile.admin’s COPY list; the two admin image tests enforce it. - No new pip dependency.
- Requires
SHIELD_APPROVAL_TOKEN_PRIVATE_KEYon both planes (already needed for tool approvals). - Tier 2 only:
SHIELD_EXCEPTION_REVIEW_MODELand its endpoint, data plane. - Extension version 1.3.0: repack the self-hosted package and update the store listing.
- Rebuild both images.
7. Failure modes & edge cases
| Case | Behaviour |
|---|---|
| Grant for a different prompt, user or destination | Ignored; the block stands; recorded as grant_mismatch |
| Grant expired before the resend | Block stands; the extension offers to request again |
| Grant already used | Block stands |
| Redis down when burning the nonce | Fail closed: the block stands. A grant that cannot be marked used could be replayed |
| Redis down when creating a request | 503; the block itself is unaffected |
| User edits the prompt after approval | Different hash: block stands. The banner says the approval was for the original text |
| Policy changed between request and approval | The grant only waives what blocked it at request time; a new failure still blocks |
| Tenant in monitor mode | Nothing is blocked, so nothing can be requested (409 not_blocked) |
| Prompt longer than 4000 characters | Stored truncated; the hash covers the whole prompt; the reviewer sees the truncation notice |
| Two approvers act at once | First decision wins; the second gets the decided record |
| Approver is the requester | 403 |
| Request never answered | expired after request_ttl_s; the user is told |
8. Test plan (Definition of Done)
- Create: blocked prompt creates a request with the server’s
blocked_by; unblocked prompt is 409; non-appealable is 403; duplicate returns the same request; the per-user limit is 429; feature off is 404. - Poll: only the requester reads it; another tenant gets 404.
- Approve and deny: status changes, grant minted with the right claims, requester cannot approve, write gate enforced, admin audit rows written.
- Guard path: a valid grant turns a block into a pass exactly once; then every row of §7 (wrong prompt, wrong user, wrong destination, expired, used, new failure, non-appealable, Redis down).
- No grant header:
/guardrails/inputresponses are byte-identical to today (existing tests unchanged). - Webhook events carry no prompt text.
- Counters: requested, approved and false-positive per policy.
- Extension unit tests: banner text, request, polling, resend with grant.
- Admin image import guard; full suite green in a clean venv; CI
pytestgate passes.
Tasks
One branch, one PR, in this order.
- Better block message in the extension: policy name and reason. Useful on its own.
- Requests: store, create and poll routes, tenant settings, limits, webhook events.
- Review and grant: approve and deny routes, grant minting, the grant
check on
/guardrails/input. - Portal: the review queue with approve, deny and false-positive.
- Extension: request form, status, resend with the grant; version 1.3.0.
- Second opinion (opt-in), after there is data from tasks 2 to 5: on create, a larger model reviews blocks that came from model-judged guardrails; an overturn issues the grant at once with the approver recorded as the model.