You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(supervisor): let request middleware mark a single request as pending a human decision #4406
I use OpenShell to run coding agents (Claude Code) in sandboxes that act on GitHub through its REST API.
I'm building supervisor middleware that sends a few risky requests, like merging a PR into the default branch, to a named person for approval, and lets routine requests through. I'm building it for OKed, an approval service for agents.
I directly hit the limits below on v0.1.2 while testing this against a real GitHub repo. I then checked the code at v0.1.3.
This affects whether teams can let agents act on real systems without approving every write in advance or blocking writes entirely.
Problem Statement
Request middleware can allow, deny, or modify a request within a bounded evaluation timeout. There is no outcome that means "this specific request needs a human decision."
So middleware has no first-class way to tell the agent "a person is deciding on this exact request; wait, then retry," and OpenShell has no way to release exactly that request once it is approved.
Impact / Why This Matters
Today the only working pattern is deny and retry, built separately in each middleware:
Deny at once with a reason code that encodes "pending, retry".
Track the approval outside OpenShell.
When the identical request comes back and has been approved, allow it once.
I built this, and it works with Claude Code. It's fragile:
The agent has to read a code. Only reason_code reaches the sandbox (1 to 64 bytes, ^[a-z][a-z0-9_]*$), so the instruction is packed into an identifier: oked_approval_pending_retry_same_request_every_60s_up_to_10m. With a bare pending code, the agent stopped and asked its user instead of retrying.
It looks like a final refusal. The sandbox sees middleware_denied. Each middleware invents its own convention.
Waiting costs model tokens. The agent waits by retrying, so every retry is another model turn.
Enforcement lives outside OpenShell. Single use, expiry, and the audit trail all sit in the middleware.
Holding the request inside the evaluation is not a way out:
The evaluation timeout is at most 30 s.
For HTTPS, OpenShell opens the upstream TCP and TLS connection at CONNECT, before it reads the request or calls middleware. GitHub closed that idle connection: holds over about 4 s broke requests to api.github.com in my tests.
Middleware runs 32 evaluations at a time per sandbox, and up to 64 more wait with no deadline. Held requests keep their sandbox connections open. At about 100 held requests, new connections to all hosts stalled, also to hosts without middleware, and DNS failed once. The code points to the 96 relay streams per sandbox.
OpenShell already has a good pattern for this, but only for rules: Policy Advisor answers at once with OpenShell-written guidance, and the agent waits on a policy.local long-poll without spending model tokens. Middleware can't start that flow for one request.
Related: #2881 asks for single-use and expiring grants, and its triage asks what one "use" means. #4201 describes another middleware that enforces single-use human approvals today. #4359 deprecates EvaluateHttpRequest for 0.2.0.
Proposed Design
Base mode: answer at once, OpenShell owns the wait. No request is held, so no evaluation slot, upstream connection, or relay stream is kept while a person decides.
Middleware author workflow.
For a request that needs a person, the middleware returns pending with structured fields only: a pending_id (same character rules as reason_code) and an expiry. No free text, consistent with OpenShell never forwarding free-form middleware text to the sandbox.
Middleware opts in with a required capability, for example openshell.supervisor-middleware.pending, the same way http-session does. Middleware that doesn't opt in behaves exactly as today.
OpenShell learns the decision from the middleware, either by asking it again with the pending_id, or through a new call that returns when the decision is made (open question below). Operator workflow. Middleware is bound as today. The policy can set the longest wait. Notional example:
network_middlewares:
approvals:
middleware: oked-approvalsendpoints:
include: ["api.github.com"]pending:
max_wait: 10m# after this, the pending request is answered as expired
What the agent sees
At once: a documented pending response with fixed text written by OpenShell, modeled on Policy Advisor's next_steps and agent_guidance. Notional example:
HTTP/1.1 403 ForbiddenContent-Type: application/jsonConnection: close
{
"error": "approval_pending",
"detail": "This request is waiting for a human decision.",
"policy": "github-policy",
"middleware": "approvals",
"pending_id": "apr_7f3c2a",
"expires_at": "2026-10-11T10:10:00Z",
"next_steps": [
{"action": "wait", "method": "GET", "url": "http://policy.local/v1/pending/apr_7f3c2a/wait?timeout=300"},
{"action": "retry", "detail": "If approved, send the identical request again."}
],
"agent_guidance": "<fixed text written by OpenShell>"
}
While waiting: the agent calls the wait URL, a long-poll like /v1/proposals/{id}/wait. It returns approved, denied, expired, or pending when its timeout ends. Waiting costs no model tokens.
Approved: the agent sends the identical request again. OpenShell matches it on sandbox, method, host, port, path, query, and a digest of the body taken before credentials are added. Exactly one matching request is released, claimed atomically before forwarding, then continues through the normal pipeline, including credential injection. Later identical requests are evaluated as new. This is the same single-use primitive feat(policy): native expiration / single-use bounds for approved policy grants (expires_at, max_uses) #2881 asks for, bound to one request.
Possible later step: a short hold. If maintainers want the agent to see only a slow request when the person answers within seconds, a hold would need OpenShell to re-dial the upstream after the decision, to cap held requests per sandbox well below the 96 relay streams, and to count held request bodies. I suggest leaving it out of the first version.
Open questions:
Status code for the pending response: 403 like Policy Advisor denials, or 503? The capacity 503 has no Retry-After on purpose (feat(middleware): inspect WebSocket text messages #2477), so the pending response needs its own error value either way.
How the decision reaches OpenShell: re-evaluation with the pending_id, or a new middleware call that returns when decided?
Which headers, if any, belong in the request match?
Should pending also be added to EvaluateHttpRequest, or only to the session hooks?
Acceptance Criteria
Middleware that opts in can return a pending result. Middleware that doesn't opt in behaves exactly as today.
The agent gets an immediate, documented pending response with pending_id, expiry, and OpenShell-written guidance. It contains no free text from the middleware and is distinct from the capacity 503.
A wait call on policy.local returns the decision when it is made, or pending when its timeout ends.
After approval, exactly one matching request is forwarded. A request that differs in method, host, port, path, query, or body is not matched.
Denied and expired outcomes return distinct, documented errors.
100 pending requests in one sandbox don't block its other traffic, including DNS.
Pending, decision, and release are logged with the pending_id and the request digest.
Alternatives Considered
Supervisor middleware as it is today. This is what I built. Deny and retry works with Claude Code, but it depends on the model reading a private code, spends tokens while waiting, and keeps single use and expiry outside OpenShell. Holding inside the evaluation fails on the timeout, the idle upstream connection, and the per-sandbox limits.
Policy Advisor. It already has the right waiting pattern. But it approves rules, not one request with its real values. The agent writes the proposal, middleware cannot start it, and a rule can't bind a request body.
Gateway interceptors. They act on gateway API calls, not sandbox egress.
Credential drivers and token grants. They store or fetch credentials. They don't see the method, path, or body, so they can't decide on one request.
Agent-side prompts (agent hooks, ACP session/request_permission, MCP elicitation). They decide per action, but they run inside the agent, which the sandbox treats as untrusted.
Updated the proposal after reading the v0.1.3 code:
Corrected the capacity numbers: 32 running, 64 queued. The full stall matches the 96 relay streams per sandbox.
Changed the base design from holding the request to answering at once, with a policy.local wait modeled on Policy Advisor. A hold is now an optional later step.
User Story
I use OpenShell to run coding agents (Claude Code) in sandboxes that act on GitHub through its REST API.
I'm building supervisor middleware that sends a few risky requests, like merging a PR into the default branch, to a named person for approval, and lets routine requests through. I'm building it for OKed, an approval service for agents.
I directly hit the limits below on v0.1.2 while testing this against a real GitHub repo. I then checked the code at v0.1.3.
This affects whether teams can let agents act on real systems without approving every write in advance or blocking writes entirely.
Problem Statement
Request middleware can allow, deny, or modify a request within a bounded evaluation timeout. There is no outcome that means "this specific request needs a human decision."
So middleware has no first-class way to tell the agent "a person is deciding on this exact request; wait, then retry," and OpenShell has no way to release exactly that request once it is approved.
Impact / Why This Matters
Today the only working pattern is deny and retry, built separately in each middleware:
I built this, and it works with Claude Code. It's fragile:
The agent has to read a code. Only
reason_codereaches the sandbox (1 to 64 bytes,^[a-z][a-z0-9_]*$), so the instruction is packed into an identifier:oked_approval_pending_retry_same_request_every_60s_up_to_10m. With a barependingcode, the agent stopped and asked its user instead of retrying.It looks like a final refusal. The sandbox sees
middleware_denied. Each middleware invents its own convention.Waiting costs model tokens. The agent waits by retrying, so every retry is another model turn.
Retry matching is guesswork.
request_idis new for every request, andoriginating_processis always empty in this version (RFC 0009, feat(supervisor): populateRequestContext.originating_processfor supervisor middleware #4269). The middleware has to fingerprint requests itself.Enforcement lives outside OpenShell. Single use, expiry, and the audit trail all sit in the middleware.
Holding the request inside the evaluation is not a way out:
The evaluation timeout is at most 30 s.
For HTTPS, OpenShell opens the upstream TCP and TLS connection at CONNECT, before it reads the request or calls middleware. GitHub closed that idle connection: holds over about 4 s broke requests to
api.github.comin my tests.Middleware runs 32 evaluations at a time per sandbox, and up to 64 more wait with no deadline. Held requests keep their sandbox connections open. At about 100 held requests, new connections to all hosts stalled, also to hosts without middleware, and DNS failed once. The code points to the 96 relay streams per sandbox.
OpenShell already has a good pattern for this, but only for rules: Policy Advisor answers at once with OpenShell-written guidance, and the agent waits on a
policy.locallong-poll without spending model tokens. Middleware can't start that flow for one request.Related: #2881 asks for single-use and expiring grants, and its triage asks what one "use" means. #4201 describes another middleware that enforces single-use human approvals today. #4359 deprecates
EvaluateHttpRequestfor 0.2.0.Proposed Design
Base mode: answer at once, OpenShell owns the wait. No request is held, so no evaluation slot, upstream connection, or relay stream is kept while a person decides.
Middleware author workflow.
pendingwith structured fields only: apending_id(same character rules asreason_code) and an expiry. No free text, consistent with OpenShell never forwarding free-form middleware text to the sandbox.openshell.supervisor-middleware.pending, the same wayhttp-sessiondoes. Middleware that doesn't opt in behaves exactly as today.pending_id, or through a new call that returns when the decision is made (open question below).Operator workflow. Middleware is bound as today. The policy can set the longest wait. Notional example:
What the agent sees
next_stepsandagent_guidance. Notional example:/v1/proposals/{id}/wait. It returnsapproved,denied,expired, orpendingwhen its timeout ends. Waiting costs no model tokens.approval_deniedandapproval_expired, so the agent knows not to retry.Logs. The pending result, the decision, and the release appear in logs and OCSF events with the
pending_idand the request digest. The digest definition can be shared with the signed action records in feat(supervisor): signed, chained action records for middleware-evaluated requests, with a middleware evidence reference #4201, so a reviewer can check that what was approved is what ran.Possible later step: a short hold. If maintainers want the agent to see only a slow request when the person answers within seconds, a hold would need OpenShell to re-dial the upstream after the decision, to cap held requests per sandbox well below the 96 relay streams, and to count held request bodies. I suggest leaving it out of the first version.
Open questions:
Retry-Afteron purpose (feat(middleware): inspect WebSocket text messages #2477), so the pending response needs its ownerrorvalue either way.pending_id, or a new middleware call that returns when decided?pendingalso be added toEvaluateHttpRequest, or only to the session hooks?Acceptance Criteria
pending_id, expiry, and OpenShell-written guidance. It contains no free text from the middleware and is distinct from the capacity 503.policy.localreturns the decision when it is made, orpendingwhen its timeout ends.pending_idand the request digest.Alternatives Considered
session/request_permission, MCP elicitation). They decide per action, but they run inside the agent, which the sandbox treats as untrusted.