Repository navigation
Supervisor middleware #1733
Description
Activity
- added a parent issue
on Jun 3, 2026 I opened an issue with a similar goal the other day: #1694
In my case the middleware needs to execute after credential injection.I opened an issue with a similar goal the other day: #1694 In my case the middleware needs to execute after credential injection.
I saw the issue come in as I was working on this and realized that the initial shape didn't solve this, as it's assuming a single pre-credential injection hook.
My initial thought was to add a post-credential hook, but limit it to be available to built-in middleware only, so the user-provided middleware couldn't hook into there, but a first-party sigv4 one could. We could remove this limitation somewhere down the road if needed, e.g. once we support deploying middlewares in a controlled environment.
I still need few hours to get the proposal to a reviewable state, so that part around first and third party middlewares is not yet fleshed out.
Reacted by Jesse Jaggars@jhjaggars thanks again for bringing this up, I added a comment in your original issue: #1694 (comment)
Please let me know if you have any questions/concerns/feedback!
Landed here from #1272 (thanks @johntmyers for the redirect). One concrete data point for the "trusted middleware / guard service" shape in the egress hook:
We maintain ATR (Agent Threat Rules, MIT) — an open detection-rule standard for agent traffic (prompt injection, tool poisoning, context/credential exfiltration, MCP attacks), Sigma/YARA-style content rules. It's a natural fit for the egress-inspection stage you describe: given outbound request content, return allow / block / annotate plus the matched rule IDs, with no model call (pure rule eval), so it stays cheap enough to run inline before forwarding.
If helpful for the v1 hook design, I can prototype an ATR-backed egress guard against whatever interface signature you land on, as a reference implementation of the "guard service" role. Keeping it to the design here rather than re-opening #1272.
- removed a parent issue
on Jun 15, 2026 - added a parent issue
on Jun 15, 2026 - added a commit that references this issue
on Jun 16, 2026 - changed the title
[-]Sandbox egress middleware RFC[/-][+]Sandbox egress middleware[/+]on Jun 26, 2026 4 remaining items
One thought that came to mind while reading both this RFC and the discussion from @eeee2345 around Agent Threat Rules (ATR).
I really like the direction ATR is taking. To me, it feels similar to what Sigma did for SIEM detections: a portable way to express AI security findings without coupling them to a specific vendor or runtime.
It did make me wonder if there is another layer that could be standardized even earlier in the pipeline.
Today, we're making decisions about concrete runtime operations (HTTP today, potentially filesystem, process execution, MCP, etc. in the future). However, downstream security, governance and observability systems often need more than what is happening, they also need why it is happening.
Would it make sense for Supervisor Middleware to define a portable AI Context object that accompanies every middleware invocation, independent of the operation being intercepted?
Something along the lines of:
context: session: id: conversation_id: trace_id: agent: id: name: version: user: id: identity: tenant: model: provider: model: endpoint: objective: goal: current_task: tool: name: invocation_id: mcp: server: tool: capability: request: operation: http destination: method: policy: decision: rationale: ...The exact schema is obviously open for discussion, but the idea would be that middleware can inspect and enrich a shared context rather than each integration inventing its own metadata model.
This also feels complementary to ATR rather than overlapping with it. ATR (or any future rule engine) could consume this richer semantic context, while OpenShell remains responsible for exposing consistent runtime context. Likewise, observability platforms, governance systems and security vendors would all benefit from a common contract without being tightly coupled to OpenShell internals.
To me this feels analogous to how Linux exposes syscalls, Windows exposes Event Logs, or OpenTelemetry defines telemetry semantics: once there is a common context model, an ecosystem can grow around it.
Curious whether others think a portable AI Context abstraction belongs at the Supervisor Middleware layer, or whether that should live as a separate specification that OpenShell simply adopts.
Reacted by matt and Adam Lin@chkp-stevegi @pimlock — rather than answer this from opinion, I built against the existing contract to see what the gap actually is. The result is at Agent-Threat-Rule/openshell-middleware-atr: an operator-run service implementing
openshell.middleware.v1.SupervisorMiddleware, bound toHTTP_REQUEST/PRE_CREDENTIALS, evaluating egress against the ATR corpus. MIT, out of tree, no changes to OpenShell. It runs —npm run verifybuilds, runs 18 tests, then starts the service and drives it over gRPC with the vendored proto.Disclosure: I author ATR and have written about it twice in this thread. Nothing below asks for anything to be adopted; it is what implementing the contract surfaced.
The direct question
Mostly a separate specification, and a smaller piece at this layer than the sketch suggests — because the contract has already answered part of it.
RequestContextexists and carriesrequest_id,sandbox_idandoriginating_process.Findingexists as a typed, audit-safe output. So the input observation and the emitted finding are already separate shapes, deliberately. Foldingpolicy.decision/policy.rationaleinto acontextobject would undo that: the second middleware in a chain would receive the first one's verdict as if it were ground truth about the world.What is genuinely missing is narrower than a full context object, and I can now say how much it costs.
What the gap costs, measured
The service prints this on startup:
768 rules loaded, 408 not reachable on this event shapeATREngineskips a rule whose declared source type does not match the event type. An HTTP egress cannot be mapped to any agent-semantic type, so 53% of the corpus is never evaluated through this hook. That is not a quality judgement about the rules — it is the cost of the vocabulary mismatch, and it was invisible to me until I ran it.Three places the mapping is lossy, encountered rather than theorised:
Concept What the contract carries What I had to do Event type nothing agent-semantic Map egress to tool_call. Lossy: rules written for tool calls now also see raw network egress that no tool call produced.Session identity sandbox_idUse it as the session id. A sandbox is not a session — several sessions can share one, and a session can outlive one — so session-scoped correlation is unavailable. Agent identity originating_process.binaryLeave it unset. That field is /usr/bin/python3; writing it as an agent id would put an interpreter path into an audit trail as if it identified an agent.The service reports the 408 as an explicit
atr_not_evaluatedfinding rather than returning a clean result, because reporting "we could not look" as "we looked and found nothing" makes every coverage number computed downstream wrong in the flattering direction. That behaviour is the part I would most like to see in the contract itself: a way for a middleware to declare which dimensions it could not observe. Nothing else can reconstruct it.Four contract details that cost me a debugging cycle each
Offered because an out-of-tree service gets no help from your CI, @pimlock, and these are all silent failures:
reason_codemust match^[a-z][a-z0-9_]{0,63}$. A dotted code such asatr.rule_matchis rejected. Worth a line in the field comment.- The 32-finding cap is per stage; a rules engine hits it easily, so a service needs its own truncation policy or it becomes a middleware failure.
reasonis discarded above 4 KiB.configarrives asgoogle.protobuf.Struct. With@grpc/proto-loaderthe wire form is{fields: {k: {stringValue: v}}}, so a parser expecting a plain object sees one unknown key calledfieldsand rejects every request. My unit tests passed throughout — only driving the actual gRPC path caught it. If a Go or Rust example is the only reference, this trap is invisible to anyone implementing in another language.
One observation rather than a complaint: on a jailbreak payload the top finding came back as
ATR-2026-00120 "SKILL.md Prompt Injection". The detection is right, but that rule was authored for skill files, so its label no longer describes the situation it fired in. Rule labels turn out to be surface-specific in a way only a cross-surface integration makes visible — which is an argument for the context object carrying enough to disambiguate surfaces, and against assuming a finding label travels well.So, concretely
Pin the two things that are specific to interception and that nobody outside this layer can supply — an outcome for the intercepted operation, and a way to declare unobserved dimensions — and reference OTel GenAI for the identity and provenance fields rather than restating them. That keeps the surface small enough to land and leaves the wider semantic-context question to a spec not tied to OpenShell's release cycle.
Happy to be told the mapping choices above are wrong; they are all in one file (
src/mapping.ts) with the reasoning attached, and it costs nothing to change them.- added a sub-issue
on Aug 10, 2026 This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 28, 2026 📋 triage-agent
Closing this umbrella as delivered. RFC 0009 landed through #1738, and the initial supervisor middleware capability landed through #2027. The issue itself records that subsequent hook families, deployment models, security boundaries, and operational improvements are tracked as focused issues, so there is no remaining independently finishable outcome on this parent.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Summary
Supervisor middleware is OpenShell's data-plane extensibility mechanism. It lets operators extend supervisor-managed data paths with trusted components that can inspect, transform, block, or annotate operations at defined hooks.
This parent issue tracks the supervisor middleware capability as a whole. Individual hook families, deployment models, security boundaries, and operational improvements are tracked as focused subissues.
Status
The initial supervisor middleware implementation landed in #2010 through PR #2027. It introduced the first hook for outbound HTTP requests before credential injection, with built-in and operator-run middleware support.
Further capabilities are tracked independently rather than as sequential implementation phases.
References
rfc/0009-supervisor-middleware/