Skip to content

Supervisor middleware #1733

Description

@pimlock

Summary

Supervisor middleware is OpenShell's data-plane extensibility mechanism. It lets operators extend supervisor-managed data paths with trusted components that can inspect, transform, block, or annotate operations at defined hooks.

This parent issue tracks the supervisor middleware capability as a whole. Individual hook families, deployment models, security boundaries, and operational improvements are tracked as focused subissues.

Status

The initial supervisor middleware implementation landed in #2010 through PR #2027. It introduced the first hook for outbound HTTP requests before credential injection, with built-in and operator-run middleware support.

Further capabilities are tracked independently rather than as sequential implementation phases.

References

Activity

  1. jhjaggars commented on Jun 4, 2026

    @jhjaggars
    Contributor

    I opened an issue with a similar goal the other day: #1694
    In my case the middleware needs to execute after credential injection.

  2. pimlock commented on Jun 4, 2026

    @pimlock
    CollaboratorAuthor

    @jhjaggars

    I opened an issue with a similar goal the other day: #1694 In my case the middleware needs to execute after credential injection.

    I saw the issue come in as I was working on this and realized that the initial shape didn't solve this, as it's assuming a single pre-credential injection hook.

    My initial thought was to add a post-credential hook, but limit it to be available to built-in middleware only, so the user-provided middleware couldn't hook into there, but a first-party sigv4 one could. We could remove this limitation somewhere down the road if needed, e.g. once we support deploying middlewares in a controlled environment.

    I still need few hours to get the proposal to a reviewable state, so that part around first and third party middlewares is not yet fleshed out.

  3. pimlock commented on Jun 5, 2026

    @pimlock
    CollaboratorAuthor

    @jhjaggars thanks again for bringing this up, I added a comment in your original issue: #1694 (comment)

    Please let me know if you have any questions/concerns/feedback!

  4. eeee2345 commented on Jun 11, 2026

    @eeee2345

    Landed here from #1272 (thanks @johntmyers for the redirect). One concrete data point for the "trusted middleware / guard service" shape in the egress hook:

    We maintain ATR (Agent Threat Rules, MIT) — an open detection-rule standard for agent traffic (prompt injection, tool poisoning, context/credential exfiltration, MCP attacks), Sigma/YARA-style content rules. It's a natural fit for the egress-inspection stage you describe: given outbound request content, return allow / block / annotate plus the matched rule IDs, with no model call (pure rule eval), so it stays cheap enough to run inline before forwarding.

    If helpful for the v1 hook design, I can prototype an ATR-backed egress guard against whatever interface signature you land on, as a reference implementation of the "guard service" role. Keeping it to the design here rather than re-opening #1272.

  5. changed the title [-]Sandbox egress middleware RFC[/-] [+]Sandbox egress middleware[/+] on Jun 26, 2026
  6. 4 remaining items

  7. changed the title [-]Sandbox egress middleware[/-] [+]Supervisor middleware[/+] on Jul 22, 2026
  8. chkp-stevegi commented on Jul 28, 2026

    @chkp-stevegi

    One thought that came to mind while reading both this RFC and the discussion from @eeee2345 around Agent Threat Rules (ATR).

    I really like the direction ATR is taking. To me, it feels similar to what Sigma did for SIEM detections: a portable way to express AI security findings without coupling them to a specific vendor or runtime.

    It did make me wonder if there is another layer that could be standardized even earlier in the pipeline.

    Today, we're making decisions about concrete runtime operations (HTTP today, potentially filesystem, process execution, MCP, etc. in the future). However, downstream security, governance and observability systems often need more than what is happening, they also need why it is happening.

    Would it make sense for Supervisor Middleware to define a portable AI Context object that accompanies every middleware invocation, independent of the operation being intercepted?

    Something along the lines of:

    context: session: id: conversation_id: trace_id: agent: id: name: version: user: id: identity: tenant: model: provider: model: endpoint: objective: goal: current_task: tool: name: invocation_id: mcp: server: tool: capability: request: operation: http destination: method: policy: decision: rationale: ...

    The exact schema is obviously open for discussion, but the idea would be that middleware can inspect and enrich a shared context rather than each integration inventing its own metadata model.

    This also feels complementary to ATR rather than overlapping with it. ATR (or any future rule engine) could consume this richer semantic context, while OpenShell remains responsible for exposing consistent runtime context. Likewise, observability platforms, governance systems and security vendors would all benefit from a common contract without being tightly coupled to OpenShell internals.

    To me this feels analogous to how Linux exposes syscalls, Windows exposes Event Logs, or OpenTelemetry defines telemetry semantics: once there is a common context model, an ecosystem can grow around it.

    Curious whether others think a portable AI Context abstraction belongs at the Supervisor Middleware layer, or whether that should live as a separate specification that OpenShell simply adopts.

  9. eeee2345 commented on Aug 9, 2026

    @eeee2345

    @chkp-stevegi @pimlock — rather than answer this from opinion, I built against the existing contract to see what the gap actually is. The result is at Agent-Threat-Rule/openshell-middleware-atr: an operator-run service implementing openshell.middleware.v1.SupervisorMiddleware, bound to HTTP_REQUEST / PRE_CREDENTIALS, evaluating egress against the ATR corpus. MIT, out of tree, no changes to OpenShell. It runs — npm run verify builds, runs 18 tests, then starts the service and drives it over gRPC with the vendored proto.

    Disclosure: I author ATR and have written about it twice in this thread. Nothing below asks for anything to be adopted; it is what implementing the contract surfaced.

    The direct question

    Mostly a separate specification, and a smaller piece at this layer than the sketch suggests — because the contract has already answered part of it. RequestContext exists and carries request_id, sandbox_id and originating_process. Finding exists as a typed, audit-safe output. So the input observation and the emitted finding are already separate shapes, deliberately. Folding policy.decision / policy.rationale into a context object would undo that: the second middleware in a chain would receive the first one's verdict as if it were ground truth about the world.

    What is genuinely missing is narrower than a full context object, and I can now say how much it costs.

    What the gap costs, measured

    The service prints this on startup:

    768 rules loaded, 408 not reachable on this event shape
    

    ATREngine skips a rule whose declared source type does not match the event type. An HTTP egress cannot be mapped to any agent-semantic type, so 53% of the corpus is never evaluated through this hook. That is not a quality judgement about the rules — it is the cost of the vocabulary mismatch, and it was invisible to me until I ran it.

    Three places the mapping is lossy, encountered rather than theorised:

    Concept What the contract carries What I had to do
    Event type nothing agent-semantic Map egress to tool_call. Lossy: rules written for tool calls now also see raw network egress that no tool call produced.
    Session identity sandbox_id Use it as the session id. A sandbox is not a session — several sessions can share one, and a session can outlive one — so session-scoped correlation is unavailable.
    Agent identity originating_process.binary Leave it unset. That field is /usr/bin/python3; writing it as an agent id would put an interpreter path into an audit trail as if it identified an agent.

    The service reports the 408 as an explicit atr_not_evaluated finding rather than returning a clean result, because reporting "we could not look" as "we looked and found nothing" makes every coverage number computed downstream wrong in the flattering direction. That behaviour is the part I would most like to see in the contract itself: a way for a middleware to declare which dimensions it could not observe. Nothing else can reconstruct it.

    Four contract details that cost me a debugging cycle each

    Offered because an out-of-tree service gets no help from your CI, @pimlock, and these are all silent failures:

    • reason_code must match ^[a-z][a-z0-9_]{0,63}$. A dotted code such as atr.rule_match is rejected. Worth a line in the field comment.
    • The 32-finding cap is per stage; a rules engine hits it easily, so a service needs its own truncation policy or it becomes a middleware failure.
    • reason is discarded above 4 KiB.
    • config arrives as google.protobuf.Struct. With @grpc/proto-loader the wire form is {fields: {k: {stringValue: v}}}, so a parser expecting a plain object sees one unknown key called fields and rejects every request. My unit tests passed throughout — only driving the actual gRPC path caught it. If a Go or Rust example is the only reference, this trap is invisible to anyone implementing in another language.

    One observation rather than a complaint: on a jailbreak payload the top finding came back as ATR-2026-00120 "SKILL.md Prompt Injection". The detection is right, but that rule was authored for skill files, so its label no longer describes the situation it fired in. Rule labels turn out to be surface-specific in a way only a cross-surface integration makes visible — which is an argument for the context object carrying enough to disambiguate surfaces, and against assuming a finding label travels well.

    So, concretely

    Pin the two things that are specific to interception and that nobody outside this layer can supply — an outcome for the intercepted operation, and a way to declare unobserved dimensions — and reference OTel GenAI for the identity and provenance fields rather than restating them. That keeps the surface small enough to land and leaves the wider semantic-context question to a spec not tied to OpenShell's release cycle.

    Happy to be told the mapping choices above are wrong; they are all in one file (src/mapping.ts) with the reasoning attached, and it costs nothing to change them.

  10. github-actions commented on Aug 28, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  11. johntmyers commented on Sep 17, 2026

    @johntmyers
    Collaborator

    📋 triage-agent

    Closing this umbrella as delivered. RFC 0009 landed through #1738, and the initial supervisor middleware capability landed through #2027. The issue itself records that subsequent hook families, deployment models, security boundaries, and operational improvements are tracked as focused issues, so there is no remaining independently finishable outcome on this parent.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

rfcstate:staleInactive item at risk of automatic closure.

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions