Skip to content

test(e2e): cover tasks and add the Tasks and upgrade guides - #3859

Open
AndrewBarba wants to merge 10 commits into
barba/tasks-9-slack-wait-postfrom
barba/tasks-10-release-readiness
Open

AndrewBarba wants to merge 10 commits into
barba/tasks-9-slack-wait-postfrom
barba/tasks-10-release-readiness

Conversation

@AndrewBarba

@AndrewBarba AndrewBarba commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

PRs 1–9 of #3821 landed the task model with a passing build but, by design, almost no new tests. This PR adds the coverage and docs the plan requires before release.

  • Mock-model e2e: steering versus sleep, approvals, and withdrawn asks; task() start, wait, and cancel; the turn rule in root, child, schedule, and MCP sessions; a serve tool continued, cancelled, and continued again; local and remote agents continued and corrected by taskId.
  • Real-model evals: a new agent-tasks fixture gates the release on task planning. The model waits when needed, fans out then waits, keeps tasks after an unrelated message, redirects or cancels an agent, continues instead of restarting, doesn't wait when told there's no rush, and doesn't run a side effect before its result.
  • Docs: a Tasks guide and an Upgrade to Tasks page, organized around execute, task, and serve.
  • Model text: the task system block and the task_wait interrupt text are tuned against the real-model evals, and the plan quotes the tuned text.
  • One decision per question: a steer or cancel could resolve a ctx.ask as cancelled while the session still accepted a person's answer, so ask_question and the channel disagreed (#3845 review). The session now decides each question once and sends the decision on the run's control hook, with its calls, interrupts, and cancels (#3821 review). task_cancel of a task() reports its questions cancelled, and question requestIds become <runId>-ask-<n>.

This PR builds on #3857 (related to #1084). Slack's post before a wait keeps its unit test from #3857, since there is no Slack e2e harness.

Validation

  • pnpm build, pnpm typecheck, pnpm lint, pnpm guard:invariants, pnpm guard:fixtures, pnpm check:deps, pnpm check:templates, pnpm docs:check, pnpm test:unit and pnpm test:integration all pass.
  • pnpm test:scenario passes, apart from the known hostless-agent-workspace npm 401, which passed when rerun alone.
  • Mock-model evals pass: agent-workflow-tools 26/26 and agent-subagents 9/9. The agent-tasks real-model evals typecheck and load, and e2e-local CI runs them.
  • The question change was checked separately with pnpm typecheck, pnpm lint, pnpm fmt, pnpm guard:invariants, pnpm docs:check, the focused workflow, session, task, and HITL unit tests (146) and integration tests (23), and the agent-workflow-tools build and typecheck. Its new sign-off-plan.cancel eval has not run yet; CI runs it.

Checklist

  • This change was requested or approved by a maintainer
  • I ran the relevant checks from CONTRIBUTING.md
  • I added tests and documentation where relevant
  • I added a changeset if this touches the published eve package
  • DCO sign-off passes for every commit (git commit --signoff)

@vercel

vercel Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
eve-docs Ready Ready Preview, v0 Sep 28, 2026 12:26am UTC
eve-pkg Ready Ready Preview, v0 Sep 28, 2026 12:26am UTC

@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Bundle + Package Summary: apps/fixtures/weather-agent

Key takeaways

  • No notable deltas vs barba/tasks-9-slack-wait-post (fbbac65).

Delta vs barba/tasks-9-slack-wait-post (fbbac65)

Area Metric Baseline Current Delta
Package Packed tarball 8.39 MB 8.40 MB +8.6 kB ⚠️
Package Unpacked publish size 30.82 MB 30.85 MB +31.3 kB ⚠️
Package Installed footprint 77.19 MB 77.22 MB +31.3 kB ⚠️
Package Published files 3848 3850 +2
Package Installed files 7815 7817 +2
Package Installed package instances 33 33 0
Package Distinct installed package names 32 32 0
Package Installed dependency edges 51 51 0
Package Installed optional peer edges 9 9 0
Runtime Unique function payloads 2 2 0
Runtime Total function bytes 18.92 MB 18.94 MB +27.4 kB ⚠️
Runtime Public routes 17 17 0
Changed function payloads vs barba/tasks-9-slack-wait-post (fbbac65) (2)
Function Status Baseline Current Delta Route changes
functions/__server.func changed 9.46 MB 9.47 MB +13.7 kB ⚠️ none
functions/.well-known/workflow/v1/flow.func changed 9.46 MB 9.47 MB +13.7 kB ⚠️ none

eve init install

Metric Baseline Current Delta
Installed footprint 115.48 MB 115.51 MB +31.3 kB ⚠️
Installed packages 97 97 0
dependencies 4 4 0
devDependencies 2 2 0
Dependency package bytes 48.21 MB 48.25 MB +31.3 kB ⚠️
devDependency package bytes 5.11 MB 5.11 MB 0 B ➖
Build Metadata
  • Preset: vercel
  • Nitro: nitro@3.0.260903-beta
  • Output directory: apps/fixtures/weather-agent/.vercel/output
  • Build metadata timestamp: 2026-09-28T00:27:18.682Z
  • Route aliases: 17 public, 1 internal (18 total aliases)
  • Vercel routes in config: 20
  • Severity legend: 🔴 dominant/large, 🟠 notable, 🟡 watch, ⚪ small
Package Drill-Down

Package Details

  • Package: eve@0.67.2
  • Package directory: packages/eve
  • Tarball: 8.40 MB (eve-0.67.2.tgz)
  • Unpacked payload: 30.85 MB across 3850 published files
  • Installed footprint: 77.22 MB across 7817 installed files
  • Installed root package: 30.57 MB
  • Installed dependencies: 46.65 MB
  • Installed package instances: 33
  • Distinct installed package names: 32
  • Installed dependency edges: 51
  • Installed optional peer edges: 9
  • Runtime dependencies: 2
  • Peer dependencies: 7 (6 optional)

Installed footprint is measured from an isolated temporary npm install of the packed tarball.
Graph metrics read only package.json files in package directories directly beneath a node_modules boundary, including nested boundaries. Each directory is one package instance; distinct names come from those manifests. Dependency edges count each unique name in dependencies or optionalDependencies per instance; optional peer edges count peerDependencies marked optional.

Heavy installed dependencies

  • eve: 30.57 MB (39.6%)
  • @rolldown/binding-linux-x64-gnu: 19.15 MB (24.8%)
  • ai: 7.73 MB (10.0%)
  • zod: 6.14 MB (8.0%)
  • undici: 3.52 MB (4.6%)
Publish payload breakdown
Published file size
🔴 dist/src/compiled/shadcn-registry/index.js       [##############..........] 9.70 MB 31.4%
🟠 dist/src/compiled/@photon-ai/chat-adapter-ime... [###.....................] 2.29 MB 7.4%
🟠 dist/src/compiled/@ai-sdk/code-mode/index.js     [#.......................] 1.03 MB 3.3%
🟡 dist/src/compiled/@vercel/blob/index.js          [#.......................] 604.9 kB 2.0%
🟡 dist/src/compiled/_chunks/workflow/signal-exi... [#.......................] 514.5 kB 1.7%
🔴 Other published files                            [########################] 16.72 MB 54.2%
Installed footprint breakdown
Installed package size
🔴 eve                             [########################] 30.57 MB 39.6%
🔴 @rolldown/binding-linux-x64-gnu [###############.........] 19.15 MB 24.8%
🔴 ai                              [######..................] 7.73 MB 10.0%
🔴 zod                             [#####...................] 6.14 MB 8.0%
🟠 undici                          [###.....................] 3.52 MB 4.6%
🟠 nitro                           [#.......................] 1.89 MB 2.5%
🔴 Other installed packages        [######..................] 8.22 MB 10.6%
Runtime dependencies (2)
Package Range Notes
nitro 3.0.260903-beta
undici 8.9.0
Peer dependencies (7)
Package Range Notes
@opentelemetry/api ^1.0.0 optional peer
ai catalog:
braintrust ^3.0.0 optional peer
chat ^4.41.0 optional peer
dd-trace ^6.13.0 optional peer
just-bash ^3.1.0 optional peer
microsandbox ^0.5.0 optional peer
eve init install drill-down

eve init install details

  • Command: eve init my-agent
  • Package manager: npm
  • Installed footprint: 115.51 MB across 9712 installed files
  • Installed packages: 97 total (91 transitive-only)
  • dependencies: 4 direct packages totaling 48.25 MB
  • devDependencies: 2 direct packages totaling 5.11 MB
  • Other transitive package files: 62.16 MB

Installed footprint is measured from an isolated temporary eve init my-agent using the current packed eve tarball.

Heavy installed dependencies

  • eve: 30.57 MB (26.5%)
  • @typescript/typescript-linux-x64: 27.95 MB (24.2%)
  • @rolldown/binding-linux-x64-gnu: 19.15 MB (16.6%)
  • zod: 9.76 MB (8.4%)
  • ai: 7.73 MB (6.7%)
Installed footprint breakdown
Installed package size
🔴 eve                              [########################] 30.57 MB 26.5%
🔴 @typescript/typescript-linux-x64 [######################..] 27.95 MB 24.2%
🔴 @rolldown/binding-linux-x64-gnu  [###############.........] 19.15 MB 16.6%
🔴 zod                              [########................] 9.76 MB 8.4%
🔴 ai                               [######..................] 7.73 MB 6.7%
🟠 undici                           [###.....................] 3.52 MB 3.0%
🔴 Other installed packages         [#############...........] 16.85 MB 14.6%
dependencies (4)
Package Range Installed size Share
@vercel/connect 2.2.0 194.7 kB 0.2%
ai ^7.0.105 7.73 MB 6.7%
eve file:eve-0.67.2.tgz 30.57 MB 26.5%
zod 4.5.4 9.76 MB 8.4%
devDependencies (2)
Package Range Installed size Share
@types/node 24.x 2.61 MB 2.3%
typescript 7.0.2 2.50 MB 2.2%
Function Drill-Down

Payload Size Graph

Unique function payload size and share of total
🔴 functions/.well-known/workflow/v1/flow.func     [########################] 9.47 MB 50.0%
🔴 functions/__server.func                         [########################] 9.47 MB 50.0%

Top Function Payloads

🟠 functions/.well-known/workflow/v1/flow.func • 1 public route • 9.47 MB
Metric Value
Public routes /.well-known/workflow/v1/flow
Runtime nodejs24.x
Handler index.mjs
Payload 9.47 MB
Function files 9.47 MB across 111 files
Traced dependencies 0 B
Signal 🟠 Bundled file _chunks/vercel.web.mjs is 2.09 MB (22.0%)

🟠 🔎 Dependency Analysis

📦 Bundled files:

Bundled file size
🟠 _chunks/vercel.web.mjs                         [############............] 2.09 MB 22.0%
🟡 _libs/undici.mjs                               [######..................] 980.8 kB 10.4%
🟡 _chunks/sandbox.mjs                            [#####...................] 811.4 kB 8.6%
🟡 _chunks/compiled-artifacts-instrumentation.mjs [####....................] 757.3 kB 8.0%
🟡 _chunks/signal-exit-DZKTacNU.mjs               [####....................] 616.2 kB 6.5%
🔴 Other bundled files                            [########################] 4.22 MB 44.6%

🧾 Vercel Config

{
  "handler": "index.mjs",
  "launcherType": "Nodejs",
  "shouldAddHelpers": false,
  "supportsResponseStreaming": true,
  "runtime": "nodejs24.x",
  "maxDuration": "max",
  "experimentalTriggers": [
    {
      "type": "queue/v2beta",
      "topic": "__eve776561746865722d6167656e74_wkf_workflow_*",
      "consumer": "default",
      "retryAfterSeconds": 5,
      "initialDelaySeconds": 0
    }
  ],
  "environment": {
    "WORKFLOW_PRECONDITION_GUARD": "1"
  }
}

🟠 functions/__server.func • 16 public routes, 1 internal alias • 9.47 MB
Metric Value
Public routes /
/.well-known/workflow/v1/webhook/[token]
/eve/v1/activity/[token]
/eve/v1/callback/[token]
/eve/v1/connections/[name]/callback/[attemptId]/[token]
/eve/v1/connections/[name]/callback/[token]
/eve/v1/health
/eve/v1/info
/eve/v1/session
/eve/v1/session/[parentSessionId]/subagents/[callId]/[childSessionId]/stream
/eve/v1/session/[sessionId]
/eve/v1/session/[sessionId]/cancel
/eve/v1/session/[sessionId]/clear
/eve/v1/session/[sessionId]/compact
/eve/v1/session/[sessionId]/reset
/eve/v1/session/[sessionId]/stream
Internal aliases /__server
Runtime nodejs24.x
Handler index.mjs
Payload 9.47 MB
Function files 9.47 MB across 111 files
Traced dependencies 0 B
Signal 🟠 Bundled file _chunks/vercel.web.mjs is 2.09 MB (22.0%)

🟠 🔎 Dependency Analysis

📦 Bundled files:

Bundled file size
🟠 _chunks/vercel.web.mjs                         [############............] 2.09 MB 22.0%
🟡 _libs/undici.mjs                               [######..................] 980.8 kB 10.4%
🟡 _chunks/sandbox.mjs                            [#####...................] 811.4 kB 8.6%
🟡 _chunks/compiled-artifacts-instrumentation.mjs [####....................] 757.3 kB 8.0%
🟡 _chunks/signal-exit-DZKTacNU.mjs               [####....................] 616.2 kB 6.5%
🔴 Other bundled files                            [########################] 4.22 MB 44.6%

🧾 Vercel Config

{
  "handler": "index.mjs",
  "launcherType": "Nodejs",
  "shouldAddHelpers": false,
  "supportsResponseStreaming": true,
  "runtime": "nodejs24.x"
}

Build Timing: e2e/fixtures/agent-tools-sandbox

This is an informational timing measurement inside eve build, from preflight through publication. Output-size measurement and profile writing are excluded.

Build mode: deployable Vercel build with sandbox template prewarm included.

  • Build pipeline: 3.80 s -> 3.78 s (-24.2 ms) vs barba/tasks-9-slack-wait-post (fbbac65).
  • Timing is informational: shared GitHub runners are too variable for a hard timing budget.
Detailed phase timings vs `barba/tasks-9-slack-wait-post (fbbac65)`
Phase Baseline Current Delta
extension.check 17.8 ms 9.3 ms -8.5 ms
project.resolve 5.7 ms 1.3 ms -4.4 ms
workspace.create 0.8 ms 1.9 ms +1.1 ms
host.prepare 692.0 ms 705.3 ms +13.3 ms
vercel.service-prefix.resolve 1.9 ms 2.0 ms +0.1 ms
nitro.create 283.4 ms 285.7 ms +2.3 ms
sandbox.prewarm 279.3 ms 268.9 ms -10.4 ms
nitro.cache.prepare 0.3 ms 0.3 ms 0.0 ms
nitro.prepare 0.9 ms 0.6 ms -0.3 ms
nitro.public-assets 0.8 ms 0.8 ms 0.0 ms
nitro.prerender 0.5 ms 0.6 ms +0.1 ms
nitro.bundle 2.37 s 2.35 s -23.5 ms
nitro.cache.write 0.4 ms 0.4 ms 0.0 ms
vercel.workflow-function.materialize 49.5 ms 51.5 ms +2.0 ms
agent-summary.emit 0.7 ms 0.6 ms -0.1 ms
connect-manifest.emit 0.4 ms 0.3 ms -0.1 ms
nitro.close 0.2 ms 0.1 ms -0.1 ms
output.publish 3.8 ms 3.5 ms -0.3 ms
workspace.remove 2.8 ms 2.6 ms -0.2 ms

Comment thread docs/tools/tasks.md Outdated
turn.expectOk();

const started = taskStarts(turn.events, "compile_report");
t.check(started.length, equals(2)).label("each report is compiled once");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This only counts two compile_report starts. task.started doesn't include the input, so two calls with topic: "churn" pass here. They also pass the report-id check at line 35: both settlements return the same RPT-… id, and the reply only has to contain it once. Since this is the release gate for fan-out, assert on the inputs from turn.toolCalls: one compile_report call per topic (churn, latency), or at least two distinct reportIds.

Comment thread e2e/fixtures/agent-workflow-tools/evals/held-turn.mcp.eval.ts
Comment thread e2e/fixtures/agent-workflow-tools/evals/mcp-client.ts
Comment thread docs/tools/tasks.md Outdated
@AndrewBarba
AndrewBarba force-pushed the barba/tasks-10-release-readiness branch from dc12872 to 4be6c15 Compare September 26, 2026 20:25
@AndrewBarba
AndrewBarba force-pushed the barba/tasks-10-release-readiness branch from 4be6c15 to 8134f8f Compare September 26, 2026 20:32
@AndrewBarba
AndrewBarba force-pushed the barba/tasks-10-release-readiness branch from 8134f8f to 433ce3b Compare September 27, 2026 00:35
@AndrewBarba
AndrewBarba force-pushed the barba/tasks-10-release-readiness branch from 433ce3b to 2b6d85a Compare September 27, 2026 00:55
The first real-model run of the task planning evals showed a model calling
task_wait after being told there was no rush, and a question asked while tasks
worked going unanswered until the tasks finished. The task system block now
says the model can reply without waiting, since its turn stays open and eve
delivers the results, and both the system block and task_wait's interrupt text
tell it to answer a new message that asks something. The text matches the
plan's quotes in research/eve-tasks.md.

Signed-off-by: Andrew Barba <barba@hey.com>
… turns

Mock-model suites in agent-workflow-tools for the plan's section 10: steering
stops sleep but not an approval-gated tool, and a question asked with
ctx.interruptSignal is withdrawn as cancelled; a task() tool starts, is waited
on, and delivers its result, and task_cancel stops one; a serve() tool is
continued by taskId, cancelled, and continued again with its state intact
because its body catches the stretch's abort; and the turn rule holds in a
root session, a child session, a schedule's session, and MCP agent_start.
Each held turn parks with turn.waiting, and a held root turn's send() result
is the final reply, not the text written before the wait. A task's question
during a hold emits input.requested, then turn.waiting, stops send() while it
is pending, and the answered turn completes once.

The scripted mock-model helper moves to @eve-e2e/config/mock-script so
fixtures share it.

Signed-off-by: Andrew Barba <barba@hey.com>
A local and a remote notebook keeper are continued by taskId across turns and
after task_cancel with their conversation intact, and corrected by taskId while
their tool runs. The correction joins the keeper's running first turn, which
reads it before completing and settles the correction's call with the
corrected reply.

Signed-off-by: Andrew Barba <barba@hey.com>
…tool

With the task system block in every root agent's prompt, a model read "Wait
for its result" as a task_wait call and spent its one remaining model call
after the approval on it. The prompt now says to reply when the tool returns;
the approval flow it gates is unchanged.

Signed-off-by: Andrew Barba <barba@hey.com>
A new agent-tasks fixture gates the release on how live models plan around
tasks: wait when the answer is needed; fan out, then wait for every result;
keep tasks after an unrelated message and answer it; correct an agent by
taskId or cancel it after a redirect; continue an agent instead of starting a
new one; don't wait when told there's no rush; and don't use a side-effect
tool before the result it depends on. Every eval is tagged real-model.

Signed-off-by: Andrew Barba <barba@hey.com>
The Tasks guide covers choosing execute(), task(), or serve(), ordering work
that depends on an agent's result, the signals a body receives, what the model
sees, held turns and turn.waiting, cancellation, events, and limits. The
upgrade guide covers execution: "background" becoming task(), dismissible
becoming ctx.interruptSignal, ctx.agent session handles, agentId becoming
taskId, task events replacing subagent.*, the removed delivery options,
turn.waiting and questions inside a call no longer ending the turn, reading
task outcomes from task.settled, stream version 26, and remote agent protocol
version 2. The workflow tool, built-in tool, streaming, and subagent pages
link to the guide instead of repeating it.

Signed-off-by: Andrew Barba <barba@hey.com>
A workflow tool run used to resolve an aborted ask as cancelled on its own,
while the session could still accept a person's answer for it, so
ask_question could report { interrupted: true } after the channel showed the
question answered. Answers also reached the run on a per-ask hook, separate
from the control hook that carries calls, interrupts, and cancels.

The session is now the one authority. It accepts an answer or a withdrawal
at the step that retires the question's route, emits input.resolved there,
and sends its decision (answer or withdrawn) on the run's control hook, the
run's one ordered inbox. An aborted signal only asks the session to
withdraw. cancel of an execute or task() call and end settle pending asks as
cancelled, because the session retires their routes when it sends them;
task_cancel of a task() run now reports its questions cancelled. A serve()
task's cancel withdraws its stretch's questions through the same round trip.

Question requestIds become <runId>-ask-<n>. The agent-workflow-tools fixture
gains a serve() tool whose question is withdrawn by task_cancel while the
task keeps serving.

Signed-off-by: Andrew Barba <barba@hey.com>
The eval follows the session with watchTurn, whose attached session repeats the turn's events in the run's event list, so task.started appeared more than once per call. Count distinct callIds and taskIds instead.

Signed-off-by: Andrew Barba <barba@hey.com>
In the agent-tasks real-model eval, Opus often called task_wait without text after an unrelated question arrived while a task worked, then answered only the task result, so the question was lost. The system block already said to answer such a message; the task_wait description and the system block now say to answer it in the same response, before waiting.

Signed-off-by: Andrew Barba <barba@hey.com>
…rrive

The agent-tasks eval still failed on Opus with the system block alone: the model waits for the result it was asked for, then answers only that result, and the question it put off sits several messages back. A task_wait that returns results now ends with a reminder to answer any message the model has not answered yet. The earlier task_wait description sentence is dropped as redundant with the system block.

Signed-off-by: Andrew Barba <barba@hey.com>

This branch was successfully deployed

2 active deployments
Preview – eve-pkg — 5d2bbf2d Deployed Sep 28, 2026 by vercel[bot]
Preview – eve-docs — 5d2bbf2d Deployed Sep 28, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant