test(e2e): cover tasks and add the Tasks and upgrade guides - #3859
AndrewBarba wants to merge 10 commits into
Conversation
Bundle + Package Summary:
|
| Area | Metric | Baseline | Current | Delta |
|---|---|---|---|---|
| Package | Packed tarball | 8.39 MB | 8.40 MB | +8.6 kB |
| Package | Unpacked publish size | 30.82 MB | 30.85 MB | +31.3 kB |
| Package | Installed footprint | 77.19 MB | 77.22 MB | +31.3 kB |
| Package | Published files | 3848 | 3850 | +2 |
| Package | Installed files | 7815 | 7817 | +2 |
| Package | Installed package instances | 33 | 33 | 0 |
| Package | Distinct installed package names | 32 | 32 | 0 |
| Package | Installed dependency edges | 51 | 51 | 0 |
| Package | Installed optional peer edges | 9 | 9 | 0 |
| Runtime | Unique function payloads | 2 | 2 | 0 |
| Runtime | Total function bytes | 18.92 MB | 18.94 MB | +27.4 kB |
| Runtime | Public routes | 17 | 17 | 0 |
Changed function payloads vs barba/tasks-9-slack-wait-post (fbbac65) (2)
| Function | Status | Baseline | Current | Delta | Route changes |
|---|---|---|---|---|---|
functions/__server.func |
changed | 9.46 MB | 9.47 MB | +13.7 kB |
none |
functions/.well-known/workflow/v1/flow.func |
changed | 9.46 MB | 9.47 MB | +13.7 kB |
none |
eve init install
| Metric | Baseline | Current | Delta |
|---|---|---|---|
| Installed footprint | 115.48 MB | 115.51 MB | +31.3 kB |
| Installed packages | 97 | 97 | 0 |
| dependencies | 4 | 4 | 0 |
| devDependencies | 2 | 2 | 0 |
| Dependency package bytes | 48.21 MB | 48.25 MB | +31.3 kB |
| devDependency package bytes | 5.11 MB | 5.11 MB | 0 B ➖ |
Build Metadata
- Preset:
vercel - Nitro:
nitro@3.0.260903-beta - Output directory:
apps/fixtures/weather-agent/.vercel/output - Build metadata timestamp: 2026-09-28T00:27:18.682Z
- Route aliases: 17 public, 1 internal (18 total aliases)
- Vercel routes in config: 20
- Severity legend: 🔴 dominant/large, 🟠 notable, 🟡 watch, ⚪ small
Package Drill-Down
Package Details
- Package:
eve@0.67.2 - Package directory:
packages/eve - Tarball: 8.40 MB (
eve-0.67.2.tgz) - Unpacked payload: 30.85 MB across 3850 published files
- Installed footprint: 77.22 MB across 7817 installed files
- Installed root package: 30.57 MB
- Installed dependencies: 46.65 MB
- Installed package instances: 33
- Distinct installed package names: 32
- Installed dependency edges: 51
- Installed optional peer edges: 9
- Runtime dependencies: 2
- Peer dependencies: 7 (6 optional)
Installed footprint is measured from an isolated temporary npm install of the packed tarball.
Graph metrics read only package.json files in package directories directly beneath a node_modules boundary, including nested boundaries. Each directory is one package instance; distinct names come from those manifests. Dependency edges count each unique name in dependencies or optionalDependencies per instance; optional peer edges count peerDependencies marked optional.
Heavy installed dependencies
eve: 30.57 MB (39.6%)@rolldown/binding-linux-x64-gnu: 19.15 MB (24.8%)ai: 7.73 MB (10.0%)zod: 6.14 MB (8.0%)undici: 3.52 MB (4.6%)
Publish payload breakdown
Published file size
🔴 dist/src/compiled/shadcn-registry/index.js [##############..........] 9.70 MB 31.4%
🟠 dist/src/compiled/@photon-ai/chat-adapter-ime... [###.....................] 2.29 MB 7.4%
🟠 dist/src/compiled/@ai-sdk/code-mode/index.js [#.......................] 1.03 MB 3.3%
🟡 dist/src/compiled/@vercel/blob/index.js [#.......................] 604.9 kB 2.0%
🟡 dist/src/compiled/_chunks/workflow/signal-exi... [#.......................] 514.5 kB 1.7%
🔴 Other published files [########################] 16.72 MB 54.2%
Installed footprint breakdown
Installed package size
🔴 eve [########################] 30.57 MB 39.6%
🔴 @rolldown/binding-linux-x64-gnu [###############.........] 19.15 MB 24.8%
🔴 ai [######..................] 7.73 MB 10.0%
🔴 zod [#####...................] 6.14 MB 8.0%
🟠 undici [###.....................] 3.52 MB 4.6%
🟠 nitro [#.......................] 1.89 MB 2.5%
🔴 Other installed packages [######..................] 8.22 MB 10.6%
Runtime dependencies (2)
| Package | Range | Notes |
|---|---|---|
nitro |
3.0.260903-beta |
|
undici |
8.9.0 |
Peer dependencies (7)
| Package | Range | Notes |
|---|---|---|
@opentelemetry/api |
^1.0.0 |
optional peer |
ai |
catalog: |
|
braintrust |
^3.0.0 |
optional peer |
chat |
^4.41.0 |
optional peer |
dd-trace |
^6.13.0 |
optional peer |
just-bash |
^3.1.0 |
optional peer |
microsandbox |
^0.5.0 |
optional peer |
eve init install drill-down
eve init install details
- Command:
eve init my-agent - Package manager:
npm - Installed footprint: 115.51 MB across 9712 installed files
- Installed packages: 97 total (91 transitive-only)
- dependencies: 4 direct packages totaling 48.25 MB
- devDependencies: 2 direct packages totaling 5.11 MB
- Other transitive package files: 62.16 MB
Installed footprint is measured from an isolated temporary eve init my-agent using the current packed eve tarball.
Heavy installed dependencies
eve: 30.57 MB (26.5%)@typescript/typescript-linux-x64: 27.95 MB (24.2%)@rolldown/binding-linux-x64-gnu: 19.15 MB (16.6%)zod: 9.76 MB (8.4%)ai: 7.73 MB (6.7%)
Installed footprint breakdown
Installed package size
🔴 eve [########################] 30.57 MB 26.5%
🔴 @typescript/typescript-linux-x64 [######################..] 27.95 MB 24.2%
🔴 @rolldown/binding-linux-x64-gnu [###############.........] 19.15 MB 16.6%
🔴 zod [########................] 9.76 MB 8.4%
🔴 ai [######..................] 7.73 MB 6.7%
🟠 undici [###.....................] 3.52 MB 3.0%
🔴 Other installed packages [#############...........] 16.85 MB 14.6%
dependencies (4)
| Package | Range | Installed size | Share |
|---|---|---|---|
@vercel/connect |
2.2.0 |
194.7 kB | 0.2% |
ai |
^7.0.105 |
7.73 MB | 6.7% |
eve |
file:eve-0.67.2.tgz |
30.57 MB | 26.5% |
zod |
4.5.4 |
9.76 MB | 8.4% |
devDependencies (2)
| Package | Range | Installed size | Share |
|---|---|---|---|
@types/node |
24.x |
2.61 MB | 2.3% |
typescript |
7.0.2 |
2.50 MB | 2.2% |
Function Drill-Down
Payload Size Graph
Unique function payload size and share of total
🔴 functions/.well-known/workflow/v1/flow.func [########################] 9.47 MB 50.0%
🔴 functions/__server.func [########################] 9.47 MB 50.0%
Top Function Payloads
🟠 functions/.well-known/workflow/v1/flow.func • 1 public route • 9.47 MB
| Metric | Value |
|---|---|
| Public routes | /.well-known/workflow/v1/flow |
| Runtime | nodejs24.x |
| Handler | index.mjs |
| Payload | 9.47 MB |
| Function files | 9.47 MB across 111 files |
| Traced dependencies | 0 B |
| Signal | 🟠 Bundled file _chunks/vercel.web.mjs is 2.09 MB (22.0%) |
🟠 🔎 Dependency Analysis
📦 Bundled files:
Bundled file size
🟠 _chunks/vercel.web.mjs [############............] 2.09 MB 22.0%
🟡 _libs/undici.mjs [######..................] 980.8 kB 10.4%
🟡 _chunks/sandbox.mjs [#####...................] 811.4 kB 8.6%
🟡 _chunks/compiled-artifacts-instrumentation.mjs [####....................] 757.3 kB 8.0%
🟡 _chunks/signal-exit-DZKTacNU.mjs [####....................] 616.2 kB 6.5%
🔴 Other bundled files [########################] 4.22 MB 44.6%
🧾 Vercel Config
{
"handler": "index.mjs",
"launcherType": "Nodejs",
"shouldAddHelpers": false,
"supportsResponseStreaming": true,
"runtime": "nodejs24.x",
"maxDuration": "max",
"experimentalTriggers": [
{
"type": "queue/v2beta",
"topic": "__eve776561746865722d6167656e74_wkf_workflow_*",
"consumer": "default",
"retryAfterSeconds": 5,
"initialDelaySeconds": 0
}
],
"environment": {
"WORKFLOW_PRECONDITION_GUARD": "1"
}
}🟠 functions/__server.func • 16 public routes, 1 internal alias • 9.47 MB
| Metric | Value |
|---|---|
| Public routes | //.well-known/workflow/v1/webhook/[token]/eve/v1/activity/[token]/eve/v1/callback/[token]/eve/v1/connections/[name]/callback/[attemptId]/[token]/eve/v1/connections/[name]/callback/[token]/eve/v1/health/eve/v1/info/eve/v1/session/eve/v1/session/[parentSessionId]/subagents/[callId]/[childSessionId]/stream/eve/v1/session/[sessionId]/eve/v1/session/[sessionId]/cancel/eve/v1/session/[sessionId]/clear/eve/v1/session/[sessionId]/compact/eve/v1/session/[sessionId]/reset/eve/v1/session/[sessionId]/stream |
| Internal aliases | /__server |
| Runtime | nodejs24.x |
| Handler | index.mjs |
| Payload | 9.47 MB |
| Function files | 9.47 MB across 111 files |
| Traced dependencies | 0 B |
| Signal | 🟠 Bundled file _chunks/vercel.web.mjs is 2.09 MB (22.0%) |
🟠 🔎 Dependency Analysis
📦 Bundled files:
Bundled file size
🟠 _chunks/vercel.web.mjs [############............] 2.09 MB 22.0%
🟡 _libs/undici.mjs [######..................] 980.8 kB 10.4%
🟡 _chunks/sandbox.mjs [#####...................] 811.4 kB 8.6%
🟡 _chunks/compiled-artifacts-instrumentation.mjs [####....................] 757.3 kB 8.0%
🟡 _chunks/signal-exit-DZKTacNU.mjs [####....................] 616.2 kB 6.5%
🔴 Other bundled files [########################] 4.22 MB 44.6%
🧾 Vercel Config
{
"handler": "index.mjs",
"launcherType": "Nodejs",
"shouldAddHelpers": false,
"supportsResponseStreaming": true,
"runtime": "nodejs24.x"
}Build Timing: e2e/fixtures/agent-tools-sandbox
This is an informational timing measurement inside eve build, from preflight through publication. Output-size measurement and profile writing are excluded.
Build mode: deployable Vercel build with sandbox template prewarm included.
- Build pipeline: 3.80 s -> 3.78 s (-24.2 ms) vs
barba/tasks-9-slack-wait-post (fbbac65). - Timing is informational: shared GitHub runners are too variable for a hard timing budget.
Detailed phase timings vs `barba/tasks-9-slack-wait-post (fbbac65)`
| Phase | Baseline | Current | Delta |
|---|---|---|---|
extension.check |
17.8 ms | 9.3 ms | -8.5 ms |
project.resolve |
5.7 ms | 1.3 ms | -4.4 ms |
workspace.create |
0.8 ms | 1.9 ms | +1.1 ms |
host.prepare |
692.0 ms | 705.3 ms | +13.3 ms |
vercel.service-prefix.resolve |
1.9 ms | 2.0 ms | +0.1 ms |
nitro.create |
283.4 ms | 285.7 ms | +2.3 ms |
sandbox.prewarm |
279.3 ms | 268.9 ms | -10.4 ms |
nitro.cache.prepare |
0.3 ms | 0.3 ms | 0.0 ms |
nitro.prepare |
0.9 ms | 0.6 ms | -0.3 ms |
nitro.public-assets |
0.8 ms | 0.8 ms | 0.0 ms |
nitro.prerender |
0.5 ms | 0.6 ms | +0.1 ms |
nitro.bundle |
2.37 s | 2.35 s | -23.5 ms |
nitro.cache.write |
0.4 ms | 0.4 ms | 0.0 ms |
vercel.workflow-function.materialize |
49.5 ms | 51.5 ms | +2.0 ms |
agent-summary.emit |
0.7 ms | 0.6 ms | -0.1 ms |
connect-manifest.emit |
0.4 ms | 0.3 ms | -0.1 ms |
nitro.close |
0.2 ms | 0.1 ms | -0.1 ms |
output.publish |
3.8 ms | 3.5 ms | -0.3 ms |
workspace.remove |
2.8 ms | 2.6 ms | -0.2 ms |
| turn.expectOk(); | ||
|
|
||
| const started = taskStarts(turn.events, "compile_report"); | ||
| t.check(started.length, equals(2)).label("each report is compiled once"); |
There was a problem hiding this comment.
This only counts two compile_report starts. task.started doesn't include the input, so two calls with topic: "churn" pass here. They also pass the report-id check at line 35: both settlements return the same RPT-… id, and the reply only has to contain it once. Since this is the release gate for fan-out, assert on the inputs from turn.toolCalls: one compile_report call per topic (churn, latency), or at least two distinct reportIds.
f50c43f to
46ecf6d
Compare
46ecf6d to
f18fdea
Compare
f18fdea to
c0b9e23
Compare
dc12872 to
4be6c15
Compare
4be6c15 to
8134f8f
Compare
8134f8f to
433ce3b
Compare
433ce3b to
2b6d85a
Compare
The first real-model run of the task planning evals showed a model calling task_wait after being told there was no rush, and a question asked while tasks worked going unanswered until the tasks finished. The task system block now says the model can reply without waiting, since its turn stays open and eve delivers the results, and both the system block and task_wait's interrupt text tell it to answer a new message that asks something. The text matches the plan's quotes in research/eve-tasks.md. Signed-off-by: Andrew Barba <barba@hey.com>
… turns Mock-model suites in agent-workflow-tools for the plan's section 10: steering stops sleep but not an approval-gated tool, and a question asked with ctx.interruptSignal is withdrawn as cancelled; a task() tool starts, is waited on, and delivers its result, and task_cancel stops one; a serve() tool is continued by taskId, cancelled, and continued again with its state intact because its body catches the stretch's abort; and the turn rule holds in a root session, a child session, a schedule's session, and MCP agent_start. Each held turn parks with turn.waiting, and a held root turn's send() result is the final reply, not the text written before the wait. A task's question during a hold emits input.requested, then turn.waiting, stops send() while it is pending, and the answered turn completes once. The scripted mock-model helper moves to @eve-e2e/config/mock-script so fixtures share it. Signed-off-by: Andrew Barba <barba@hey.com>
A local and a remote notebook keeper are continued by taskId across turns and after task_cancel with their conversation intact, and corrected by taskId while their tool runs. The correction joins the keeper's running first turn, which reads it before completing and settles the correction's call with the corrected reply. Signed-off-by: Andrew Barba <barba@hey.com>
…tool With the task system block in every root agent's prompt, a model read "Wait for its result" as a task_wait call and spent its one remaining model call after the approval on it. The prompt now says to reply when the tool returns; the approval flow it gates is unchanged. Signed-off-by: Andrew Barba <barba@hey.com>
A new agent-tasks fixture gates the release on how live models plan around tasks: wait when the answer is needed; fan out, then wait for every result; keep tasks after an unrelated message and answer it; correct an agent by taskId or cancel it after a redirect; continue an agent instead of starting a new one; don't wait when told there's no rush; and don't use a side-effect tool before the result it depends on. Every eval is tagged real-model. Signed-off-by: Andrew Barba <barba@hey.com>
The Tasks guide covers choosing execute(), task(), or serve(), ordering work that depends on an agent's result, the signals a body receives, what the model sees, held turns and turn.waiting, cancellation, events, and limits. The upgrade guide covers execution: "background" becoming task(), dismissible becoming ctx.interruptSignal, ctx.agent session handles, agentId becoming taskId, task events replacing subagent.*, the removed delivery options, turn.waiting and questions inside a call no longer ending the turn, reading task outcomes from task.settled, stream version 26, and remote agent protocol version 2. The workflow tool, built-in tool, streaming, and subagent pages link to the guide instead of repeating it. Signed-off-by: Andrew Barba <barba@hey.com>
A workflow tool run used to resolve an aborted ask as cancelled on its own,
while the session could still accept a person's answer for it, so
ask_question could report { interrupted: true } after the channel showed the
question answered. Answers also reached the run on a per-ask hook, separate
from the control hook that carries calls, interrupts, and cancels.
The session is now the one authority. It accepts an answer or a withdrawal
at the step that retires the question's route, emits input.resolved there,
and sends its decision (answer or withdrawn) on the run's control hook, the
run's one ordered inbox. An aborted signal only asks the session to
withdraw. cancel of an execute or task() call and end settle pending asks as
cancelled, because the session retires their routes when it sends them;
task_cancel of a task() run now reports its questions cancelled. A serve()
task's cancel withdraws its stretch's questions through the same round trip.
Question requestIds become <runId>-ask-<n>. The agent-workflow-tools fixture
gains a serve() tool whose question is withdrawn by task_cancel while the
task keeps serving.
Signed-off-by: Andrew Barba <barba@hey.com>
The eval follows the session with watchTurn, whose attached session repeats the turn's events in the run's event list, so task.started appeared more than once per call. Count distinct callIds and taskIds instead. Signed-off-by: Andrew Barba <barba@hey.com>
In the agent-tasks real-model eval, Opus often called task_wait without text after an unrelated question arrived while a task worked, then answered only the task result, so the question was lost. The system block already said to answer such a message; the task_wait description and the system block now say to answer it in the same response, before waiting. Signed-off-by: Andrew Barba <barba@hey.com>
…rrive The agent-tasks eval still failed on Opus with the system block alone: the model waits for the result it was asked for, then answers only that result, and the question it put off sits several messages back. A task_wait that returns results now ends with a reminder to answer any message the model has not answered yet. The earlier task_wait description sentence is dropped as redundant with the system block. Signed-off-by: Andrew Barba <barba@hey.com>
Summary
PRs 1–9 of #3821 landed the task model with a passing build but, by design, almost no new tests. This PR adds the coverage and docs the plan requires before release.
sleep, approvals, and withdrawn asks;task()start, wait, and cancel; the turn rule in root, child, schedule, and MCP sessions; aservetool continued, cancelled, and continued again; local and remote agents continued and corrected bytaskId.agent-tasksfixture gates the release on task planning. The model waits when needed, fans out then waits, keeps tasks after an unrelated message, redirects or cancels an agent, continues instead of restarting, doesn't wait when told there's no rush, and doesn't run a side effect before its result.execute,task, andserve.task_waitinterrupt text are tuned against the real-model evals, and the plan quotes the tuned text.ctx.askascancelledwhile the session still accepted a person's answer, soask_questionand the channel disagreed (#3845 review). The session now decides each question once and sends the decision on the run's control hook, with its calls, interrupts, and cancels (#3821 review).task_cancelof atask()reports its questionscancelled, and questionrequestIds become<runId>-ask-<n>.This PR builds on #3857 (related to #1084). Slack's post before a wait keeps its unit test from #3857, since there is no Slack e2e harness.
Validation
pnpm build,pnpm typecheck,pnpm lint,pnpm guard:invariants,pnpm guard:fixtures,pnpm check:deps,pnpm check:templates,pnpm docs:check,pnpm test:unitandpnpm test:integrationall pass.pnpm test:scenariopasses, apart from the knownhostless-agent-workspacenpm 401, which passed when rerun alone.agent-workflow-tools26/26 andagent-subagents9/9. Theagent-tasksreal-model evals typecheck and load, and e2e-local CI runs them.pnpm typecheck,pnpm lint,pnpm fmt,pnpm guard:invariants,pnpm docs:check, the focused workflow, session, task, and HITL unit tests (146) and integration tests (23), and theagent-workflow-toolsbuild and typecheck. Its newsign-off-plan.canceleval has not run yet; CI runs it.Checklist
CONTRIBUTING.mdevepackagegit commit --signoff)