[pull] main from langwatch:main#279
Open
pull[bot] wants to merge 2484 commits into
Open
Conversation
rogeriochaves
force-pushed
the
main
branch
5 times, most recently
from
January 21, 2026 01:15
1e7b14c to
2209258
Compare
…4738) Bumps [uvicorn](https://github.com/Kludex/uvicorn) from 0.38.0 to 0.49.0. - [Release notes](https://github.com/Kludex/uvicorn/releases) - [Changelog](https://github.com/Kludex/uvicorn/blob/main/docs/release-notes.md) - [Commits](Kludex/uvicorn@0.38.0...0.49.0) --- updated-dependencies: - dependency-name: uvicorn dependency-version: 0.49.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…-server (#4880) chore(deps-dev): bump typescript-eslint in /mcp-server Bumps [typescript-eslint](https://github.com/typescript-eslint/typescript-eslint/tree/HEAD/packages/typescript-eslint) from 8.60.0 to 8.62.0. - [Release notes](https://github.com/typescript-eslint/typescript-eslint/releases) - [Changelog](https://github.com/typescript-eslint/typescript-eslint/blob/main/packages/typescript-eslint/CHANGELOG.md) - [Commits](https://github.com/typescript-eslint/typescript-eslint/commits/v8.62.0/packages/typescript-eslint) --- updated-dependencies: - dependency-name: typescript-eslint dependency-version: 8.61.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…pescript-sdk (#4773) chore(deps-dev): bump vitest-mock-extended in /typescript-sdk Bumps [vitest-mock-extended](https://github.com/eratio08/vitest-mock-extended) from 3.1.1 to 4.0.0. - [Release notes](https://github.com/eratio08/vitest-mock-extended/releases) - [Commits](eratio08/vitest-mock-extended@v3.1.1...v4.0.0) --- updated-dependencies: - dependency-name: vitest-mock-extended dependency-version: 4.0.0 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [immer](https://github.com/immerjs/immer) from 10.2.0 to 11.1.8. - [Release notes](https://github.com/immerjs/immer/releases) - [Commits](immerjs/immer@v10.2.0...v11.1.8) --- updated-dependencies: - dependency-name: immer dependency-version: 11.1.8 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…0 in /typescript-sdk (#4316) chore(deps-dev): bump @opentelemetry/sdk-trace-web in /typescript-sdk Bumps [@opentelemetry/sdk-trace-web](https://github.com/open-telemetry/opentelemetry-js) from 2.6.0 to 2.8.0. - [Release notes](https://github.com/open-telemetry/opentelemetry-js/releases) - [Changelog](https://github.com/open-telemetry/opentelemetry-js/blob/main/CHANGELOG.md) - [Commits](open-telemetry/opentelemetry-js@v2.6.0...v2.8.0) --- updated-dependencies: - dependency-name: "@opentelemetry/sdk-trace-web" dependency-version: 2.7.1 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [pnpm/action-setup](https://github.com/pnpm/action-setup) from 4.3.0 to 6.0.9. - [Release notes](https://github.com/pnpm/action-setup/releases) - [Commits](pnpm/action-setup@v4.3.0...0ebf471) --- updated-dependencies: - dependency-name: pnpm/action-setup dependency-version: 6.0.8 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [coolname](https://github.com/alexanderlukanin13/coolname) from 2.2.0 to 5.0.0. - [Changelog](https://github.com/alexanderlukanin13/coolname/blob/master/HISTORY.rst) - [Commits](alexanderlukanin13/coolname@2.2.0...5.0.0) --- updated-dependencies: - dependency-name: coolname dependency-version: 5.0.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [pnpm/action-setup](https://github.com/pnpm/action-setup) from 4.3.0 to 6.0.9. - [Release notes](https://github.com/pnpm/action-setup/releases) - [Commits](pnpm/action-setup@v4.3.0...0ebf471) --- updated-dependencies: - dependency-name: pnpm/action-setup dependency-version: 6.0.9 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [turbo](https://github.com/vercel/turborepo) from 2.9.14 to 2.10.0. - [Release notes](https://github.com/vercel/turborepo/releases) - [Changelog](https://github.com/vercel/turborepo/blob/main/RELEASE.md) - [Commits](vercel/turborepo@v2.9.14...v2.10.0) --- updated-dependencies: - dependency-name: turbo dependency-version: 2.9.16 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [superjson](https://github.com/blitz-js/superjson) from 1.13.3 to 2.2.6. - [Release notes](https://github.com/blitz-js/superjson/releases) - [Commits](https://github.com/blitz-js/superjson/commits) --- updated-dependencies: - dependency-name: superjson dependency-version: 2.2.6 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…tch (#4601) Bumps [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) from 6.0.1 to 6.0.3. - [Release notes](https://github.com/vitejs/vite-plugin-react/releases) - [Changelog](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@6.0.3/packages/plugin-react) --- updated-dependencies: - dependency-name: "@vitejs/plugin-react" dependency-version: 6.0.2 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…n /python-sdk (#4918) chore(deps-dev): bump langchain-google-vertexai in /python-sdk Bumps [langchain-google-vertexai](https://github.com/langchain-ai/langchain-google) from 3.2.3 to 3.2.4. - [Release notes](https://github.com/langchain-ai/langchain-google/releases) - [Commits](langchain-ai/langchain-google@libs/vertexai/v3.2.3...libs/vertexai/v3.2.4) --- updated-dependencies: - dependency-name: langchain-google-vertexai dependency-version: 3.2.4 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
#5348) * perf(ui): actually defer GTM/Pendo/Crisp instead of injecting on mount next-script.tsx accepted a `strategy` prop but ignored it entirely, injecting every third-party <script> synchronously the instant its component mounted. Fix the shim to honor strategy: beforeInteractive injects immediately, everything else waits for requestIdleCallback (with a window-load/timeout fallback for Safari). Part of #5330 (Tier 1) · parent #5329 · closes #5346 * test(compat): expand next-script coverage to 100% Adds coverage for the src-script path (async, onLoad/onError wiring), non-string children, missing id, arbitrary attribute pass-through, dangerouslySetInnerHTML exclusion, the already-complete-document idle fallback branch, the requestIdleCallback timeout arg, and StrictMode's double-effect-invoke not double-injecting the script. * fix(lint): satisfy biome import-order and JSX line-length rules CI's lint job flagged an unsorted import and single-line JSX that exceeded the line-length rule in the expanded test coverage. * fix(compat): cancel pending script injection on unmount, address review - next-script.tsx: runWhenIdle now returns a cancel function, wired to the effect's cleanup. Without this, a deferred script scheduled via requestIdleCallback/setTimeout/window-load could still fire after the component unmounts — e.g. if SignedInExtraFooterComponents unmounts before idle (logout, impersonation flip, fast navigation away), the scheduled Pendo/Crisp snippet could inject with stale per-user data. (CodeRabbit review, P2) - Renamed test file .unit.test.tsx -> .integration.test.tsx: it renders the real Script component tree via @testing-library/react, so per project convention it's an integration test, not a unit test. - beforeEach now clears all <script> nodes, not just script[id] — the "given no id" case was leaking an untracked node across tests. - vi.useRealTimers() now runs unconditionally in afterEach so a failed assertion mid-test can't leave fake timers active for later tests. - Removed a dead requestIdleCallback stub in the StrictMode test (that path uses strategy="beforeInteractive", which never calls runWhenIdle). - Added regression coverage for the new cancellation behavior across all three defer paths (requestIdleCallback, Safari-fallback timeout, Safari-fallback load-listener), including the cancelIdleCallback- unavailable case. 100% statement/branch/function/line coverage maintained on next-script.tsx. * fix(compat): fix StrictMode remount regression, split oversized test file - next-script.tsx: the "already injected" guard was set at schedule time (before runWhenIdle) rather than at actual-injection time. React StrictMode's dev-only setup->cleanup->setup cycle cancels the pending deferred inject on the synthetic cleanup, then re-runs the effect — with the guard set early, that remount saw it already tripped and never rescheduled, so the script silently never injected under StrictMode for any deferred strategy (afterInteractive/lazyOnload/ unset). Fixed by moving the flag into inject() itself, only set once it actually runs. (CodeRabbit review) - Removed a redundant inner "already injected" check inside inject() itself — tracing through runWhenIdle's cancellation logic, inject() can never be invoked twice given the cancelled-flag guard already in place, so the extra check was unreachable defensive code. - Split next-script.integration.test.tsx (350 non-comment lines) into three focused files per the project's 300-line SRP convention (.coderabbit.yaml): timing (strategy/idle/cancellation), rendering (src/children/attributes output), and strictmode (double-invoke correctness). Added the StrictMode + deferred-strategy regression test the prior single-file suite was missing — this is exactly the case that would have caught the bug above. 100% coverage maintained on next-script.tsx (27 tests total, up from 26). * fix(lint): rename injected ref to hasInjectedRef for boolean-naming convention CodeRabbit flagged the newly introduced boolean-holding ref lacking an is/has prefix, per the repo's boolean-naming path instructions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(analytics): retry gtag/Reo instead of dropping when GTM is still idle-deferred The /review pass on this PR found a real regression: window.gtag and window.Reo are defined by GTM's container once it loads (not by the inline snippet this PR defers), but ExtraFooterComponents.tsx, analyticsClient.ts, and tracking.ts all checked for them once and gave up permanently if absent. Deferring GTM's own script widened that race from tens-of-ms to up to 4s (or until window.load on Safari), making the open_dashboard gtag event, Reo.identify, the analytics client's google provider, and trackEventOnce's one-shot events plausible to silently and permanently drop. - New shared src/utils/pollForGlobal.ts (bounded poll + cancel fn, matching the runWhenIdle/startSessionRecordingWhenIdle idiom already used in this perf initiative). - ExtraFooterComponents.tsx's Reo/gtag effects now poll instead of checking once. - tracking.ts's trackEvent/trackEventOnce poll for gtag; trackEventOnce no longer marks an event "sent" in localStorage until it actually sends, so a miss is retried instead of lost forever. - analyticsClient.ts takes an explicit isGtagReady flag instead of reading window.gtag once at construction time; new useIsGtagReady() hook (src/hooks/) polls and flips it, letting AppProviders re-render once GTM's container is ready. 100% coverage on all new/changed logic. Live-verified: GTM/Pendo/Crisp defer timing is unchanged (git-stash A/B against the pre-fix state), no new console errors introduced (confirmed a pre-existing "Maximum update depth exceeded" warning is unrelated), and window.Reo genuinely loads with an observable delay in a live signed-in session — proving the retry path exercises for real, not just in mocked tests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(analytics): type test mocks instead of casting through any /code-review flagged as any usages in analyticsClient.unit.test.ts. Typed the PostHog and Provider.send test fixtures properly instead. --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(billing): wire evaluation runs to billable events meter
The billable-events meter subscribed to lw.evaluation.scheduled and
lw.evaluation.started, neither of which is ever emitted in production.
The live evaluation pipeline only emits lw.evaluation.reported (via
reportEvaluation / ReportEvaluationCommand / ExecuteEvaluationCommand),
so evaluation executions contributed zero billable rows.
Subscribe the meter to lw.evaluation.reported instead and drop the two
dead subscriptions. The reported event's idempotencyKey is
${tenantId}:${evaluationId}:reported, so retries and replays collapse
to exactly one billable unit per evaluation.
Regression test drives the real ProjectionRouter eventType filter + map
+ store, asserting a billable record is produced for a reported event
and that scheduled/started are ignored — not a shape assertion.
Also extracts the duplicated createMockQueueManager test factory into
the shared testHelpers.
Fixes #5124
* test(billing): fix biome format + add BDD given nesting in evaluation regression test
- multi-line the spy store object spread (biome reviewdog lint failure)
- wrap the three when-blocks under a given describe (CodeRabbit BDD nesting)
#5137) Bumps [haystack-ai](https://github.com/deepset-ai/haystack) from 2.30.1 to 2.30.2. - [Release notes](https://github.com/deepset-ai/haystack/releases) - [Commits](deepset-ai/haystack@v2.30.1...v2.30.2) --- updated-dependencies: - dependency-name: haystack-ai dependency-version: 2.30.2 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
… /langwatch (#5058) chore(deps-dev): bump @testcontainers/redis in /langwatch Bumps [@testcontainers/redis](https://github.com/testcontainers/testcontainers-node) from 11.14.0 to 12.0.3. - [Release notes](https://github.com/testcontainers/testcontainers-node/releases) - [Commits](testcontainers/testcontainers-node@v11.14.0...v12.0.3) --- updated-dependencies: - dependency-name: "@testcontainers/redis" dependency-version: 12.0.2 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [motion](https://github.com/motiondivision/motion) from 12.38.0 to 12.42.0. - [Changelog](https://github.com/motiondivision/motion/blob/main/CHANGELOG.md) - [Commits](motiondivision/motion@v12.38.0...v12.42.0) --- updated-dependencies: - dependency-name: motion dependency-version: 12.40.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…4579) Bumps [react-markdown](https://github.com/remarkjs/react-markdown) from 9.1.0 to 10.1.0. - [Release notes](https://github.com/remarkjs/react-markdown/releases) - [Changelog](https://github.com/remarkjs/react-markdown/blob/main/changelog.md) - [Commits](remarkjs/react-markdown@9.1.0...10.1.0) --- updated-dependencies: - dependency-name: react-markdown dependency-version: 10.1.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
The @opentelemetry/resources 2.7.0 -> 2.8.0 bump (#5055) merged badly and left the lockfile with duplicate YAML mapping keys plus a dangling snapshot reference. Duplicate keys are rejected by both pnpm's parser and Dependabot's, which is why CI fails `pnpm install --frozen-lockfile` (ERR_PNPM_BROKEN_LOCKFILE, duplicated mapping key 800:3) and every typescript-sdk Dependabot PR reports "can't parse your pnpm-lock.yaml". - Collapse the three identical sdk-metrics@2.8.0 entries under `packages:` to one - Drop the two stale sdk-metrics@2.8.0 snapshots that still pinned resources 2.7.0 - Repoint otlp-transformer's resources dep at 2.8.0, the only resources snapshot that still exists No version bumps, no re-resolution. `pnpm install --frozen-lockfile` now exits 0 and a strict duplicate-key scan comes back clean.
…/mcp-server (#5048) chore(deps): bump @opentelemetry/sdk-node in /mcp-server Bumps [@opentelemetry/sdk-node](https://github.com/open-telemetry/opentelemetry-js) from 0.217.0 to 0.219.0. - [Release notes](https://github.com/open-telemetry/opentelemetry-js/releases) - [Changelog](https://github.com/open-telemetry/opentelemetry-js/blob/main/CHANGELOG.md) - [Commits](open-telemetry/opentelemetry-js@experimental/v0.217.0...experimental/v0.219.0) --- updated-dependencies: - dependency-name: "@opentelemetry/sdk-node" dependency-version: 0.219.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…itive floors (#5610) Both reach the mcp-server python evaluation env transitively (python-liquid via the langwatch SDK path dependency, python-dotenv via the tooling stack). Pin them in [tool.uv] constraint-dependencies so re-resolution cannot pull a vulnerable version back in: - python-dotenv 1.1.1 -> 1.2.2 (advisory #934) - python-liquid 2.2.0 -> 2.2.2 (advisory #1528) Co-authored-by: langwatch-agent <langwatch-agent@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…t in prod (#5856) * fix(scenarios): restore pino as direct dep so child bundle resolves it in prod The pre-compiled scenario child bundle (dist/scenario-child-process.js) keeps `pino`/`pino-pretty` EXTERNAL (build-scenario-child-process.mjs), emitting a runtime `require('pino')`. pino reaches the child only via `child-logger -> @langwatch/observability`, and #2404 (5fc0876) moved the pino family out of the app manifest into that workspace package. Because pino is then neither a direct langwatch dep nor public-hoisted, pnpm never symlinks it into `langwatch/node_modules/pino`, so the bundle's external require cannot resolve pino from `dist/` at prod runtime -> MODULE_NOT_FOUND, every scenario run reported `failure`. `pnpm prune --prod` (Dockerfile) compounds it. Re-add `pino` and `pino-pretty` (the two bundle externals in the child's import graph) to the app's direct dependencies. pnpm then top-links them into `langwatch/node_modules/`, where the external require resolves, surviving `pnpm prune --prod`. Verified: pino now resolves from the dist/ dir. Not app-source regression range: 60b9efc..HEAD changed nothing here; the delta is #2404, which predates the clean pins (rollback to 60b9efc would NOT fix it). Refs #5855 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sji6H8TmNQ1CDGLaUnQKiA * test(scenarios): guard that every child-bundle external resolves in prod layout Regression guard for #5855. The pre-compiled child bundle emits runtime require() calls for its externalized deps; each MUST resolve from the bundle's own directory (the prod resolution root). Extracts every externalized bare require("x") from the built bundle and asserts it resolves — pino included. Falsifiable: with pino de-linked from langwatch/node_modules (the pre-fix / prod state) the guard fails `expected ['pino'] to deeply equal []`; with the fix it passes. Proven RED->GREEN this session. Refs #5855 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sji6H8TmNQ1CDGLaUnQKiA * test(scenarios): execute the bundle in the guard (address CodeRabbit review) Per the coding guideline (runtime regression tests must execute the affected code path and observe the crash; string assertions are supplementary), make the child-bundle guard SPAWN the compiled bundle and assert it boots without MODULE_NOT_FOUND — the exact #5855 failure — as the PRIMARY assertion. The static externalized-require enumeration is retained as SUPPLEMENTARY coverage so a failure still names the exact unresolved module. Falsifiable via the runtime path: with pino de-linked the spawned bundle emits MODULE_NOT_FOUND and the primary assertion fails (proven RED->GREEN). Refs #5855 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sji6H8TmNQ1CDGLaUnQKiA --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…rward negateFilters/traceIds (#5858) * fix(analytics): stop fabricating 0% pass rate for days an evaluator never ran The timeseries row parser defaulted every bucket missing a series value to 0. Average-type metrics (pass rate, score) are only ever missing when the evaluator produced no processed runs in that bucket, so the 0 invented a failing day. Combined with the monitor card averaging daily buckets with equal weight, an evaluator with 6/6 passed runs on one of four active days showed 25% while the analytics donut said 100%. The parser now only zero-fills additive aggregations (counts and sums, where no rows really is zero). The monitor card headline now reads a single full-period bucket, which is run-weighted by construction and matches the donut exactly. * fix(analytics): forward negateFilters and traceIds through the read path The legacy shim dropped both fields when forwarding to buildTimeseriesQuery, so the Negate Filters toggle and trace-scoped graphs were silently ignored on every analytics query. The routed fast paths (slim and rollup tables) do not implement either parameter, so the route table now sends any query carrying them to the legacy fallback table for its source instead of serving non-negated or unscoped results. * chore(scripts): drop removed Project fields from seed-local-admin piiRedactionLevel and captured visibility moved off the Project model in the scoped data-privacy refactor; the seed script still passed them and failed on a fresh run. * fix(analytics): full-period summaries stop coalescing empty averages to 0; bind specs The timeScale full scalar path wrapped every metric in coalesce(x, 0), so an evaluator with no runs in the period read as a 0% pass rate instead of no data (the same fabrication the daily path had). The coalesce now applies only to additive series, driven by a shared isZeroWhenAbsentSeries predicate used by both the summary builder and the row parser. Also binds the new feature-file scenarios to their tests for the feature-parity check and covers the never-ran evaluator with an integration test. * fix(analytics): review round — real-value error fallback, always-fill additive keys, user-visible specs The monitor card's error fallback now averages the real daily values instead of the sparkline's filled data (which substitutes 1 for empty days on pass rates and inflated the average). Additive series that are null in every bucket now still default to 0, matching the normaliser's documented contract. Feature scenarios reworded to describe chart behavior rather than parser and routing internals. * refactor(analytics): delete the dead legacy timeseries path ClickHouseAnalyticsService.getTimeseries and its private parse and normalisation copies had no callers since the app-layer rewrite took over timeseries reads, but kept the retired zero-fill-everything semantics alive as a second implementation that could drift or be resurrected. Timeseries reads now have exactly one parser and one zero-when-absent predicate; AnalyticsBackend shrinks to the non-timeseries reads it actually serves.
…n a design concept, made the go agent well tuff (#5741) * feat(langy): rework assistant architecture and experience * fix langy CLI build inputs and restore assets * fix generated docs and Go lint * fix Go lint spelling * fix Langy CI regressions * fix Mintlify Anthropic heading * fix remaining Langy CI checks * preserve experiment SDK error classes * fix Langy rollout fallbacks and runtime image * fix chart app startup diagnostics * rename Langy runtime migration
--- updated-dependencies: - dependency-name: mcp dependency-version: 1.28.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
hardening(langy): constrain worker processes
…#5871) * fix(experiments): route workflow-agent targets to executeWorkflowCell A workflow built in Studio and saved as an agent (agent.type === "workflow") has no code of its own, just a pointer to the linked workflow. The Experiments Workbench was routing it through the same path as code agents, which built an empty DSL code node and handed it to the DSPy code-execution runner, producing a generic "user code must define one of..." error on every row of the target column. Resolve and dispatch the linked workflow the same way a direct workflow target already does (dataLoader.ts, orchestrator.ts, workflowBuilder.ts throws loudly if that resolution is ever skipped). Also fixes the target header, which showed a code icon and a dead "Edit" action for these targets: useOpenTargetEditor.ts read config.workflowId (camelCase, never populated) instead of the agent's own workflowId column, so opening the linked workflow silently did nothing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(experiments): address review feedback on workflow-agent target fix - Fix a spec/implementation mismatch CodeRabbit caught: the spec said editing a workflow-agent target opens a sidebar drawer and never a new tab, but the actual (correct) implementation opens the linked Studio workflow in a new tab — a full graph editor can't be edited meaningfully inside a narrow sidebar. Update the spec to describe what's actually implemented. - Use a named object parameter for loadPublishedWorkflow instead of positional args, per repo convention. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(experiments): add a proper editor drawer for workflow-agent targets (#5882) * feat(experiments): add a proper editor drawer for workflow-agent targets Editing a workflow-agent target previously just did window.open — no drawer, no way to map dataset columns to the workflow's input fields like every other agent target type gets. Add AgentWorkflowTargetEditorDrawer: - A "Workflow" card showing the linked workflow's name with an "Open Workflow" link to the real Studio graph editor in a new tab (a full graph editor can't be edited meaningfully inside a narrow sidebar). - Below it, the same VariablesSection mapping UI code/HTTP agent targets already get, fed by the workflow's real entry-node inputs (via getMappingSurfaceInputs, the same extraction AgentWorkflowEditorDrawer already uses for Scenarios). - Mappings persist immediately via the existing setTargetMapping / removeTargetMapping flow callback, so there's no separate save step. Stacked on fix/workflow-agent-target-execution. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(experiments): stop swallowing the Open Workflow link's data-testid Running the new drawer's integration test against real local infra (native ClickHouse/Redis/Postgres, not testcontainers) caught a real bug the earlier typecheck-only pass missed: composing ChakraLink via asChild + NextLink only forwards style-related props to the rendered child, not arbitrary data attributes, so data-testid="open-workflow-link" never actually landed in the DOM. target="_blank" is always a hard navigation into a new tab, so there was nothing to gain from routing this through the app's client-side Link in the first place — render a plain ChakraLink with href/target directly instead, which renders its own anchor and keeps every prop that's set on it. (AgentWorkflowEditorDrawer.tsx, the pre-existing Scenarios-context drawer this pattern was copied from, has the identical bug on its own "open-workflow-editor-link" testid — untested there too, so nobody's hit it. Left alone: unrelated component, out of scope here.) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * style(experiments): use the shared WorkflowCardDisplay for the linked workflow Swap the plain bordered name+link row for WorkflowCardDisplay — the same icon/name/timestamp card component EvaluatorEditorShared.tsx already uses for a workflow-backed evaluator, wrapped in the same Link isExternal pattern for consistent styling across the app instead of a bespoke box just for this drawer. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * style(agents): use the shared WorkflowCardDisplay in the Scenarios workflow-agent drawer too Same treatment as AgentWorkflowTargetEditorDrawer: swap the plain bordered name+link row for WorkflowCardDisplay (icon, name, last-updated), and render the "Open editor" link via Link isExternal directly instead of ChakraLink asChild + NextLink. This drawer had the identical latent bug on its own open-workflow-editor-link testid — untested here too (no test asserted it), so it was never caught. Fixed alongside the styling change since both come from copying the same pattern. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(experiments): address second round of CodeRabbit feedback - Preserve each workflow input's real DSL type instead of hardcoding "str" (extractWorkflowInputs), so mapping validation can enforce compatibility instead of deferring the failure to execution. - Don't show the synthetic "input" mapping field when the agent or its linked workflow fails to load / hasn't resolved yet — that fallback is only valid once we know for a fact the workflow loaded successfully and genuinely declares zero entry inputs. Render an explicit error state instead. - Capture activeDatasetId (and the dataset-source predicate) at drawer-open time instead of re-reading the store live inside onInputMappingsChange — this drawer isn't modal, so a user could switch the active dataset while it's still open, and a live read would then write the mapping into the wrong dataset's bucket. Same fix applied to the pre-existing http/code/evaluator branches too, since all four shared the identical pattern. - Restructure the drawer's integration tests into given/when blocks with action-based descriptions per project convention. - Extract the query/state derivation into useWorkflowTargetAgentData, a plain .ts hook, out of the drawer component. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…node drawer action menu (#4974) * fix(studio): remove selected node or connection with the Delete key ReactFlow defaults deleteKeyCode to Backspace only, so pressing Delete on a selected node or connection did nothing on the canvas. Bind both Backspace and Delete so removal works with either key on macOS and Windows/Linux. Spec: specs/optimization-studio/canvas-connection-deletion.feature * feat(studio): add a Duplicate/Delete action menu to the node drawer The node already exposed a duplicate/delete overflow menu on the canvas, but it was easy to miss. Mirror it into the node drawer header so duplicating or deleting a node is discoverable while its drawer is open. Structural entry and end nodes do not show the menu. Spec: specs/optimization-studio/node-duplicate-delete-menu.feature * test(studio): bind delete-key and node-menu scenarios for feature-parity Add @Scenario annotations linking the new canvas-connection-deletion and node-duplicate-delete-menu scenarios to their tests, and drop the non-standard # Bindings comments. check:feature-parity passes. * fix(studio): portal the node drawer action menu so it floats over content The drawer Duplicate/Delete menu imported Menu from @chakra-ui/react directly, which renders Menu.Content inline in the header flow with no Portal or Positioner. The menu pushed the header controls sideways instead of floating over the drawer body. Switch to the shared ~/components/ui/menu wrapper, which portals the content, positions it via floating-ui, and forces a high z-index so it sits above the drawer. * test(studio): assert drawer menu Duplicate and Delete actions fire Open the node action menu and click each item, asserting duplicateNode and deleteNode are called and that Delete also deselects and closes the drawer. Previously the test only checked the trigger was visible.
…bled (#4911) The code evaluator drawer disabled its Create/Save button whenever the form was incomplete (missing name, empty code, or no inputs) with no explanation, so it read as a dead button. Surface the reason inline in the footer next to the button, naming the missing requirements, and clearing it as they are satisfied. Adds a pure codeEvaluatorDisabledReason helper (unit-tested for the wording), an integration test for the rendered behavior, and a BDD scenario.
* Fix comparison judge error handling * test(experiments-v3): bind comparison-error-handling scenarios to their tests The feature-parity check flagged both scenarios in comparison-error-handling.feature as unbound. Annotate the existing regression tests with @Scenario so the spec binds and the required check passes. Swap the two scenario pyramid tags so each matches the type of the test it now binds to (unit serialization test / integration cell-render test). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(experiments-v3): golden/input-adaptive comparison prompts and workbench fixes Golden field now defaults to None with a matching adaptive default judge prompt (golden x input presence), include_metrics now actually instructs the judge to weigh cost/duration instead of silently ignoring them, and assorted comparison workbench fixes (charts, drawers, evaluator editor) land alongside their tests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: adapt to DomainError -> HandledError rename from main merge Main renamed DomainError/SerializedDomainError to HandledError/ SerializedHandledError (domain-error.ts -> handled-error.ts) after this branch's base commit, a semantic break git's line-merge could not catch. Update the four call sites this branch added, and re-derive evaluationResults.ts's wire parser against the new flat traceId/spanId/traceUrl shape instead of the old nested telemetry object it replaced. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: address CodeRabbit findings on comparison judge PR Fixes real bugs found in the second CodeRabbit pass: - resolveScopedRowIndices/generateComparisonCells now take named object parameters, matching repo convention and fixing 26 test call sites that silently compiled without the required scopedRowIndices property (typecheck:tests only, not caught by the default typecheck exclude). - Reopening the comparison evaluator editor after a page reload now updates the existing target in place instead of creating a duplicate column, by wiring the same target-bound callbacks openTargetEditor uses. - useTargetOutputs no longer falls back to the known-invalid schema-less output copy when a prompt lookup resolves with no data. - Comparison columns now use their own 24%/14% default/minimum width in every sizing path (drag start, drag clamp, double-click reset, total width sum), not just the render path. - specs/experiments/comparison.feature scenarios now match the actual has_golden_answer default and describe observable comparison behavior instead of prompt-rendering internals. - Ruff/type-annotation cleanup in select_best_compare.py and its tests, as const on the shipped judge-prompt list, describe/when test naming, a stronger assertion on the "Whole output" label, and corrected comments about which drawer setter notifies subscribers. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: finish golden-answer default rollout, add Zod boundary parsing - EMPTY_COMPARISON_CONFIG and comparisonEvaluatorConfigSchema still defaulted a brand-new Comparison evaluator to hasGoldenAnswer=true, contradicting the Golden-field-defaults-to-None behavior and the rewritten comparison.feature spec. Both now default to false, matching the Python evaluator's has_golden_answer default. The legacy pairwiseEvaluatorConfigSchema default is untouched (protects existing pairwise_compare rows). - readSerializedDomainError now validates its untyped wire payload with a Zod schema (+ infer) instead of manual type assertions, per the repo's own coding guideline. Same behavior, including the code/kind back-compat derivation; added unit coverage for the payload shapes it handles. - Added an explicit assertion that the include_metrics judge instruction is present when metrics are on and absent when they are off, so the metrics-injection tests actually cover the feature they are named for, not just the candidate-block formatting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test: nest given/when describe blocks in the new sizing/rehydration tests Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(comparison): reuse VariableMappingInput for Golden/Input field pickers Golden field and Input field were hand-built Chakra Menu dropdowns duplicating what VariableMappingInput (the same source.field picker prompt variable mapping and HTTP agent mapping already use) does for a single dataset source: a searchable dropdown with a closable "Source.field" chip. Swapped both to VariableMappingInput, which also picks up its keyboard nav and search-filter for free. The variant per-candidate output picker keeps its own implementation — it selects a JSON-schema subfield path within one variant's own structured output with no separate source concept, which does not map onto VariableMappingInput's source+path model. Also fixed a real gap this surfaced: reopening an EXISTING comparison column's editor never threaded the active dataset's name into comparisonContext (only the fresh "Add Comparison" flow did), so a previously-saved Golden field showed a generic "Dataset." prefix instead of the real dataset name on re-edit. Verified end-to-end against a real local dev server (native ClickHouse/ Postgres/Redis): created a Comparison evaluator, picked a golden column, saved, hard-reloaded, and re-opened the editor to confirm both the value and the dataset-qualified label persist correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(comparison): reuse VariableMappingInput for the per-variant output picker too Extends the Golden/Input field pickers' VariableMappingInput swap to the per-variant structured-output picker in VariantCard, so all three field selectors in the Comparison evaluator now share the same chip/dropdown UX used elsewhere in the app. The unwrap-path business logic is fully preserved: a single "output" field's stored path stays unwrapped (getVariantOutputOptions) even though its qualified label spans two segments ("output.answer"). The mapping's onMappingChange looks up the real path by the selected option's label rather than trusting VariableMappingInput's raw path, since the label and path deliberately diverge for that case. Removed now-dead code: qualifyFieldLabel (no remaining call sites) and the WHOLE_OUTPUT sentinel (no longer needed now that "no selection" is expressed via VariableMappingInput's own undefined mapping). Verified: typecheck and typecheck:tests clean, all 47 tests in ComparisonConfigForm.integration.test.tsx pass (9 updated for the new testids/behavior), and the broader experiments-v3 suite is unchanged (359/360 non-skipped). Confirmed in source that the picker only appears once a variant emits 2+ structured output fields, matching the Golden/Input pickers' visual pattern. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(comparison): read the unsaved prompt draft for the variant field picker The per-variant output picker in the Comparison evaluator derives its field list from useTargetOutputs, which only ever read target.outputs. Clicking "Apply" in the prompt editor writes an output-schema edit to target.localPromptConfig instead — target.outputs only refreshes on a real prompt "Save"/"Update to v2" (see promptEditorCallbacks.ts). The hook never checked the draft, so a variant's just-applied JSON schema stayed invisible to the picker until the prompt was saved as a new version, even though every other part of the UI already showed it. useTargetOutputs now prefers target.localPromptConfig.outputs (when present and non-empty) over the target's saved copy, before falling back to the live prompt query as before. Verified live: giving both support-detailed and support-concise a multi-property JSON schema and clicking Apply (no save) now surfaces the VariableMappingInput picker for both, with the same source-header + nested-field dropdown as the Golden/Input pickers. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…#5889) * chore: remove committed PR screenshots, close the gitignore gap Three directories of browser-QA screenshots (42 PNGs, ~5.8MB) reached main across PRs #5142, #5452 and #5523: langwatch/.pr-screenshots/ 20 PNGs + .gitkeep langwatch/docs/pairwise-bugfixes/ 10 PNGs docs/pr5106/ 12 PNGs None are referenced by any code, spec, doc or test — they were PR-body evidence, which by convention lives in the langwatch/pr-screenshots repo and is linked by raw URL, never committed to the product tree. Deleting them loses nothing: all three PR bodies link the images by *branch* name, and those branches were deleted on merge, so every one of those URLs already 404s today. langwatch/docs/ held nothing but the artifact dump, so it goes with it. The .gitignore rule that should have caught this was path-anchored to `langwatch/pr-screenshots/` and missed the dot-prefixed spelling actually used. Replaced with `**/pr-screenshots/` + `**/.pr-screenshots/`. * ci: fail PRs that commit screenshots outside image homes Adds the check that would have stopped the 42 PNGs this PR removes. Fails a PR that adds an image outside the directories where images legitimately live (docs/images, docs/media, langwatch/public, assets, specs, python-sdk/examples), and points the author at the pr-screenshots repo instead. A path allowlist rather than a reference scan on purpose: "is this image used anywhere?" reads like the better rule but flags plenty of legitimate docs images, which are referenced only via the docs site's own path conventions. Verified against the history it is meant to catch — it fails #5142, #5367 and #5523, and passes a real docs-image commit (#4796) and this PR's own deletions. * ci: fail the image guard loudly on git-diff errors CodeRabbit: with set -o pipefail, a git diff failure (e.g. a bad BASE_REF) was swallowed by the || true that exists only to absorb grep's exit-1 on no match, so the guard silently exited 0 — the one thing a guard must never do. Split git diff onto its own line so its failure trips set -e; grep keeps its guarded no-match. Verified: bad ref now exits 128, valid ref still passes.
* gate Langy to internal staff * refactor(langy): unify the internal-only access gate across every surface Rework of the internal-only gating so one authoritative decision guards every customer-facing Langy surface, instead of three drifting copies. - `hasLangyAccess` is the single decision (staff bypass, else `release_langy_enabled`), kept transport-free; `enforceLangyAccess` is its one tRPC adapter, applied to the `langy`, `langyGithub` and `langyEgress` routers. A denial maps to NOT_FOUND so the gate can't probe for existence. - Close the gaps the first pass missed: `langyGithub` (getInstallStatus/listRepos/disconnect) and `langyEgress` (get/set) were reachable by a non-staff customer with the flag off; they now sit behind the gate too, so Langy is actually internal-only. - Delete the deprecated `POST /api/langy/chat` Hono route and unmount it — the browser drives turns through the `langy.createConversation` / `continueConversation` tRPC mutations; nothing else called it. Its route-only tests go with it. - Share `LANGY_RELEASE_FLAG` between the server gate and the `useShowLangy` client hook so the two can never drift onto different keys. - Tests: a `hasLangyAccess` decision suite and an `enforceLangyAccess` adapter suite (single mocked boundary); drop the redundant, heavily-mocked router test that re-covered the same branches. * fix(langy): forward organizationId to the access gate on GitHub install The REST install route gated on hasLangyAccess without organizationId, so an org-scoped release_langy_enabled rule would mis-evaluate against distinctId only, wrongly denying users granted at the org level. The tRPC langyGithub router already forwards organizationId via enforceLangyAccess; align the REST route with it. The 404 gate stays ahead of the 400 org-required check so a denied account still can't probe the endpoint. * fix(langy): re-check the access gate on the GitHub setup callback The /setup callback persists the installation (recordInstallation) after only session + nonce + membership checks — it never re-evaluated hasLangyAccess. An install begun while release_langy_enabled was on could therefore complete and persist the connection after the rollout was disabled or the caller's access revoked, making the kill switch non-immediate on a customer-facing path. Re-run the gate after the membership check and before persisting; the nonce is already burned by then so a denied caller can't retry the signed state. Also lock both gate paths with tests: an org-targeted allow/deny on /install (proving organizationId reaches the flag) and a revoked-access /setup that must refuse to persist. * fix(test): export LANGY_RELEASE_FLAG from the isLangwatchStaff mock The refactor made useShowLangy import LANGY_RELEASE_FLAG from ~/utils/isLangwatchStaff, but ProjectLangyLayout's integration test mocked that module with only isLangwatchStaff, so useShowLangy threw "No LANGY_RELEASE_FLAG export is defined on the mock" and the render crashed (failing test-integration shard 5). Use importOriginal + spread to keep the real constant while still driving the staff predicate from the test's gate. * fix(langy): prove org membership before evaluating its rollout flag authorizeInResolver only defers the generic permission check; the real membership check ran inside the resolver, AFTER enforceLangyAccess. So a signed-in non-member could pass a guessed organizationId and tell that org's Langy rollout state apart from the response — FORBIDDEN when the flag is on (gate passes, membership fails) vs NOT_FOUND when it is off (gate denies) — a cross-tenant probe of an arbitrary tenant's rollout, contrary to the non-probing boundary. Lift ensureOrganizationMember into an enforceOrganizationMembership middleware and chain it before enforceLangyAccess on getInstallStatus / listRepos / disconnect. On the REST /install path, check membership before the org-scoped flag (same order). /setup already checks membership first. A non-member's response is now independent of the org's flag value. Tests: a router suite proving non-member -> FORBIDDEN with the flag never consulted (both flag states) plus member allow/deny, and matching non-member flag-independence tests on the REST /install route.
Add configurable gateway endpoint and workload hardening, expand event-sourcing and service telemetry, and keep self-hosted defaults permissive while hosted cloud opts into stricter enforcement.
…legacy outbox retired (ADR-052) (#5911) * feat(automations): specify process-manager dispatch topology * feat(automations): move dispatch onto process managers * fix(automations): consume idempotent matches once * docs(automations): finalize reaction topology * docs(automations): retire legacy outbox contracts * fix(automations): harden process delivery and webhook secrets * chore(automations): clean retired reactor helper * test(automations): cover process-manager dispatch * fix(webhooks): redact logs and prune per tenant * fix(automations): handle flag-off webhook edits * fix(automations): re-arm matches by settle window * refactor(automations): consolidate server code under app-layer/automations + pipelines/automations PR #5911 review round 1: the triggers name is retired. Services, repos and dispatch live in app-layer/automations; delivery senders (webhook/slack/email transports) in app-layer/automations/delivery; process managers, subscribers and the dispatch wiring in event-sourcing/pipelines/automations. The queue retry contract (dispatchError) moves into the event-sourcing substrate that consumes it. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * refactor(automations): split providers three ways — shared, features, app-layer PR #5911 review round 1: src/automations dissolves. Pure Zod/data provider definitions land in shared/automations/providers, the UI halves + client registry in features/automations/providers, and the server halves + registry in server/app-layer/automations/providers. ServerDef grows real persistActionParams/redactActionParams hooks that absorb the secret.ts modules, so the tRPC router reads/writes actionParams through the registry instead of switching on action, and secret handling can never sit next to client code again. Persist-time validation failures become typed HandledErrors (MissingSlackBotTokenError, InvalidActionParamsError). Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * feat(automations): slim the webhook delivery log to an outcome facts table PR #5911 review round 1: WebhookDelivery stores outcome, status, latency, a capped transport error and a failure classification — never the request URL, headers, or response body. Redacted-header storage is gone entirely (redactHeadersForLog deleted); HTTP failures store no error text because the classified message embeds a response snippet. failureKind (blocked_url / timeout / network / rate_limited / client_error / server_error) drives plain-language operator guidance in the drawer's delivery list. Spec updated: the delivery log never stores request content. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * fix(automations): no cron jobs, no match loss, no cutover table drop PR #5911 review round 1, remaining P1s + agent passes: - The chart ships no automation CronJobs at all: the new webhook delivery-log prune CronJob is gone too; pruning runs as a daily scheduled process manager (webhookDeliveryPrune) on the same ADR-052 substrate as the graph-alert sweep. The /api/cron/webhook_delivery_cleanup route is removed. - The settlement PM's pending-match cap no longer discards customer matches: past the bound the oldest matches flush to immediate dispatch intents (degraded batching, never loss); state field renamed overflowFlushed and the log intent reports flush counts. Spec + ADR updated. - The ReactorOutbox drop migration is deferred one release: dropping the table while old worker replicas still read it would crash them mid-drain during a rolling deploy. ADR-052 documents the expand/contract split. - Subscriber idempotency is pinned by contract comments + regression tests: every idempotency-key input derives from the committed event, never handler-time wall clock; tests fail if Date.now() creeps in. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * feat(automations): inline PM topology in the pipeline; encrypted failure responses in the delivery log PR #5911 review round 1, follow-ups from live discussion: - pipeline.ts now authors the full process-manager topology inline (pm => pm.state().intent().on().onWake().outbox()) exactly as ADR-052's builder API shows; only executor deps are injected. The *PM factory indirection is deleted and tests pull definitions from the real registered pipeline via a shared harness. - The string-matching failure classifier is gone. The HTTP status is the fact; the drawer derives guidance from the status bucket. For debugging, a failed attempt keeps the receiver's truncated response (body + headers + Retry-After) AES-encrypted at rest, scrubbed of our configured header values even when the receiver echoes them back, decrypted server-side by the service for the drawer, and deleted with the row by the 30-day prune. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * refactor(automations): plaintext truncated failure responses — industry baseline Per review discussion: field-level encryption is beyond what GitHub/Stripe/ Svix do for webhook delivery logs, so the failure response ({body, headers, retryAfterMs}) stores as a truncated plaintext JSONB column. The scrub of our configured header values (even when the receiver echoes them back) and the 30-day prune both stay. If volume ever makes this table a bloat concern, the designated escape hatch is moving the whole delivery log to ClickHouse with native TTL — not splitting the blob out alone. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * refactor(automations): routes go through services — no raw prisma in the router PR #5911 review round 1, service/repo pattern enforcement: - TriggerRepository gains findById/findAllByProjectId/findFirstByCustomGraphId/ create/update; TriggerService gains getById/getAllForProject/ getByCustomGraphId/create/update/softDeleteById (owning the soft-delete rule). All 19 direct ctx.prisma calls in the automations tRPC router now route through services; remaining prisma uses are service factories only. - New automations-owned AutomationCustomGraphService + repository (getById, existsInProject tenancy guard, getAllNamesByIds) and a MonitorService + repository (getAllByIds) following the same pattern. - automationDispatch.wiring loads triggers/custom graphs/projects through the injected services instead of raw prisma closures. - pipelineRegistry passes the pipeline's executor deps ({dispatch, sweep, prune}) matching the inline PM topology; httpDestination's captured response headers use the ssrfSafeFetch-derived Headers type. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * feat(automations): extract @langwatch/automations workspace package The shared automation domain leaves the app: provider definitions (Zod schemas + metadata), notification cadences, and the Liquid templating engine now live in langwatch/packages/automations as a source-only private workspace package (same convention as @langwatch/observability), consumable by the CLI, MCP server, and web surfaces with no database, React, or server-only dependencies. - The Prisma enums are mirrored in the package (enums.ts) so consumers need no @prisma/client; prismaEnumParity.unit.test.ts pins both directions at runtime and at the type level. - App imports switch from ~/shared/automations + ~/shared/templating to @langwatch/automations/* subpaths; the React halves (features/automations/ providers) and server halves (app-layer/automations/providers) stay in the app by design. - check:feature-parity scans langwatch/packages so moved scenario bindings keep counting; package ships its own vitest + tsconfig (160 tests). - pnpm minimumReleaseAgeExclude temporarily lists ws/@browserbasehq/sdk/ postcss: adding a workspace importer forces a full re-resolve and those shipped releases inside the 7-day gate; remove once they age. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * chore(automations): renumber new migrations after main's latest 20260715* sorted before main's 20260716140000_add_langy_runtime_schema — prisma migrate deploy applies strictly in order, so the webhook migrations move to 20260718* timestamps. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * test(automations): fix tests-project typecheck fallout from the review rework The typecheck (tests) CI job caught three mechanical breaks the app-project tsgo run does not see: HttpDestinationResponse mocks missing the new required responseHeaders field, an import of SavedTriggerRow from its pre-package location, and two subscriber tests destructuring a CommandHandlerResult without awaiting the possibly-async union. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * build(typecheck): explicit incremental tsbuildinfo + CI cache for tsgo tsgo (@typescript/native-preview 7.0.0) supports --incremental and --tsBuildInfoFile, and writes/reuses the buildinfo even under --noEmit. Point the tsgo app, tsgo tests, package, and editor tsconfigs at explicit tsBuildInfoFile paths under node_modules/.cache/tsbuildinfo so the cache location is stable and CI-cacheable (the tests config overrides the path it inherits from the app config to avoid a collision). Cache those buildinfo files in the langwatch-app-ci typecheck job with a rolling per-SHA key + prefix restore-keys, restored after install so pnpm cannot clobber it. Warm cache cuts the app check ~13s->1.3s and the tests check ~19s->1.6s locally. Project references were evaluated and skipped: CI runs tsgo in --project (non-build) mode where references give no skip/caching benefit, and composite would force .d.ts emit + a build step onto the deliberately source-only @langwatch/automations package. Its sources are already checked inline in the app graph and captured in the app's incremental buildinfo. Docker build cache (type=gha,mode=max, per image+arch) and pnpm-store caching (setup-node cache:pnpm) were already in place and left unchanged. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * fix(automations): address verified audit P0/P1s + Kimi review round Audit round (20 packets, all applied): - P0: deleteDispatchedBefore routed through a raw DELETE with the sanctioned tenancy opt-out marker — the prune tick no longer trips the multitenancy guard; cross-project regression tests on both store adapters - security: webhook header names validated as RFC 7230 tokens (smuggling attempts dropped, not mangled); Block Kit markdown/header blocks sanitised (mrkdwn escaping, caps); trigger.name header-injection strip in all three template context builders plus a central subject strip in the mailer; graphs router redacts via the generic per-provider redactor (fail-closed for unknown actions); daily email-cap dedup key includes triggerId; processWakeWorker logs via toSafeFailureDiagnostic - behavior: terminal webhook failures keep the open-incident claim (no re-POST loop to dead endpoints); webhook prefill no longer latches while the feature flag is loading; triggerSettlement outbox retention rides the daily prune wake - structure: RecordTriggerMatchPort moved OSS-side (EE imports inward); passesTraceOriginGuards single-sourced - docs: ADR-039 marked superseded; ADR-040 amendment matches the shipped slim delivery log incl. the best-effort scrub residual; ADR-052 package location + ADR-051 citation corrected - tests: wiring suites for both pipelines' subscribers, process-runtime guards + schedule arming, store retention on both adapters, delivery drill-down XSS rendering, webhook provider client, terminal-vs-retryable dispatch Kimi K3 (high) cross-review of the batch prompted the token validation, mailer-boundary strip, fail-closed redaction, header truncation-marker fix, and retention-log diagnostic; its remaining notes are tracked in the PR conversation. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1 * refactor(automations): store the receiver's webhook failure response verbatim — no redaction Decision (now explicit in ADR-040 §6): the response side of a failed delivery is the receiver's own output; masking what an endpoint echoes would hide exactly what the delivery log exists to show, and any exact-match scrub is evadable (re-encoded/cased echoes) so it added code and false safety without a guarantee. The boundary stays one level up — our request content (URL, headers, body) is never persisted in any form, and stored responses are pruned after 30 days. Deletes scrubHeaderValues and its call sites; error text and the truncated response body/headers are stored as received. Spec scenario and service/repository docs updated to match. Claude-Session: https://claude.ai/code/session_01Rxp9AeY1d18xEW7KnNCEt1
The code-block executor spawned the Python runner with cmd.Env unset, so Go inherited the full pod environment into the child. Any os.environ read in user code could see every variable in the pod — AWS credentials, service-account token paths, LANGWATCH_* internals, and DB/Redis/ClickHouse secrets. Set cmd.Env from a secure-by-default allowlist (interpreter/locale/TLS plumbing only). cmd.Env is always a non-nil slice, so exec can never fall back to inheriting os.Environ(). Project secrets already reach user code via Request.Secrets over stdin, so withholding the environment does not break the secrets contract. - nil allowlist -> defaultEnvAllowlist (secure default) - non-nil empty -> pass nothing (maximally locked down) - populated -> pass exactly those names, when present Regression tests assert a parent-set secret is invisible to user code while allowlisted PATH remains, and that an empty allowlist passes nothing. Scope: this covers the in-process Go codeblock path. The per-project Lambda execution path (langwatch-nlp-role) injects creds as env vars into the function's own process and needs a separate mitigation, tracked separately. Claude-Session: https://claude.ai/code/session_013EZb3uXpGq34e1GyDbQoPk
…nes without the content (#5914) * fix(aigateway): pair customer trace exports with the project that owns them The customer trace bridge paired Bundle.ProjectID with Config.ProjectOTLPToken. Those fields ride different refresh clocks — the former on the auth JWT (~15min soft expiry), the latter on the config fetch (60s TTL) — and refreshConfigBackground grafts a fresh Config onto a stale Bundle without touching ProjectID. resolveTraceProject returns the scoped project only while a VK has exactly one PROJECT scope, and otherwise falls back to the org's internal_governance project. Adding a second scope therefore flips both the id and the token, but the gateway picked up only the token. Ingest routes purely on the token (langwatch.project_id is never read there), so within the skew window a project's full gen_ai.input.messages and gen_ai.output.messages were stored under a different project, across a team boundary. The control plane already emits project_id beside project_otlp_token; the Go wire struct simply dropped it. Parse it as BundleConfig.TraceProjectID and take both halves from Config, so they are always the pair that was materialised together. The span attribute is stamped from the same field, so the routing key and the bucket key cannot diverge. An empty id with a non-empty token fails closed and logs. Separately, exporterFor returned the cached exporter before consulting the registry and nothing invalidated that cache, which made any wrong pairing permanent and blackholed a customer's traces for the process lifetime after an API key rotation. The registry is now consulted first and a cached exporter is reused only while it still serves the destination the registry names. Export failures were also discarded wholesale (_ = ExportSpans, an unconditional nil return, swallowed construction errors, no logger in the file), so a customer receiving nothing was undetectable. Failures now log and are returned, with counters for spans dropped by reason. Adds transport_test.go, which had no coverage at all — verified to fail against the previous behaviour on rotation, endpoint change, and clearing. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * feat(aigateway): trace the gateway itself, without the prompt bodies The gateway's own span carried origin, method, path and status and nothing else. No model, no provider, no token usage, no outcome — none of the metadata gateway tracing exists to provide. All of it went only to the customer's project, so operationally we were blind to our own data plane. StampInternalGenAI puts the safe half on our span: request and response model, provider, prompt/completion/total/cache tokens, cost, virtual key, gateway request id, and the upstream error classifier. It reads scalars the gateway computed and never touches params.RequestBody or params.ResponseBody, so prompts and completions stay on the customer-bound span alone. It is wired at the composition root as a decorator over the trace emitter, keeping app/ free of adapter imports. The three content keys (gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions) were declared in the internal tracer's own package, unused, in the same const block as the safe keys — one autocomplete away from being stamped on internal telemetry by exactly the change above. They are deleted; the customer bridge keeps its own private copies where content legitimately belongs. internal_span_test.go pins the boundary by stamping a span from params holding real prompt, completion and system text and asserting none of it survives by key or by value. Also fixes the sample-ratio footgun shared by aigateway, langyagent and nlpgo: each defaulted the ratio to 1.0 and then rewrote any 1.0 outside local development to 0.1, which cannot tell a deliberate 100% from an unset field. An operator asking for full sampling in production silently got 10%. Unset is now its own value and explicit ratios are honoured; the non-local default is unchanged at 0.1. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * feat(langy): keep an operational copy of worker telemetry, without the content Worker spans went only to the customer's project, re-parented onto the manager's internal turn trace id. LangWatch had no visibility into workers it runs, and the customer received spans whose parent span exists only in our backend. The relay now dual-exports. The customer path is byte-for-byte unchanged; a second, content-stripped copy goes to LangWatch's own collector when one is configured, and is simply not installed when none is. services/langyagent/otel builds that copy. It deep-copies the batch (the customer's is never mutated), keeps a strict allowlist of shape-and-cost span attributes, drops every span event, clears status descriptions, clears link attributes, and replaces the worker's resource with our service identity. The allowlist is the point. opencode's spans are the highest-density prompt/completion surface in the system — the relay already accepts-and-drops worker logs and metrics for that reason. A denylist fails open on the next key its instrumentation invents, so an unrecognised key is assumed to carry content and dropped. The second export is detached: LangWatch's observability must never delay or fail a customer's telemetry. Covered end to end — a batch carrying prompt text reaches the customer intact and arrives at our collector with the content gone and the model still present. Spec: specs/langy/langy-otel-tracing.feature Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * fix internal OTEL privacy and bounded export * refactor internal span usage attributes * harden internal trace boundary * fix(telemetry): harden internal trace metadata * fix(telemetry): fail closed on telemetry config that could invert a tenant boundary Three guards, all enforced at boot or pinned by tests: - OTel.Validate() (wired into aigateway, nlpgo, langyagent LoadConfig): rejects an off-box OTEL_DEBUG_COLLECTOR_ENDPOINT — the debug collector is a tenant-agnostic copy of every span, safe only when it cannot leave the developer's machine, so the check is on the destination address, never on what an environment calls itself. Also rejects out-of-range and NaN sample ratios (strconv.ParseFloat accepts "NaN", which compares false past a range check and lands on NeverSample). - Registry.SetFromBundle now treats a bundle with a cleared OTLP token or endpoint as a revocation and drops the cached entry, instead of exporting under the dead token until LRU eviction. - New tests pin the telemetry boundary in both directions: the gateway's internal span stamps exactly the reviewed key set and adds no events or status text; the Langy relay's customer forward carries the session key and never the internal collector credential, and the internal export carries the internal credential and never the customer's. Claude-Session: https://claude.ai/code/session_015FXg8aqvNPLk7PpwTdMtqL * feat(telemetry): one correct way — official OTel env vars across every service Every LangWatch process now reads the OFFICIAL OpenTelemetry environment variables (OTEL_EXPORTER_OTLP_ENDPOINT/HEADERS, the signal-specific traces pair, OTEL_TRACES_SAMPLER + OTEL_TRACES_SAMPLER_ARG, OTEL_TRACES_EXPORTER, OTEL_EXPORTER_OTLP_PROTOCOL, OTEL_SDK_DISABLED) through one resolver, pkg/config.OTel.Resolve, and the SDK is always handed explicit options — no exporter setting is ever read implicitly by the SDK behind the config's back, and an unset endpoint means OFF, never the SDK's localhost default. Resolution is fail-closed: - official + deprecated name (OTEL_OTLP_ENDPOINT/HEADERS, OTEL_SAMPLE_RATIO) set to different values -> boot error; precedence is never guessed - deprecated name alone keeps working, with a rename warning at boot - unsupported values of supported vars (protocol grpc, unknown sampler, ratio kinds without an arg, NaN/out-of-range args) -> boot error - recognised-but-ineffective vars (metrics/logs endpoint overrides) warn instead of disappearing silently The OTEL_* namespace is LangWatch's own telemetry ONLY. Customer trace destinations stay product configuration (per-project tokens, the customer trace bridge, LANGWATCH_ENDPOINT): nlpgo's per-tenant router deliberately does NOT read the official exporter vars — in a dev shell they point at the local observability stack, and reading them would divert customer studio traces into it. nlpgo warns loudly when they are set. Provenance marker: single-tenant providers (aigateway, langyagent) and the TS app stamp langwatch.origin=platform_internal on their resource — matching the langy relay's existing langwatch.origin=langy_worker — so internal telemetry is identifiable wherever it lands. The one guard that survives a valid-but-wrong destination, because it rides in the data, not the config. Multi-tenant resources are never marked: nlpgo's resource reaches customer projects via the tenant router. Dev ergonomics: when the debug collector IS the primary collector (the post-unification dev default — haven exports both names at the local stack), the duplicate span processor and metric reader are skipped; the debug OTLP log pipeline stays, as no primary log pipeline exists. The gateway chart now emits the official name alongside the deprecated one (equal values are accepted) so old images and new images both keep tracing through the upgrade. Claude-Session: https://claude.ai/code/session_015FXg8aqvNPLk7PpwTdMtqL * fix(telemetry): close three review findings on the env-var unification From the adversarial review round (tenancy + rollout reviewers): - nlpgo: drop the OTEL_OTLP_ENDPOINT fallback for the customer router. The unification made that name the deprecated alias for the INTERNAL collector, so one var carried two contradictory meanings — an operator pointing "nlpgo's own telemetry" at the internal stack would have silently routed customer studio content there (LANGWATCH_ENDPOINT unset). Customer routing now comes from LANGWATCH_ENDPOINT only; OTel endpoint vars present without it produce a loud customer-export-is-OFF warning. Losing telemetry is recoverable; misrouting it is not. - debug collector: resolve *.localhost names and require every answer to be loopback. RFC 6761 recommends but does not guarantee the TLD stays on-box — a corporate wildcard or /etc/hosts entry could point one elsewhere, and on nlpgo the debug collector carries full customer span content. Non-resolving names are refused rather than trusted on suffix. - langy relay: strip langwatch.origin from the customer forward's resource. The worker is prompt-injectable; it must not be able to brand its spans with LangWatch's provenance marker in the customer's project, and future ingest-side enforcement keyed on the marker must never trust a worker-supplied value. Legitimate worker resource attributes survive. Claude-Session: https://claude.ai/code/session_015FXg8aqvNPLk7PpwTdMtqL * fix(telemetry): clear the Go lint gate on the telemetry change - Wrap the .localhost resolver error with %w and split the "resolved to nothing" case out, so the failure keeps its cause. - Use require for the error assertions the linter flagged. - Rename bodyDropped to isBodyDropped for the boolean prefix rule. - Drop two spellings the misspell linter rejects. * fix(langy): strip forged provenance the worker can repeat or move The worker writes its own OTLP bytes, so it is not limited to the shapes pdata's Map API can express. Two ways it could keep a forged langwatch.origin=platform_internal on the customer forward: - Repeat the key. pcommon.Map.Remove returns after the first match and the OTLP unmarshaler preserves repeated pairs verbatim, so the twin rode through. Every removal now sweeps all entries. - Move the key onto a span. Ingest resolves span origin BEFORE falling back to the resource, so stripping the resource alone left the higher-precedence claim intact. Spans and scopes are stripped too, on every batch including one that arrives before the turn exists. The same first-match weakness let a repeated langwatch.thread.id or tag.tags survive the manager's stamp, so those are swept before being set rather than overwritten in place. Tests build the payloads at the proto level and assert through ReparentOTLP; each one fails against the previous code. * fix(telemetry): honour OTEL_TRACES_EXPORTER=none on the direct POST path The SDK exporter respected the off-switch, but PrimaryOTLP kept handing out the collector base URL. Callers that POST spans themselves — the langy relay — bypass the SDK exporter entirely, so they carried on shipping spans after an operator had turned span export off. Clear the base endpoint alongside the traces endpoint when the exporter is none. Metrics are unaffected: OTEL_TRACES_EXPORTER governs traces. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP
…fault (#5924) * feat(haven): keep the managed ClickHouse's own logs lightweight by default Stock ClickHouse never bounds its self-telemetry: the system log tables have no TTL and the server log runs at trace with a 1000M x 10 rotation. Measured on one laptop after ~5 days of ordinary dev: 610 MiB of system log tables (303 MiB of it text_log, 70M rows in asynchronous_metric_log) plus a 1.3 MiB and growing server log. haven already caps ClickHouse memory two ways, so unbounded log disk was the odd one out. By default haven now disables the high-volume system logs (text_log, trace_log, metric_log, asynchronous_metric_log, processors_profile_log, query_metric_log), caps the ones worth keeping (query_log, part_log, error_log, crash_log) at 7 days, and quiets the server log to warnings with a 50M x 2 rotation. HAVEN_CLICKHOUSE_FULL_LOGS=1 restores stock behaviour; HAVEN_CLICKHOUSE_LOG_TTL_DAYS tunes the cap. Two things this needed to actually take effect: - The config is now written in Ensure, not start(). A container already running the correct image was never reconfigured, so any tuning change reached new machines only. A changed config now forces a recreate; the data dir is a bind mount, so databases survive. - Config governs table creation, not existing tables, so an already- running server kept its unbounded ones. applySystemLogPolicy drops the disabled tables and MODIFY TTLs the kept ones. Best-effort by design: reclaiming disk must never be what stops a stack coming up. Verified live: 610 MiB -> 16.7 MiB of system logs, server log 1.3M -> 20K, all 9 stack databases intact, no config errors, doctor green. Also fixes a misreport in the brew adapters: a cancelled startup context kills the `brew list` probe too, which was surfaced as "X is not installed" and sent you off to brew install packages that were already running. Cancellation now reads as cancellation. Claude-Session: https://claude.ai/code/session_015BxdnnCo9aT1sy1nNzMEJT * fix(haven): address review — env truthiness, TTL single-source, retrofit every Ensure - parse HAVEN_CLICKHOUSE_FULL_LOGS with envTruthy ("1"/"true") instead of any-non-empty, so FULL_LOGS=0 no longer restores full logs - rename ClickHouseLimits.LightweightLogs -> LightweightLogsEnabled per the boolean naming rule - centralize the TTL fallback in EffectiveSystemLogTTLDays and derive the retrofit DDLs in domain (SystemLogRetrofitStatements) beside the rendered config, so the two can never drift - run the retrofit on every Ensure: gating it on configChanged made it one-shot — any failure between the config write and the DDL loop left the file looking current and the disk reclaim never retried - spec scenario reworded to observable outcomes and tagged @Unit so the parity checker enforces its binding; feature narrative now covers log disk - tests: writeConfig change detection (incl. legacy-file cleanup), env wiring for the log flags, retrofit statement derivation Claude-Session: https://claude.ai/code/session_01WtpQtQxDUV4frj2ChAkLEY
…s (ADR-053) (#5922) * Harden tenant-aware outbound egress * Pin Langy egress to validated addresses * feat(ssrf): one shared address-classification rule set across Go and TS Every service that makes tenant-directed outbound requests re-implemented "is this IP safe to egress to?" independently, and they had drifted: - the Langy egress proxy blocked neither NAT64 nor 6to4; - the AI gateway blocked local-use NAT64 but not the well-known prefix; - the NLP HTTP block missed CGNAT/benchmarking/reserved AND Azure WireServer (168.63.129.16) entirely; - the TS app never covered CGNAT, benchmarking or documentation ranges. A tenant who controls DNS can steer a request into whichever gap a given service left open. This makes the rule set one thing, expressed once per language and held to a single shared conformance corpus. - pkg/ssrf (Go): Classify / IsPublicAddress / Blocked over the union of the two IANA Special-Purpose Address Registries; every prefix cited to its RFC. - @langwatch/ssrf (TS): byte-for-byte equivalent, all prefixes in one file. - pkg/ssrf/testdata/address_vectors.json: the shared corpus both test suites load — if the languages ever disagree on a vector, one suite fails. Rewired onto the shared package: the Langy egress proxy, the AI gateway customer-endpoint validator, the NLP HTTP block, and the TS app's ssrfProtection. Fixes the review finding that langyagent omitted NAT64/6to4 and closes the Azure WireServer gap on the NLP HTTP path. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * fix(ssrf): keep aigateway always-blocking unspecified + link-local The shared pkg/ssrf refactor accidentally relaxed the gateway's customer-endpoint policy: 0.0.0.0/::, 169.254.0.0/16 and fe80::/10 were refused unconditionally by the old isAlwaysBlockedEndpointIP, but the rewrite folded them into the generic "special" category, so a self-hosted operator with BLOCK_LOCAL off could reach them again. Restore the original posture as a thin policy layer over ssrf.Classify: classification stays shared and corpus-tested; the stricter always-block subset (metadata + unspecified + link-local) is the gateway's own. SaaS (blockLocal on) is unchanged. Pin it with a permissive-mode test. Also re-tighten the app ssrfProtection module doc, which still listed only the classic private ranges after delegating to @langwatch/ssrf. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * fix(ssrf): make the wider egress deny set opt-in, and log what it would refuse Folding every non-globally-routable address into the nlpgo HTTP block's deny set was a silent breaking change for self-hosters. A workflow that reaches an internal service over Tailscale (100.64.0.0/10) — or any other range that was permitted here before — would start returning ssrf_blocked on a patch upgrade, with nothing in the logs or the release notes to explain why. Split "what an address IS" from "what this deployment does about it". pkg/ssrf still owns the classification; the HTTP block now decides: - cloud metadata is refused unconditionally, so the Azure WireServer hole (168.63.129.16) closes for everyone — no legitimate traffic behind it; - private/loopback/link-local/unspecified stay refused, as before; - everything else non-public is refused only under strict egress. Strict egress is off by default and arrives as a dedicated egress knob on EngineConfig, passed explicitly into SSRFOptions. Nothing under app/engine/blocks reads the environment for it — what an egress boundary refuses is deployment policy, so it is declared once and plumbed. When strict egress is off, every non-public address the strict set would have refused is logged with the range, its RFC, and the ALLOWED_PROXY_HOSTS hint. An operator can read their logs, allow-list what they actually depend on, and turn strict on deliberately — rather than discovering the list from a broken production workflow. Refusals log the same detail; the error stays "ssrf_blocked" so a tenant learns nothing about the network behind the boundary. Adds ssrf.Describe for the operator-facing range labels, and fixes the two lint failures in these files (exhaustive switch, US spelling). * fix(telemetry): one consistent service.name scheme across the services The Go mono-binary reported bare service names (aigateway / langyagent / nlpgo) while the app reports langwatch-app and the sandboxed Langy worker already reports langwatch-service-langyworker. Grafana therefore listed the same fleet under two naming schemes. Map every mono-binary subcommand to the canonical langwatch-service-<cmd> at the single chokepoint (cmd/service/main.go), so the OTel resource service.name (traces/logs/metrics) AND the per-line `service` field (pkg/clog) — both read from ServiceInfo.Service — line up as: langwatch-app · langwatch-service-aigateway · langwatch-service-langyagent · langwatch-service-nlpgo · langwatch-service-langyworker An explicit OTEL_SERVICE_NAME now wins on the Go side too, matching the app's precedence and the official OTel convention. Refresh the stale langwatch-backend / bare-name references in CLAUDE.md and haven's observability comment so the gcx examples match what is emitted. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * test(specs): mark unbound egress-isolation scenarios @unimplemented The new tenant-aware-egress-isolation.feature describes target-state behaviour with no TS step bindings yet, so the feature-parity checker (which scans @integration scenarios for a matching TS test) failed the PR. Tag every @integration scenario @unimplemented — the same convention specs/nlp-go/proxy.feature uses for not-yet-built Go surfaces — so the gate passes while the bindings remain a follow-up. @deployment/@load/ @resilience scenarios are not scanned and are left as-is. Addresses the CodeRabbit critical review comment on #5922. Claude-Session: https://claude.ai/code/session_01CLc1dDjucq1kkbpFUHGjfP * docs(adr-053): keep live deployment posture out of the public ADR The ADR carried a dated read-only inspection of the production cluster: absent gateway env flags, unselected NetworkPolicies, credential and service-account-token specifics, and the Langy sandbox egress state. This repo is public and Phase 0 containment has not shipped, so that section read as a current-state map of an unmitigated path. Replace it with the structural rationale the tracks actually design against, and drop the same class of detail from the current-shape diagram and the phase headings. Deployment-specific review and remediation tracking stay in the private infrastructure repo. * fix(telemetry): stamp service fields on the dispatch logger clog.New reads ServiceInfo off the context, but the logger was built from a bare context before serviceTelemetryName(cmd) was known, so its records carried no service/version/environment. Set ServiceInfo first, then rebuild the logger. The bootstrap logger keeps the pre-dispatch failures, where there is no service name to stamp yet. * refactor(ssrf): boolean prefixes, named params, behavioural spec wording Review follow-ups: prefix the new booleans (shouldBlockLocal / isAllowlisted), move parsePrefix/prefixContains/blocked onto named object params per the house rule, inline the single-use legacy predicate, and restate three integration scenarios in behavioural terms rather than naming the infrastructure mechanism that enforces them. bytesEqual keeps positional args: it is a symmetric equality primitive where {a, b} names add nothing. * ci(ssrf): run both halves of the conformance corpus in one job pkg/ssrf and @langwatch/ssrf are held to one corpus, but their suites ran in two unrelated workflows. A cross-language divergence therefore surfaced as two red Xs in different places, or as none at all when only one side's paths changed and the other was never re-run. Add a path-gated workflow that runs the Go and TypeScript halves against address_vectors.json in a single job, so "the two languages disagree" is one unambiguous failure. The TypeScript step runs even when the Go step already failed, so a divergence report names both sides in one log. * refactor(ssrf): named params for bytesEqual, document remaining helpers Closes out the review thread on positional args: bytesEqual now takes a named object like its neighbours, so the rule holds across the whole module rather than almost all of it. Also documents the symbols the docstring check flagged - the two address parsers, the Prefix shape and parsePrefix, the Langy egress log monitor, and the NLP resolver seam. Each says why it is shaped that way (strict IPv4 parsing and mapped-IPv6 collapse are bypass defences; parsePrefix throws because a malformed constant must fail at import, not silently drop a rule) rather than restating the signature. * fix(nlpgo): take the egress-policy logger from the context The startup line reached for deps.Logger directly while the rest of the service pulls its logger off the context via clog.Get. Same underlying zap logger, but the context form picks up whatever clog.With has since attached, so the line stays consistent with every other log site. * build(ssrf): incremental typecheck cache, and make it a rule The new package's tsconfig had neither incremental nor tsBuildInfoFile, so every typecheck re-checked it cold. Match the pattern the app and packages/automations already use, writing to a package-scoped path so workspace packages cannot clobber each other's cache. Record it in CLAUDE.md as well - it is the kind of thing that is only ever noticed once a package is already slow, so it belongs in the rules table rather than in reviewers' heads. * revert(ssrf): keep the blockLocal / allowlisted parameter names 164ee39 renamed these to shouldBlockLocal / isAllowlisted on a bot suggestion that the PR author had answered "no" to. Put the original names back. The named-object-parameter change on blocked() stays - that was a separate thread and is not what was rejected.
…+ shared package (#5917) * feat(errors): handled-error remediation channel (tips/docsUrl/fault) + shared package - Extract HandledError core into shared source-only package @langwatch/handled-error (packages/handled-error); the app imports it via the app-layer shim, mcp-server bundles it via tsup noExternal. - Add additive tips/docsUrl/fault to HandledError, SerializedHandledError, and the Go herr ErrorBody (reserved Meta keys promoted on the wire, lossless Body/FromBody round-trips). - Fault-axis logging: handled customer errors log at warn (spike-watch), platform/provider at error; PostHog capture reserved for unhandled 5xx. - Boundaries: tRPC (via serialize), REST (additive keys only), packages/api formatter, evaluationResults zod schema, MCP server parses error bodies and renders Tips/Docs in tool errors, SSE error frames carry domainError and mask unhandled messages. - ClickHouse: resilient client translates MEMORY_LIMIT_EXCEEDED / TIMEOUT_EXCEEDED / connection failures into typed errors with remediation tips (raw error preserved in reasons so retry classifiers keep working). - Annotate traces/api-key/evaluations/langy error classes with tips + docs links; aigateway budget_exceeded Go pilot. - ADR-045 amendment + spec scenarios. * refactor(errors): address PR review - Remove the app-layer re-export shim; all consumers import @langwatch/handled-error directly; trace-URL provider wired from server/handled-error-wiring.ts (server.mts + workers.ts) - Central remediation registry (error-remediation.ts): all tips + docs links keyed by code, with a CI test asserting every docsPath exists - Replace NlpgoHandledError with plain HandledError via handledErrorFromHerr (now accepts an httpStatus override) - Preserve wire trace/span ids on herr-adapted errors; emit fault + trace ids on serialized reasons - Fault-aware log levels on every boundary: Hono REST (logHttpRequest), SSE, and Go telemetry join tRPC's fault axis - herr: fault validated against the three-value contract, lifting code simplified (metaString/metaStrings helpers) - classifiers: single-pass reasons unwrap, cleaner memory-limit matcher - resultMapper: typed meta.reason for evaluator 401/403 - package: vitest as a real devDep (no ../ hack), incremental tsc - lockfile regenerated; minimumReleaseAgeExclude entries version-pinned * ci: retrigger workflows (pull_request synchronize did not fire on previous push) * fix(deps): dedupe minimumReleaseAgeExclude after main merge, keep version pins * fix(errors): CI + review-pr findings - mcp-server standalone install: add packages/handled-error to its workspace, regenerate its lockfile, COPY the package in its Dockerfile - Fix two stale ../../handled-error imports missed by the shim removal - gocritic rangeValCopy: index the FromBody reasons loop - REST + packages/api envelope tests for tips/docsUrl/fault - Duck-type handled-error detection at the tRPC logger and SSE boundary (second-copy-of-module defense, same as the Hono handler) - nlpgo envelope parses trace_id/span_id; MCP prefers code over error - Biome fixes on PR lines; spec/ADR text synced (shim references, @unimplemented tags lifted where now implemented, Go fault-default asymmetry documented) - Resilient-client translation wiring test; identity assertion in the double-translate test; registry entries must carry a channel * refactor(errors): drop unused handled-error re-exports from app-layer barrel * fix(errors): static wiring import in server.mts (repo rule), guard reasons length
* chore(cleanup): delete orphaned dead code (0 importers, ~3k lines)
Removes files verified to have zero importers (static, dynamic import(),
and path-based), across superseded and abandoned clusters:
- superseded chat/playground hooks (usePlaygroundStore, useLoadChatMessages,
useChatWithSubscription, useMemoizedChatWindowIds)
- welcome cards removed from the layout (IntegrationChecksCard,
AgentSimulationTesting, ResourcesCard)
- orphaned components/hooks/utils (PromptsList, ModelProviderSelector,
SimulationZoomGrid, ScenarioInfoCard, usePersonalIngestionBinding,
NewEvaluationsButton, EnterpriseLockedKpi, ScenarioLibraryToolbar,
ProjectIntegration, useAnimatedFocusElementById, easyCatch,
toFixedWithoutRounding, typescript/assertNever, headers/normalizeHeaderValue)
- experiments-v3 orphans (useEvaluatorMappings, ConfigPanel)
- unadopted Chakra v3 wrappers (ui/steps, ui/listbox)
- superseded in-process scenario runner (simulation-runner.service)
- unreferenced prompt-config DTO + optimization_studio types/modules
* refactor(mailer): consolidate @react-email/* subpackages into @react-email/components
All 6 mailer templates imported Button/Container/Heading/Html/Img/Section/Text
from 7 individual @react-email/* subpackages; @react-email/components (already a
dependency) re-exports all of them. Switch the imports and drop the 7 redundant
subpackage deps. @react-email/render kept (separate render-to-HTML concern).
* refactor: replace lodash.{clonedeep,debounce,isequal} with lodash-es
lodash-es (already a dependency) provides cloneDeep/debounce/isEqual with
identical call sites. Converts the 7 default imports across 6 files to named
lodash-es imports and drops the 3 single-purpose lodash.* packages plus their
@types. (isEqual also overlaps the existing fast-deep-equal.)
* chore(deps): remove unused dependencies (prod + dev)
Verified unused via byte-level import scan + a dynamic-reference sweep
(no require()/import()/string-target references anywhere):
prod: @ai-sdk/{amazon-bedrock,azure,google} + @aws-sdk/client-bedrock-runtime
(LLMs route through @ai-sdk/openai-compatible); both unused OTLP exporters
+ @opentelemetry/sdk-node (only -proto is wired); all three @tiptap/*;
react-international-phone/-helmet-async/-collapsed; react-select (redundant
transitive of chakra-react-select); handlebars, micro, path-to-regexp;
cookie/debug/flat (only string literals matched, never the packages);
lodash.omit, lodash.throttle; fetch-h2 (+ its stale test mock)
dev: webpack + raw-loader + string-replace-loader + turbo (Vite is the bundler,
no webpack/turbo build); jest-mock-extended (tests use vi.mock); watch;
testcontainers (only @testcontainers/{clickhouse,redis} are used)
types: @types/{cookie,debug,lodash.omit,lodash.throttle,mermaid}
Kept despite tooling flags: pino-pretty + pino-opentelemetry-transport (runtime
string transport targets), import/require-in-the-middle (OTel loaders),
concurrently (used by start.sh dev path), @langwatch/scenario (62 imports).
@hono/node-server, dropped in the original pass, is kept: the Hono bridge in
src/start.ts imports getRequestListener from it.
deps 193->170, devDeps 53->41. Trims the pruned production image (bedrock/
aws-sdk/otel-exporters/tiptap and their transitive closures).
* chore(deps): move @vitejs/plugin-react to devDependencies
Only imported by vite.config.ts (build-time). The Dockerfile runs
'pnpm prune --prod' after 'vite build', so a build-only plugin sitting in
dependencies ships in the runtime image needlessly.
* docs(featureFlag): swap retired flag key in examples for a live one
release_ui_simulations_menu_enabled was fully retired (read by zero components)
but survived in JSDoc/comment examples. Point the examples at a live flag
(release_ui_ai_gateway_menu_enabled) so the docs don't reference a dead key.
* refactor: inline escape-string-regexp, remove the dependency
Single call site (llmModelCost model-name matching). Inlines the exact v5
implementation — crucially preserving the hyphen -> \x2d encoding the
surrounding regex logic depends on (it later maps \x2d back to '-').
* refactor(scenarios): drop deprecated serialized.adapters re-export shim
Repoints the 2 consumers (registry + test) at ./serialized-adapters directly
and deletes the @deprecated re-export shim, closing a 'never re-export'
violation.
* docs: correct misleading comment on inlined escapeStringRegexp
The \x2d encoding is normalized back to literal hyphens immediately
below, so the matching does not rely on that form.
* fix(playground): unwrap immer draft before cloneDeep in splitTab
lodash-es@4.18's cloneDeep walks Ctor.prototype via isPrototype, which
violates the immer draft Proxy invariant ('get' on proxy TypeError).
The old lodash.clonedeep@4.5.0 predated that code path. Unwrap the
draft with current() before cloning.
* docs(playground): drop unreproducible crash claim from splitTab clone
The comment above the splitTab clone asserted that lodash-es@4.18's
cloneDeep walks Ctor.prototype, violates the immer draft Proxy invariant
and throws. That does not reproduce: against the installed immer@11.1.8
and lodash-es@4.18.1, cloneDeep on a raw draft returns a correct plain
clone, and the whole store suite passes with current() removed.
Keep current() — taking a plain snapshot before leaving a recipe is the
correct way to do this, and it stops a draft proxy reaching committed
state — but describe what it actually does instead of a crash nobody can
trigger.
Adds a characterization test for the real invariant: the split tab's data
must be plain and share no object identity with the source. It is labelled
as characterization, not regression, because it passes with and without
current() — there is no failing case to pin, and a test that cannot fail
should not be dressed up as one.
* chore(cleanup): repoint specs at surviving modules, drop stale turbo.json
Deleting dead code left documentation pointing at files that no longer
exist. Fixes the citations this PR breaks:
- specs/scenarios/simulation-runner.feature documented SimulationRunnerService,
which is deleted here. The behaviour is still required and now lives in
src/server/scenarios/execution/ (data-prefetcher loads the scenario and
carries situation/criteria, serialized-adapters resolves the targets,
execution-pool + scenario-child-process run it), so the feature file stays
and records where it went. Rewording its steps off the old class name is
left as follow-up rather than guessed at here.
- specs/scenarios/AUDIT_MANIFEST.md: three KEEP rows cited line numbers inside
the deleted service; repointed at the modules that implement them today.
- specs/home/AUDIT_MANIFEST.md: the DELETE verdict rested on a grep hit in
welcome/ResourcesCard.tsx, which is deleted here — the grep now finds
nothing, which only strengthens the verdict.
- specs/experiments-v3/AUDIT_MANIFEST.md: four KEEP rows cite
useEvaluatorMappings.ts, deleted here with no replacement, so those verdicts
rest on an implementation that is gone and need re-deriving. Flagged in place
rather than silently rewritten.
Also drops langwatch/turbo.json: the turbo devDependency is removed in this
PR, nothing invokes turbo, and the config still described a Next build
(next.config.mjs inputs, .next/** outputs) for what is now a Vite app.
* chore(ci): drop the dead Turbo cache step
Nothing invokes turbo, so langwatch/.turbo is never produced and this
cache step has always been a no-op lookup. The turbo devDependency and
turbo.json are removed in this PR; this is the last reference.
…ir config (#5866) * feat(workflows): remove workflow-level default_llm, LLM nodes own their config Creating a workflow on a fresh environment (no ModelDefaultConfig rows seeded at any scope) persisted default_llm.model = "" because the creation form used the cascade-resolved default verbatim and the cascade legitimately had nothing to return. Every component run then 500'd with the opaque 'Model provider not configured: ' error (empty provider name, the split of an empty model string). default_llm had no reachable UI left (WorkflowPropertiesPanel was dead code) and existed only as an execution-time fallback for LLM nodes without their own config, so remove the concept outright instead of patching fallbacks around it. DSL spec_version bumps to 1.5: - saveOrCommitWorkflowVersion materializes a model on every modelless llm parameter: the payload's legacy default_llm first, then the cascade-resolved workflows.create_default feature (new registry key, role DEFAULT), then the registry flagship DEFAULT_MODEL. Seeding defaults is never a precondition for a runnable workflow. - migrateDSLVersion folds legacy default_llm into modelless llm params on read and drops the field; runWorkflow now migrates published versions before dispatch so pre-1.5 workflows keep running. - addEnvs throws the typed LlmModelNotSetError (mapped to a 422 by the post_event route, no error capture) instead of the misleading provider error when stale client state ships a modelless node. - nlpgo drops Workflow.DefaultLLM and its resolveLLMConfig fallback; a modelless signature node fails with the typed llm_model_not_set NodeError naming the node. - The studio seeds freshly dragged signature nodes from the resolved default, templates ship modelless llm params on purpose, and the dead WorkflowPropertiesPanel + workflowSelected store state are gone. See specs/workflows/workflow-node-owned-llm.feature. * fix(workflows): address review round 1 - Update dslAdapter and workflowBuilder contract tests to LATEST_SPEC_VERSION (the 1.5 bump broke their 1.4 assertions) - Drop the now-unused execReq param from runSignature (golangci unparam) - Bind all ten workflow-node-owned-llm scenarios to tests, including a new runWorkflow integration test proving a pre-1.5 published version dispatches with the folded model, and a new LlmSignatureNodeDraggable unit test for drag-time seeding; reword @integration scenarios to user-visible behavior - Optional-chain model access in useModelProviderKeys and rename it to .ts (hook files must not be .tsx) - materializeNodeLlmConfigs only treats ModelNotConfiguredError as 'nothing configured'; infrastructure failures propagate instead of silently pinning the flagship, with tests for both paths - Gate the scenario data-prefetcher default-model fallback to legacy (pre-1.5) DSLs so a modelless 1.5 node fails with the typed engine error instead of running silently - Customer-facing copy for the workflows.create_default registry entry, as-const LATEST_SPEC_VERSION, when-style describe naming, and the current Haiku 4.5 identifier in tests * fix(scenarios): compare DSL spec versions by components, not parseFloat parseFloat("1.10") is 1.1, which would misclassify a future spec version as legacy and silently inject DEFAULT_MODEL. Compare major/minor numerically instead, keeping the conservative invalid-version -> legacy fallback.
…ion (#5932) * fix(errors): make HandledError boundary guards survive class duplication `instanceof` compares class identity, so it breaks when a bundler includes the handled-error module twice — Next.js/turbopack does this across route and server boundaries. An error raised from the second copy is a genuine HandledError, but `instanceof HandledError` returns false, which silently downgrades a handled 4xx into an unhandled 500 and replaces the vetted copy with "An unknown error occurred". #5917 already hit this and hand-rolled duck-typed workarounds in sse.ts and trpc.ts ("a bundler can load a second copy of the package"). This centralises that instead of spreading it further. Every instance already carries a `readonly isHandled = true` brand, so match on that. Hardening the guard fixes every call site at once: - `isHandled`: `instanceof` OR the brand - `isUnhandled` / `toUserMessage` / `serializeReason` route through `isHandled` - 12 call sites across trpc, sse, langy, automations, experiments, clickhouse-trace, event-sourcing and error-handler now ask `isHandled` - the two ad-hoc duck-typed blocks collapse into it One `instanceof HandledError` now remains repo-wide: the canonical one inside `isHandled` itself. `is()` stays instanceof-only on purpose: subclass narrowing cannot be brand-based (the brand says "handled", not "which subclass"), so cross-boundary callers compare `code`. The brand also deliberately does not match a deserialised wire payload — a plain object with no prototype, so the methods this guard promises would not exist; those go through `handledErrorFromHerr` / `isHandledErrorLike`. error-handler keeps its `code`+`httpStatus` tail so the accepted set does not narrow; it only reads fields there. Verified the 4 new duplication tests fail against the old bare-instanceof guard and pass after. Claude-Session: https://claude.ai/code/session_01GofbFj1QUWFKdeZmTBDEhp * fix(errors): require a real Error for the handled-error brand Addresses review on #5932. `hasHandledErrorBrand` did not do what its own doc comment claimed. It checked `typeof error === "object"`, so a plain `{ isHandled: true }` matched — and that is reachable, not hypothetical: the brand is an own *enumerable* class field, so `JSON.parse(JSON.stringify(err))` or a worker `postMessage` structured clone yields exactly such an object. It would then have been treated as a handled error and crashed on `.serialize()` in trpc.ts and sse.ts. Requiring `error instanceof Error` rejects those while still admitting bundler duplicates, since `Error` is the realm's shared global. Also guards the `in` operator in error-handler against a primitive or null throw, which would have thrown a TypeError inside the error handler itself. New test pins it: "rejects a cloned payload carrying the brand" fails against the old typeof check and passes now. Claude-Session: https://claude.ai/code/session_01GofbFj1QUWFKdeZmTBDEhp * fix(sse): drop raw request input from the handler error log The catch-all logged the full tRPC input payload, which can carry PII. The observable error path in the same file already omits it — align the two. Error identity, path, and the handled-error code/fault fields remain. Claude-Session: https://claude.ai/code/session_01GofbFj1QUWFKdeZmTBDEhp
…atch (#5936) An OriginResolvedEvent was committed with an empty aggregateId, which then flowed into the automations pipeline and failed validation on the lw.automation.trigger.record_match command (traceId requires min length 1), poisoning the triggerMatch reactor job. Guard at three points: - resolveOriginCommandDataSchema now requires a non-empty traceId, so the bad value can never reach the event store in the first place. - originGate reactor skips scheduling deferred resolution when the trace aggregate has an empty id, and warns instead of silently propagating. - traceAlertTriggerMatch subscriber skips events with an empty aggregateId rather than throwing, so already-committed bad events drain instead of poisoning the job. Claude-Session: https://claude.ai/code/session_01W3C3yCw3dLrERt5tEnc3iD
… runner, accent polish (#5929)
… device login works out of the box (#5901) * fix(governance): enable the governance flag by default so self-hosted device login works out of the box * chore(governance): update stale ships-dark flag comments to default-on * fix(specs): describe the operator override behaviorally per review * docs(governance): drop disable-flag instructions from feature docs * docs(governance): write availability notes from the fresh-reader perspective * docs(governance): ADR status stays terse, SaaS notes name the device-login gate
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )