Skip to content

ansible: raise jenkins agent heap on mnx x64 workspace machines - #4493

Merged
ryanaslett merged 1 commit into
nodejs:mainfrom
ryanaslett:jenkins-workspace-mnx-heap-size
Sep 28, 2026
Merged

ryanaslett merged 1 commit into
nodejs:mainfrom
ryanaslett:jenkins-workspace-mnx-heap-size

Conversation

@ryanaslett

Copy link
Copy Markdown
Contributor

Problem

test-mnx-ubuntu2204-x64-1 (jenkins-workspace-9) hit a JVM heap OutOfMemoryError in its Jenkins agent process on 2026-09-28:

java[209827]: java.lang.OutOfMemoryError: Java heap space
java[209827]: SEVERE: [JNLP4-connect connection to ci.nodejs.org/...] Reader thread killed by OutOfMemoryError

This killed the JNLP connection's reader thread, which the Jenkins controller saw as the agent going offline mid-build. Five node-test-linter builds (and their upstream node-test-pull-request runs) hung for 65–77 minutes before timing out and failing with ClosedChannelException / "Agent went offline during the build".

The agent runs with -Xmx{{ server_ram|default('128m') }} (ansible/roles/jenkins-worker/templates/systemd.service.j2), and no host currently overrides server_ram, so all jenkins-workspace machines run with a 128m heap.

Change

Set server_ram: 512m for the two mnx x64 workspace machines (jenkins-workspace-9 and jenkins-workspace-10). Both have 16G RAM and 72 cores, so there's ample headroom.

Why not jenkins-workspace-6 too

jenkins-workspace-6 (test-ibm-ubuntu2204-x64-3) is a much smaller instance — 3.8G RAM, 2 vCPUs. I checked its journal going back to 2026-07-17 and found no JVM heap OOM there; the only OOM event in that window was an unrelated host-level OOM-killer event that killed a git build process (a symptom of the whole box being short on RAM, not the agent's heap ceiling). Raising the agent's heap on that box wouldn't address a problem it's had, and would take RAM away from build processes on an already memory-constrained host, so it's left at the 128m default for now.

🤖 Generated with Claude Code

test-mnx-ubuntu2204-x64-1 (jenkins-workspace-9) hit a JVM heap
OutOfMemoryError in its Jenkins agent process on 2026-09-28, which
killed the JNLP connection thread and caused several node-test-linter
builds to time out and disconnect ('Agent went offline during the
build'). The agent runs with the default -Xmx128m from
roles/jenkins-worker/templates/systemd.service.j2.

Both mnx x64 workspace machines (jenkins-workspace-9 and
jenkins-workspace-10) have 16G of RAM and plenty of headroom, so raise
server_ram to 512m for them.

jenkins-workspace-6 (test-ibm-ubuntu2204-x64-3) is a much smaller
instance (3.8G RAM, 2 vCPUs) and its logs show no JVM heap OOM in over
two months of history, only an unrelated host-level OOM kill of a git
process. Left at the 128m default since there's no evidence it needs
more and the box can't spare the memory.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@ryanaslett
ryanaslett merged commit 50a435b into nodejs:main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants