You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[Ubuntu 24.04][Sandbox] After reboot: sandbox container exit-0 crash-loops (reliable) and connect recovery can recreate the container and lose /sandbox/.openclaw/workspace/ (observed once) #7404
Two related post-reboot failures on the same sandbox. Common root cause: after a host reboot the container is restored by Docker but nothing re-runs nemoclaw-start, so the in-sandbox OpenClaw gateway does not come back on its own.
Bug A (reliable): on some reboots the container exit-0 crash-loops — entrypoint starts, exits 0 within milliseconds, unless-stopped restarts it, forever. Observed RestartCount of 278 and 299 on two separate reboots. Because it exits inside the 45s healthcheck StartPeriod, it never reaches unhealthy and never serves.
Bug B (observed once): on a different reboot the container stayed up but the gateway was dead; running nemoclaw <name> connect printed "Recovering... Recreating the sandbox container..." and replaced the container (new ID confirmed). All /sandbox/.openclaw/workspace/ files were reset to image defaults — IDENTITY.md back to template, memory/ gone, all timestamps at container-creation time.
Bug B is not reliably reproducible. Across three deliberate reboots I could not re-trigger the recreate; connect instead took a non-destructive "re-establishing dashboard port forward" branch and left the workspace intact (verified with a marker file). A reboot without running connect also preserved the marker — so reboot alone does not lose state; only the connect recreate path did. The original loss followed a DGX system upgrade, which may be relevant.
This report covers two related failures that both surface after a host reboot on the same sandbox. They share a root cause — after reboot the container is restored but nemoclaw-start is not re-run, so the in-sandbox OpenClaw gateway does not restart automatically — but they manifest differently and interfere with each other, so I'm filing them together.
Bug A — container exit-0 crash-loop after reboot (reliably reproducible)
On some reboots the sandbox container cannot stay up. The entrypoint starts, exits cleanly (exit 0) within milliseconds, and Docker's unless-stopped policy restarts it immediately. This repeats indefinitely — I observed RestartCount values of 278 and 299 on two separate reboots. docker inspect shows FinishedAt in the same millisecond as StartedAt with ExitCode=0.
Because each restart re-enters the 45-second healthcheck StartPeriod, the container never transitions to unhealthy and the agent never serves. docker ps shows a perpetual Up <seconds> (health: starting) with an uptime that keeps resetting.
Bug B — connect recovery can recreate the container and lose workspace state (observed once)
On a different reboot the container stayed up but the gateway was dead. Running nemoclaw <name> connect printed "Recovering... Recreating the sandbox container with its managed startup command..." and replaced the container (new container ID confirmed via docker inspect). All files under /sandbox/.openclaw/workspace/ were reset to image defaults — IDENTITY.md, SOUL.md, USER.md, AGENTS.md, and the memory/ directory of daily notes. Sandbox registration, model, provider, policy presets, and channel config were preserved; only workspace state was lost, with no warning or confirmation.
Reproducibility caveats for Bug B:
The gateway-down precondition reproduces on every reboot.
The destructive recreate itself does not. In three deliberate reboot tests, connect took a non-destructive "Dashboard port forward is missing or dead. Re-establishing..." branch and left the container ID unchanged and the workspace intact (verified with a marker file).
A reboot with the marker present but no connect also preserved it — so reboot alone does not lose state.
The recreate branch appears to require the container to be unhealthy when connect runs (the original incident had ~1h of failed healthchecks). On reboots where Bug A crash-loops, the container never reaches unhealthy, so I could not set up Bug B's precondition on demand.
The original loss followed a DGX system upgrade, which may have influenced container state.
Expected
Bug A: after reboot, the container should start its agent and stay up, or fail loudly — not exit 0 and churn silently forever.
Bug B: no connect branch should silently discard workspace state. Either run the same backup/restore phase rebuild uses, or warn and require confirmation before a destructive recreate, or defer to rebuild.
Bug B is the same defect class as #4179 (closed by #4197), which fixed the installer-driven recreate path.
Reproduction Steps
Bug A (reliable)
Onboard a sandbox (local vLLM inference, GPU enabled).
Reboot the host.
After boot: docker inspect --format 'restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} exit={{.State.ExitCode}}' <container>
Observed: high and climbing restarts, finished ≈ started (same ms), exit=0.
docker ps --filter name=openshell-<name> --format '{{.Status}}\t{{.RunningFor}}'
Observed: Up <seconds> (health: starting) with RunningFor of hours — i.e. uptime keeps resetting.
It never reaches unhealthy.
Not every reboot triggers Bug A — 2 of my reboots crash-looped, others left the container up with a dead gateway. This is a consistent issue.
Bug B (precondition reliable, recreate NOT reliably reproducible)
On a reboot where the container stays up (Bug A did NOT trigger), confirm the gateway is down:
Write a marker so loss is detectable: docker exec -u sandbox <container> bash -c 'echo marker > /sandbox/.openclaw/workspace/MARKER.md'
Wait for the container to become unhealthy (docker ps → (unhealthy)). <--- This is the part I can't reproduce, which is why I didn't put it as a separate report; unclear to me what leads to this state
The recreate branch appears to need this; on a healthy/starting container connect takes the non-destructive branch instead.
Record the container ID: docker inspect --format '{{.Id}}' <container>
Run nemoclaw <name> connect and read the banner:
"Recovering... Recreating the sandbox container..." = destructive (observed once).
"Dashboard port forward is missing or dead. Re-establishing..." = non-destructive (observed 3×).
If recreated: cat /sandbox/.openclaw/workspace/MARKER.md (gone in the original incident) and re-check the container ID (new in the original incident).
I could not force step 5 down the recreate branch on demand — partly because Bug A prevents the container from reaching the unhealthy state step 3 needs.
The original Bug B incident occurred on a reboot triggered by a DGX system upgrade. NemoClaw/OpenClaw versions and the sandbox image tag were unchanged across that reboot (verified in ~/.nemoclaw/sandboxes.json before/after), but I cannot rule out that the upgrade affected container/sandbox state.
Same Docker build (29.2.1, a5c7197) and aarch64 / Ubuntu 24.04 NVIDIA-hardware profile as #4179.
Debug Output
[debug] Collecting diagnostics for sandbox ''...
[debug] Quick mode: true
═══ System ═══
Wed Jul 22 03:43:46 PM MDT 2026
Linux 6.17.0-1026-nvidia #26-Ubuntu SMP PREEMPT_DYNAMIC Thu Jun 25 00:57:17 UTC 2026 aarch64 aarch64 aarch64 GNU/Linux
15:43:46 up 1:44, 6 users, load average: 0.41, 0.57, 0.53
total used free shared buff/cache available
Mem: 124610 80341 18291 109 27190 44268
Swap: 16383 0 16383
NVIDIA-SMI 580.159.03 Driver Version: 580.159.03 CUDA Version: 13.0
GPU 0: NVIDIA GB10, 37C, P0, 9W
Processes: 20108 C VLLM::EngineCore 68669MiB
═══ Docker ═══
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
8fe01f908de0 6217934eb80f "/opt/openshell/bin/…" 37 minutes ago Up 37 minutes (healthy) openshell--21d5c401-cc45-4e8c-bf8e-cb08d890cf32
c2028dde71b9 9204569b17ee "/bin/bash -lc 'pip …" 26 hours ago Up 2 hours 0.0.0.0:8000->8000/tcp, [::]:8000->8000/tcp nemoclaw-vllm
NOTE: container 8fe01f908de0 was created by the connect auto-recovery at
21:06 UTC. The pre-recovery container was 62c97a5d7b2f. The sandbox itself
was created 2026-07-21 22:06:56.
═══ OpenShell ═══
Server Status
Gateway: nemoclaw
Server: https://127.0.0.1:8080
Status: Connected
Version: 0.0.85
NAME CREATED PHASE
2026-07-21 22:06:56 Ready
Log excerpt (post-recovery steady state; the incident window predates this
capture). Representative lines only — the full output is ~200 lines of
these repeating:
NOTE: all in-sandbox processes start at 21:06 — the recovery time, not the
original sandbox creation time. PID 1 (the supervisor) included, which
confirms a fresh container rather than a restart.
═══ Kernel Messages ═══
(skipped: dmesg_restrict=1)
Logs
## Bug A — crash-loop
$ docker inspect --format 'restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} exit={{.State.ExitCode}}' openshell-<name>-...
restarts=299 started=2026-07-22T23:15:16.829Z finished=2026-07-22T23:15:16.687Z exit=0
$ docker ps -a --filter name=openshell-<name> --format '{{.Status}}\t{{.RunningFor}}'
Up 10 seconds (health: starting) 2 hours ago
`finished` precedes `started` by ~140ms across the restart boundary;exit 0;
299 restarts. A separate earlier reboot showed the same signature at
restarts=278.
Healthcheck definition — StartPeriod (45s) exceeds the container's lifetime,so it can never leave `starting`:Interval=30s Timeout=5s StartPeriod=45s Retries=3## Bug B — recreate + workspace loss (original incident)$ docker inspect --format '{{.Id}} created={{.Created}}' openshell-<name>-...8fe01f908de0... created=2026-07-22T21:06:04Z# prior container ID (docker events, pre-recovery): 62c97a5d7b2f...$ ls -la /sandbox/.openclaw/workspace/-rw-r--r-- 1 sandbox sandbox 8110 Jul 22 21:06 AGENTS.md-rw-r--r-- 1 sandbox sandbox 697 Jul 22 21:06 IDENTITY.md # unmodified template-rw-r--r-- 1 sandbox sandbox 1807 Jul 22 21:06 SOUL.md-rw-r--r-- 1 sandbox sandbox 538 Jul 22 21:06 USER.md$ ls /sandbox/.openclaw/workspace/memory/ls: cannot access ...: No such file or directory## Bug B — non-destructive branch (later reboot, could not reproduce loss)$ nemoclaw <name> connect Dashboard port forward to '<name>' is missing or dead. Re-establishing... ✓ Dashboard port forward re-established.$ cat /sandbox/.openclaw/workspace/MARKER.mdrepro-marker-... # survived$ docker inspect --format '{{.Id}}' openshell-<name>-...8fe01f908de0... # unchangedReboot with marker but no `connect`: marker preserved, gateway confirmeddead, container ID unchanged — reboot alone does not lose workspace state.Related: #2042 (recreated pod runs without re-running startup; recovery is aside-effect of `connect`), #910 (request for a `reconnect` command), #4179 /#4197 (installer recreate path lost the same files; fixed).
Checklist
I confirmed this bug is reproducible
I searched existing issues and this is not a duplicate
Investigation Summary
nemoclaw-start, so the in-sandbox OpenClaw gateway does not come back on its own.unless-stoppedrestarts it, forever. ObservedRestartCountof 278 and 299 on two separate reboots. Because it exits inside the 45s healthcheck StartPeriod, it never reachesunhealthyand never serves.nemoclaw <name> connectprinted "Recovering... Recreating the sandbox container..." and replaced the container (new ID confirmed). All/sandbox/.openclaw/workspace/files were reset to image defaults —IDENTITY.mdback to template,memory/gone, all timestamps at container-creation time.connectinstead took a non-destructive "re-establishing dashboard port forward" branch and left the workspace intact (verified with a marker file). A reboot without runningconnectalso preserved the marker — so reboot alone does not lose state; only theconnectrecreate path did. The original loss followed a DGX system upgrade, which may be relevant.unhealthy, but Bug A prevents the container from ever reachingunhealthyon the reboots where it crash-loops — so the two bugs interfere, and I could not drive the container to the state Bug B needs on demand. Bug B is the same defect class as [Ubuntu 24.04][Upgrade] cross-version v0.0.49 -> v0.0.50 with NEMOCLAW_RECREATE_SANDBOX=1 silently loses /sandbox/.openclaw/workspace/ files #4179 (installer recreate path, fixed by fix(onboard): back up workspace state before sandbox recreate #4197).Description
This report covers two related failures that both surface after a host reboot on the same sandbox. They share a root cause — after reboot the container is restored but
nemoclaw-startis not re-run, so the in-sandbox OpenClaw gateway does not restart automatically — but they manifest differently and interfere with each other, so I'm filing them together.Bug A — container exit-0 crash-loop after reboot (reliably reproducible)
On some reboots the sandbox container cannot stay up. The entrypoint starts, exits cleanly (exit 0) within milliseconds, and Docker's
unless-stoppedpolicy restarts it immediately. This repeats indefinitely — I observedRestartCountvalues of 278 and 299 on two separate reboots.docker inspectshowsFinishedAtin the same millisecond asStartedAtwithExitCode=0.Because each restart re-enters the 45-second healthcheck
StartPeriod, the container never transitions tounhealthyand the agent never serves.docker psshows a perpetualUp <seconds> (health: starting)with an uptime that keeps resetting.Bug B —
connectrecovery can recreate the container and lose workspace state (observed once)On a different reboot the container stayed up but the gateway was dead. Running
nemoclaw <name> connectprinted "Recovering... Recreating the sandbox container with its managed startup command..." and replaced the container (new container ID confirmed viadocker inspect). All files under/sandbox/.openclaw/workspace/were reset to image defaults —IDENTITY.md,SOUL.md,USER.md,AGENTS.md, and thememory/directory of daily notes. Sandbox registration, model, provider, policy presets, and channel config were preserved; only workspace state was lost, with no warning or confirmation.Reproducibility caveats for Bug B:
connecttook a non-destructive "Dashboard port forward is missing or dead. Re-establishing..." branch and left the container ID unchanged and the workspace intact (verified with a marker file).connectalso preserved it — so reboot alone does not lose state.unhealthywhenconnectruns (the original incident had ~1h of failed healthchecks). On reboots where Bug A crash-loops, the container never reachesunhealthy, so I could not set up Bug B's precondition on demand.Expected
connectbranch should silently discard workspace state. Either run the same backup/restore phaserebuilduses, or warn and require confirmation before a destructive recreate, or defer torebuild.Bug B is the same defect class as #4179 (closed by #4197), which fixed the installer-driven recreate path.
Reproduction Steps
Bug A (reliable)
docker inspect --format 'restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} exit={{.State.ExitCode}}' <container>Observed: high and climbing
restarts,finished≈started(same ms),exit=0.docker ps --filter name=openshell-<name> --format '{{.Status}}\t{{.RunningFor}}'Observed:
Up <seconds> (health: starting)withRunningForof hours — i.e. uptime keeps resetting.It never reaches
unhealthy.Not every reboot triggers Bug A — 2 of my reboots crash-looped, others left the container up with a dead gateway. This is a consistent issue.
Bug B (precondition reliable, recreate NOT reliably reproducible)
docker exec <container> ps -ef | grep -E 'openclaw|nemoclaw-start'→ emptydocker exec -u sandbox <container> bash -c 'echo marker > /sandbox/.openclaw/workspace/MARKER.md'unhealthy(docker ps→(unhealthy)). <--- This is the part I can't reproduce, which is why I didn't put it as a separate report; unclear to me what leads to this stateThe recreate branch appears to need this; on a healthy/starting container
connecttakes the non-destructive branch instead.docker inspect --format '{{.Id}}' <container>nemoclaw <name> connectand read the banner:cat /sandbox/.openclaw/workspace/MARKER.md(gone in the original incident) and re-check the container ID (new in the original incident).I could not force step 5 down the recreate branch on demand — partly because Bug A prevents the container from reaching the
unhealthystate step 3 needs.Environment
The original Bug B incident occurred on a reboot triggered by a DGX system upgrade. NemoClaw/OpenClaw versions and the sandbox image tag were unchanged across that reboot (verified in
~/.nemoclaw/sandboxes.jsonbefore/after), but I cannot rule out that the upgrade affected container/sandbox state.Same Docker build (29.2.1, a5c7197) and aarch64 / Ubuntu 24.04 NVIDIA-hardware profile as #4179.
Debug Output
[debug] Collecting diagnostics for sandbox ''...
[debug] Quick mode: true
═══ System ═══
Wed Jul 22 03:43:46 PM MDT 2026
Linux 6.17.0-1026-nvidia #26-Ubuntu SMP PREEMPT_DYNAMIC Thu Jun 25 00:57:17 UTC 2026 aarch64 aarch64 aarch64 GNU/Linux
15:43:46 up 1:44, 6 users, load average: 0.41, 0.57, 0.53
total used free shared buff/cache available
Mem: 124610 80341 18291 109 27190 44268
Swap: 16383 0 16383
═══ Processes ═══
20108 2634 VLLM::EngineCore 2.8 4.7
139901 139166 openclaw 0.3 1.2
2219 1 /usr/bin/dockerd -H fd:// - 0.0 0.4
2049 1 /usr/bin/containerd 0.0 0.3
138960 138936 /opt/openshell/bin/openshel 0.0 0.1
107560 1 openshell-gateway[nemoclaw= 0.0 0.1
(desktop/system processes omitted)
═══ GPU ═══
NVIDIA-SMI 580.159.03 Driver Version: 580.159.03 CUDA Version: 13.0
GPU 0: NVIDIA GB10, 37C, P0, 9W
Processes: 20108 C VLLM::EngineCore 68669MiB
═══ Docker ═══
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
8fe01f908de0 6217934eb80f "/opt/openshell/bin/…" 37 minutes ago Up 37 minutes (healthy) openshell--21d5c401-cc45-4e8c-bf8e-cb08d890cf32
c2028dde71b9 9204569b17ee "/bin/bash -lc 'pip …" 26 hours ago Up 2 hours 0.0.0.0:8000->8000/tcp, [::]:8000->8000/tcp nemoclaw-vllm
NOTE: container 8fe01f908de0 was created by the
connectauto-recovery at21:06 UTC. The pre-recovery container was 62c97a5d7b2f. The sandbox itself
was created 2026-07-21 22:06:56.
═══ OpenShell ═══
Server Status
Gateway: nemoclaw
Server: https://127.0.0.1:8080
Status: Connected
Version: 0.0.85
NAME CREATED PHASE
2026-07-21 22:06:56 Ready
Sandbox:
Id: 21d5c401-cc45-4e8c-bf8e-cb08d890cf32
Name:
Phase: Ready
Resource version: 1384
Policy source: sandbox
Revision: 10
Policy filesystem_policy:
include_workdir: true
read_write:
landlock:
compatibility: best_effort
(full network_policies block omitted for brevity — presets: npm, pypi,
huggingface, brew, brave, local-inference, openclaw-pricing, discord)
Log excerpt (post-recovery steady state; the incident window predates this
capture). Representative lines only — the full output is ~200 lines of
these repeating:
[sandbox] [WARN ] [openshell_supervisor_process::ssh] data on unknown channel ChannelId(27)
[sandbox] [OCSF ] NET:OTHER [INFO] ALLOWED gateway.discord.gg:443 [policy:discord engine:l7-websocket]
[sandbox] [OCSF ] NET:OPEN [INFO] ALLOWED inference.local:443
[sandbox] [INFO ] [openshell_router] routing proxy inference request endpoint=http://host.openshell.internal:8000/v1 method=GET path=/v1/models
═══ Onboard Session ═══
{
"version": 1,
"sessionId": "(redacted)",
"status": "complete",
"resumable": false,
"mode": "interactive",
"startedAt": "2026-07-21T21:38:26.963Z",
"updatedAt": "2026-07-21T22:10:19.393Z",
"sandboxName": "",
"provider": "vllm-local",
"model": "nvidia/Qwen3.6-35B-A3B-NVFP4",
"endpointUrl": "http://host.openshell.internal:8000/v1",
"credentialEnv": null,
"preferredInferenceApi": "openai-completions",
"toolDisclosure": "progressive",
"observabilityEnabled": false,
"policyPresets": ["npm","pypi","huggingface","brew","brave",
"local-inference","openclaw-pricing","discord"],
"gpuPassthrough": true,
"lastCompletedStep": "policies",
"failure": null,
"steps": { (all complete; agent_setup skipped) }
}
═══ Sandbox Internals ═══
UID PID PPID C STIME TTY TIME CMD
root 1 0 0 21:06 ? 00:00:03 /opt/openshell/bin/openshell-sandbox
sandbox 111 1 0 21:06 ? 00:00:00 bash /usr/local/bin/nemoclaw-start
sandbox 628 111 1 21:06 ? 00:00:27 openclaw
sandbox 643 111 0 21:06 ? 00:00:00 bash /usr/local/bin/nemoclaw-start
sandbox 645 643 0 21:06 ? 00:00:00 tail -n +1 -F /tmp/gateway.log
sandbox 659 111 0 21:06 ? 00:00:00 python3 -u -
NOTE: all in-sandbox processes start at 21:06 — the recovery time, not the
original sandbox creation time. PID 1 (the supervisor) included, which
confirms a fresh container rather than a restart.
═══ Kernel Messages ═══
(skipped: dmesg_restrict=1)
Logs
Checklist