Skip to content

[Ubuntu 24.04][Sandbox] After reboot: sandbox container exit-0 crash-loops (reliable) and connect recovery can recreate the container and lose /sandbox/.openclaw/workspace/ (observed once) #7404

Description

@rlacatus

Investigation Summary

  • Two related post-reboot failures on the same sandbox. Common root cause: after a host reboot the container is restored by Docker but nothing re-runs nemoclaw-start, so the in-sandbox OpenClaw gateway does not come back on its own.
  • Bug A (reliable): on some reboots the container exit-0 crash-loops — entrypoint starts, exits 0 within milliseconds, unless-stopped restarts it, forever. Observed RestartCount of 278 and 299 on two separate reboots. Because it exits inside the 45s healthcheck StartPeriod, it never reaches unhealthy and never serves.
  • Bug B (observed once): on a different reboot the container stayed up but the gateway was dead; running nemoclaw <name> connect printed "Recovering... Recreating the sandbox container..." and replaced the container (new ID confirmed). All /sandbox/.openclaw/workspace/ files were reset to image defaults — IDENTITY.md back to template, memory/ gone, all timestamps at container-creation time.
  • Bug B is not reliably reproducible. Across three deliberate reboots I could not re-trigger the recreate; connect instead took a non-destructive "re-establishing dashboard port forward" branch and left the workspace intact (verified with a marker file). A reboot without running connect also preserved the marker — so reboot alone does not lose state; only the connect recreate path did. The original loss followed a DGX system upgrade, which may be relevant.
  • The recreate branch appears gated on the container being unhealthy, but Bug A prevents the container from ever reaching unhealthy on the reboots where it crash-loops — so the two bugs interfere, and I could not drive the container to the state Bug B needs on demand. Bug B is the same defect class as [Ubuntu 24.04][Upgrade] cross-version v0.0.49 -> v0.0.50 with NEMOCLAW_RECREATE_SANDBOX=1 silently loses /sandbox/.openclaw/workspace/ files #4179 (installer recreate path, fixed by fix(onboard): back up workspace state before sandbox recreate #4197).

Description

This report covers two related failures that both surface after a host reboot on the same sandbox. They share a root cause — after reboot the container is restored but nemoclaw-start is not re-run, so the in-sandbox OpenClaw gateway does not restart automatically — but they manifest differently and interfere with each other, so I'm filing them together.

Bug A — container exit-0 crash-loop after reboot (reliably reproducible)

On some reboots the sandbox container cannot stay up. The entrypoint starts, exits cleanly (exit 0) within milliseconds, and Docker's unless-stopped policy restarts it immediately. This repeats indefinitely — I observed RestartCount values of 278 and 299 on two separate reboots. docker inspect shows FinishedAt in the same millisecond as StartedAt with ExitCode=0.

Because each restart re-enters the 45-second healthcheck StartPeriod, the container never transitions to unhealthy and the agent never serves. docker ps shows a perpetual Up <seconds> (health: starting) with an uptime that keeps resetting.

Bug B — connect recovery can recreate the container and lose workspace state (observed once)

On a different reboot the container stayed up but the gateway was dead. Running nemoclaw <name> connect printed "Recovering... Recreating the sandbox container with its managed startup command..." and replaced the container (new container ID confirmed via docker inspect). All files under /sandbox/.openclaw/workspace/ were reset to image defaults — IDENTITY.md, SOUL.md, USER.md, AGENTS.md, and the memory/ directory of daily notes. Sandbox registration, model, provider, policy presets, and channel config were preserved; only workspace state was lost, with no warning or confirmation.

Reproducibility caveats for Bug B:

  • The gateway-down precondition reproduces on every reboot.
  • The destructive recreate itself does not. In three deliberate reboot tests, connect took a non-destructive "Dashboard port forward is missing or dead. Re-establishing..." branch and left the container ID unchanged and the workspace intact (verified with a marker file).
  • A reboot with the marker present but no connect also preserved it — so reboot alone does not lose state.
  • The recreate branch appears to require the container to be unhealthy when connect runs (the original incident had ~1h of failed healthchecks). On reboots where Bug A crash-loops, the container never reaches unhealthy, so I could not set up Bug B's precondition on demand.
  • The original loss followed a DGX system upgrade, which may have influenced container state.

Expected

  • Bug A: after reboot, the container should start its agent and stay up, or fail loudly — not exit 0 and churn silently forever.
  • Bug B: no connect branch should silently discard workspace state. Either run the same backup/restore phase rebuild uses, or warn and require confirmation before a destructive recreate, or defer to rebuild.

Bug B is the same defect class as #4179 (closed by #4197), which fixed the installer-driven recreate path.

Reproduction Steps

Bug A (reliable)

  1. Onboard a sandbox (local vLLM inference, GPU enabled).
  2. Reboot the host.
  3. After boot: docker inspect --format 'restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} exit={{.State.ExitCode}}' <container>
    Observed: high and climbing restarts, finished ≈ started (same ms), exit=0.
  4. docker ps --filter name=openshell-<name> --format '{{.Status}}\t{{.RunningFor}}'
    Observed: Up <seconds> (health: starting) with RunningFor of hours — i.e. uptime keeps resetting.
    It never reaches unhealthy.

Not every reboot triggers Bug A — 2 of my reboots crash-looped, others left the container up with a dead gateway. This is a consistent issue.

Bug B (precondition reliable, recreate NOT reliably reproducible)

  1. On a reboot where the container stays up (Bug A did NOT trigger), confirm the gateway is down:
    • docker exec <container> ps -ef | grep -E 'openclaw|nemoclaw-start' → empty
  2. Write a marker so loss is detectable:
    docker exec -u sandbox <container> bash -c 'echo marker > /sandbox/.openclaw/workspace/MARKER.md'
  3. Wait for the container to become unhealthy (docker ps → (unhealthy)). <--- This is the part I can't reproduce, which is why I didn't put it as a separate report; unclear to me what leads to this state
    The recreate branch appears to need this; on a healthy/starting container connect takes the non-destructive branch instead.
  4. Record the container ID: docker inspect --format '{{.Id}}' <container>
  5. Run nemoclaw <name> connect and read the banner:
    • "Recovering... Recreating the sandbox container..." = destructive (observed once).
    • "Dashboard port forward is missing or dead. Re-establishing..." = non-destructive (observed 3×).
  6. If recreated: cat /sandbox/.openclaw/workspace/MARKER.md (gone in the original incident) and re-check the container ID (new in the original incident).

I could not force step 5 down the recreate branch on demand — partly because Bug A prevents the container from reaching the unhealthy state step 3 needs.

Environment

Device:        NVIDIA DGX Spark
OS:            Ubuntu 24.04.4 LTS
Kernel:        6.17.0-1026-nvidia
Architecture:  aarch64
Node.js:       v22.23.1
Docker:        Docker version 29.2.1, build a5c7197
OpenShell CLI: openshell 0.0.85 (docker driver)
NemoClaw:      v0.0.88
OpenClaw:      2026.6.10
NVIDIA driver: 580.159.03 (CUDA 13.0), GB10
Inference:     vllm-local, nvidia/Qwen3.6-35B-A3B-NVFP4

The original Bug B incident occurred on a reboot triggered by a DGX system upgrade. NemoClaw/OpenClaw versions and the sandbox image tag were unchanged across that reboot (verified in ~/.nemoclaw/sandboxes.json before/after), but I cannot rule out that the upgrade affected container/sandbox state.

Same Docker build (29.2.1, a5c7197) and aarch64 / Ubuntu 24.04 NVIDIA-hardware profile as #4179.

Debug Output

[debug] Collecting diagnostics for sandbox ''...
[debug] Quick mode: true

═══ System ═══

Wed Jul 22 03:43:46 PM MDT 2026
Linux 6.17.0-1026-nvidia #26-Ubuntu SMP PREEMPT_DYNAMIC Thu Jun 25 00:57:17 UTC 2026 aarch64 aarch64 aarch64 GNU/Linux
15:43:46 up 1:44, 6 users, load average: 0.41, 0.57, 0.53
total used free shared buff/cache available
Mem: 124610 80341 18291 109 27190 44268
Swap: 16383 0 16383

═══ Processes ═══

PID    PPID CMD                         %MEM %CPU

20108 2634 VLLM::EngineCore 2.8 4.7
139901 139166 openclaw 0.3 1.2
2219 1 /usr/bin/dockerd -H fd:// - 0.0 0.4
2049 1 /usr/bin/containerd 0.0 0.3
138960 138936 /opt/openshell/bin/openshel 0.0 0.1
107560 1 openshell-gateway[nemoclaw= 0.0 0.1
(desktop/system processes omitted)

═══ GPU ═══

NVIDIA-SMI 580.159.03 Driver Version: 580.159.03 CUDA Version: 13.0
GPU 0: NVIDIA GB10, 37C, P0, 9W
Processes: 20108 C VLLM::EngineCore 68669MiB

═══ Docker ═══

CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
8fe01f908de0 6217934eb80f "/opt/openshell/bin/…" 37 minutes ago Up 37 minutes (healthy) openshell--21d5c401-cc45-4e8c-bf8e-cb08d890cf32
c2028dde71b9 9204569b17ee "/bin/bash -lc 'pip …" 26 hours ago Up 2 hours 0.0.0.0:8000->8000/tcp, [::]:8000->8000/tcp nemoclaw-vllm

NOTE: container 8fe01f908de0 was created by the connect auto-recovery at
21:06 UTC. The pre-recovery container was 62c97a5d7b2f. The sandbox itself
was created 2026-07-21 22:06:56.

═══ OpenShell ═══

Server Status

Gateway: nemoclaw
Server: https://127.0.0.1:8080
Status: Connected
Version: 0.0.85
NAME CREATED PHASE
2026-07-21 22:06:56 Ready

Sandbox:
Id: 21d5c401-cc45-4e8c-bf8e-cb08d890cf32
Name:
Phase: Ready
Resource version: 1384
Policy source: sandbox
Revision: 10

Policy filesystem_policy:
include_workdir: true
read_write:

  • /tmp
  • /dev/null
  • /dev/pts
  • /sandbox/.openclaw
  • /sandbox/.nemoclaw
  • /home/linuxbrew
  • /sandbox
  • (nvidia device nodes)
  • /proc
    landlock:
    compatibility: best_effort

(full network_policies block omitted for brevity — presets: npm, pypi,
huggingface, brew, brave, local-inference, openclaw-pricing, discord)

Log excerpt (post-recovery steady state; the incident window predates this
capture). Representative lines only — the full output is ~200 lines of
these repeating:

[sandbox] [WARN ] [openshell_supervisor_process::ssh] data on unknown channel ChannelId(27)
[sandbox] [OCSF ] NET:OTHER [INFO] ALLOWED gateway.discord.gg:443 [policy:discord engine:l7-websocket]
[sandbox] [OCSF ] NET:OPEN [INFO] ALLOWED inference.local:443
[sandbox] [INFO ] [openshell_router] routing proxy inference request endpoint=http://host.openshell.internal:8000/v1 method=GET path=/v1/models

═══ Onboard Session ═══

{
"version": 1,
"sessionId": "(redacted)",
"status": "complete",
"resumable": false,
"mode": "interactive",
"startedAt": "2026-07-21T21:38:26.963Z",
"updatedAt": "2026-07-21T22:10:19.393Z",
"sandboxName": "",
"provider": "vllm-local",
"model": "nvidia/Qwen3.6-35B-A3B-NVFP4",
"endpointUrl": "http://host.openshell.internal:8000/v1",
"credentialEnv": null,
"preferredInferenceApi": "openai-completions",
"toolDisclosure": "progressive",
"observabilityEnabled": false,
"policyPresets": ["npm","pypi","huggingface","brew","brave",
"local-inference","openclaw-pricing","discord"],
"gpuPassthrough": true,
"lastCompletedStep": "policies",
"failure": null,
"steps": { (all complete; agent_setup skipped) }
}

═══ Sandbox Internals ═══

UID PID PPID C STIME TTY TIME CMD
root 1 0 0 21:06 ? 00:00:03 /opt/openshell/bin/openshell-sandbox
sandbox 111 1 0 21:06 ? 00:00:00 bash /usr/local/bin/nemoclaw-start
sandbox 628 111 1 21:06 ? 00:00:27 openclaw
sandbox 643 111 0 21:06 ? 00:00:00 bash /usr/local/bin/nemoclaw-start
sandbox 645 643 0 21:06 ? 00:00:00 tail -n +1 -F /tmp/gateway.log
sandbox 659 111 0 21:06 ? 00:00:00 python3 -u -

NOTE: all in-sandbox processes start at 21:06 — the recovery time, not the
original sandbox creation time. PID 1 (the supervisor) included, which
confirms a fresh container rather than a restart.

═══ Kernel Messages ═══

(skipped: dmesg_restrict=1)

Logs

## Bug A — crash-loop


$ docker inspect --format 'restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} exit={{.State.ExitCode}}' openshell-<name>-...
restarts=299 started=2026-07-22T23:15:16.829Z finished=2026-07-22T23:15:16.687Z exit=0

$ docker ps -a --filter name=openshell-<name> --format '{{.Status}}\t{{.RunningFor}}'
Up 10 seconds (health: starting)    2 hours ago


`finished` precedes `started` by ~140ms across the restart boundary; exit 0;
299 restarts. A separate earlier reboot showed the same signature at
restarts=278.

Healthcheck definition — StartPeriod (45s) exceeds the container's lifetime,
so it can never leave `starting`:


Interval=30s Timeout=5s StartPeriod=45s Retries=3


## Bug B — recreate + workspace loss (original incident)


$ docker inspect --format '{{.Id}} created={{.Created}}' openshell-<name>-...
8fe01f908de0... created=2026-07-22T21:06:04Z
# prior container ID (docker events, pre-recovery): 62c97a5d7b2f...

$ ls -la /sandbox/.openclaw/workspace/
-rw-r--r-- 1 sandbox sandbox 8110 Jul 22 21:06 AGENTS.md
-rw-r--r-- 1 sandbox sandbox  697 Jul 22 21:06 IDENTITY.md   # unmodified template
-rw-r--r-- 1 sandbox sandbox 1807 Jul 22 21:06 SOUL.md
-rw-r--r-- 1 sandbox sandbox  538 Jul 22 21:06 USER.md
$ ls /sandbox/.openclaw/workspace/memory/
ls: cannot access ...: No such file or directory


## Bug B — non-destructive branch (later reboot, could not reproduce loss)


$ nemoclaw <name> connect
  Dashboard port forward to '<name>' is missing or dead.
  Re-establishing...
  ✓ Dashboard port forward re-established.
$ cat /sandbox/.openclaw/workspace/MARKER.md
repro-marker-...          # survived
$ docker inspect --format '{{.Id}}' openshell-<name>-...
8fe01f908de0...           # unchanged


Reboot with marker but no `connect`: marker preserved, gateway confirmed
dead, container ID unchanged — reboot alone does not lose workspace state.

Related: #2042 (recreated pod runs without re-running startup; recovery is a
side-effect of `connect`), #910 (request for a `reconnect` command), #4179 /
#4197 (installer recreate path lost the same files; fixed).

Checklist

  • I confirmed this bug is reproducible
  • I searched existing issues and this is not a duplicate

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area: onboardingOnboarding FSM, provider setup, sandbox launch, or first-run flowarea: sandboxOpenShell sandbox lifecycle, runtime, config, or recoveryplatform: containerAffects Docker, containerd, Podman, or imagesplatform: ubuntuAffects Ubuntu Linux environments

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions