Skip to content

feat: add native Windows support (server + windows/amd64 llama-cpp backend) - #11429

Open
LionelColaso wants to merge 1 commit into
mudler:masterfrom
LionelColaso:windows-backends-llama-cpp
Open

LionelColaso wants to merge 1 commit into
mudler:masterfrom
LionelColaso:windows-backends-llama-cpp

Conversation

@LionelColaso

@LionelColaso LionelColaso commented Aug 9, 2026 •

Copy link
Copy Markdown

Description

Adds native Windows support end to end: LocalAI now releases a windows/amd64 server binary, and the llama-cpp backend is built natively for windows/amd64 under MSYS2 UCRT64 and packaged as an OCI image tar that LocalAI installs and runs as a native process - no docker daemon or WSL required on the host.

Server side

  • goreleaser: add windows (amd64/arm64) to the release targets
  • Makefile: download the win64 protoc zip and rename protoc.exe to protoc, resolve code-gen plugins via --plugin instead of PATH, force SHELL=sh and name the binary local-ai.exe on Windows, ignore protoc.exe
  • build-test.yaml: add a native windows-latest build gate that installs GNU make via Chocolatey, adds Git for Windows' usr/bin to PATH and builds with CGO_ENABLED=0
  • pkg/downloader: close the write handle before removing or renaming a partial download so Windows file locks do not break the resume and error paths; guard the POSIX-permission and symlink tests on non-Windows
  • tests: Windows guards and path fixes for core/gallery, video_internal, loader and the testcontainers database setup

Backend side

  • scripts/build/llama-cpp-windows.sh: builds gRPC from source (pinned v1.59.0, with mingw-w64 fixes for c-ares, boringssl and zlib), then the three llama.cpp variants (cpu-all with GGML_CPU_ALL_VARIANTS + Vulkan, rpc, and the AVX-off fallback), bundles the mingw runtime DLLs and ships an OCI tar via local-ai util create-oci-image. Re-runnable and auto-dispatches into MSYS2 when launched from Git for Windows' bash. JOBS override supported for memory-limited hosts.
  • core/gallery: backends install skips the OCI-registry digest lookup for ocifile:// streams (a local tarball is not a registry reference), so the local-build install step no longer prints a confusing could not parse reference digest warning
  • backend/cpp/llama-cpp/run.ps1: PowerShell launcher (mirrors run.sh) that pkg/model/process.go starts on Windows
  • backend/index.yaml: windows/amd64 backend entry and variants
  • pkg/system/capabilities.go: windows engine preference rules so the gallery picks the native build on Windows hosts
  • .github/backend-matrix.yml + backend_build_windows.yml: windows matrix entries and a reusable windows build workflow; backend.yml and backend_pr.yml wire the windows backend jobs (build on PR, publish on master)
  • docs/content/getting-started/windows.md plus related page updates (GPU-acceleration, install)

Notes for Reviewers

  • The Windows backend image ships three llama-cpp executables picked by the run.ps1 launcher: llama-cpp-cpu-all.exe (all ggml CPU variants, auto-detects a Vulkan device at runtime and falls back to CPU), llama-cpp-grpc.exe (gRPC-RPC build, selected when LLAMACPP_GRPC_SERVERS is set) and llama-cpp-fallback.exe (static, AVX-off fallback). The mingw runtime DLLs are bundled in the image, so no MSYS2 install is needed on the host.
  • macOS is intentionally unaffected; the Darwin matrix entries are unchanged.
  • Review follow-ups (richiejp): backend processes now run inside a Windows Job Object with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE (pkg/model/windows_job_windows.go), so model unloads, graceful shutdown, and an abrupt local-ai.exe exit all reap the wrapper + backend tree; it degrades to a logged warning when the host already nests the process in a non-breakaway job.
  • Windows smoke suite (tests/e2e/windows): builds and boots the real local-ai.exe with the mock-backend laid out exactly as a gallery install ships it (run.sh for discovery, run.ps1 to launch) and asserts a chat completion via run.ps1 plus job-object tree reaping after hard-killing local-ai.exe. Self-skips on non-Windows hosts; wired as make test-windows-smoke and a windows-latest job in tests-e2e.yml. 2/2 specs pass locally on a Windows host and on the tests-windows-smoke CI job.
  • backend_build_windows.yml's publish job now signs the windows images keyless with cosign (same flow as backend_merge.yml); the gallery verification: block stays unpopulated until it can cover every published variant.
  • Commits are DCO-signed and carry an Assisted-by: trailer per .agents/ai-coding-assistants.md.

Signed commits

  • Yes, I signed my commits.
  • Documentation updated (docs/content/) for user-facing changes, or not applicable

Closes #2368

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 2 times, most recently from 6caf372 to b347f46 Compare August 9, 2026 18:16
@LionelColaso LionelColaso changed the title feat: add native Windows support with a windows/amd64 llama-cpp backend feat: add native Windows support (server + windows/amd64 llama-cpp backend) Aug 9, 2026
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from b347f46 to c16c463 Compare August 9, 2026 18:25
Comment thread backend/cpp/llama-cpp/run-windows/main.go Outdated
Comment thread pkg/model/process.go Outdated
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 3 times, most recently from b6bbe35 to 4e16c60 Compare August 10, 2026 02:45

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mudler the Windows launcher concerns are resolved on the rebased head: the unused run.cmd is removed, the native run.exe path is now explained, and backends/llama-cpp-windows is serialized in .NOTPARALLEL. I reviewed the range-diff from the prior head; git diff --check, the launcher build, pkg/system, and all 48 backend-filter tests pass. Good from my side once the newly restarted Windows/backend CI completes.

@LionelColaso
LionelColaso requested a review from mudler August 11, 2026 13:41
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 2 times, most recently from 319917b to b269dde Compare August 12, 2026 09:29
@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

@mudler the requested launcher change is addressed on the rebased head: the compiled run.exe helper is gone, run.ps1 mirrors the binary selection and PATH setup, and pkg/model/process.go invokes it through Windows PowerShell. I isolated the patch from the rebase and verified git diff --check, all 48 backend-filter tests, and pkg/system; the contributor also reports native Windows launcher coverage for selection, argument forwarding, paths with spaces, and exit propagation. Good from my side; only DCO is currently reported, so the Windows/backend workflow result is still pending.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from b269dde to 7705d38 Compare August 15, 2026 06:47
@localai-org-maint-bot
localai-org-maint-bot force-pushed the windows-backends-llama-cpp branch 2 times, most recently from e9f247d to 547ff9d Compare August 19, 2026 15:13
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 3 times, most recently from 9c1fe20 to 788603f Compare August 21, 2026 05:34
@localai-org-maint-bot
localai-org-maint-bot force-pushed the windows-backends-llama-cpp branch 2 times, most recently from e353b89 to c334688 Compare September 8, 2026 22:05
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from c334688 to 4c2f267 Compare September 9, 2026 04:41

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Found a runtime blocker: pkg/model/process.go sets process.WithName("powershell.exe") (a bare filename) together with process.WithWorkDir(workDir). On Windows, Go's syscall.StartProcess resolves the application name relative to attr.Dir when both are set — so it looks for <workDir>\powershell.exe instead of searching PATH, and CreateProcess with an explicit lpApplicationName does no PATH lookup. Every Windows backend launch will fail with ERROR_FILE_NOT_FOUND.

Fix: resolve the full path before calling WithName, e.g. exec.LookPath("powershell.exe") (which searches PATH and finds System32), or hardcode C:\\Windows\\System32\\WindowsPowerShell\\v1.0\\powershell.exe.

The rest of the PR looks sound — capability detection ordering, docs, CI workflows. Once this is fixed, it should be good to merge.

@richiejp

Copy link
Copy Markdown
Collaborator

I would just resolve the merge conflict and enable auot-merge, but there appear to be a few issues

• P1 — Backend survives shutdown. backend/cpp/llama-cpp/run.ps1:30 launches the backend as a PowerShell child. LocalAI tracks PowerShell’s PID, and the pinned process
manager kills only that process. Unloading models or shutting down can leave inference processes holding memory and ports. Launch the executable directly or manage the
process tree with a Windows Job Object.

• P2 — Misleading firewall advice. docs/content/getting-started/windows.md:92 claims localhost-only listening and recommends allowing firewall access. The actual default
is :8080—all interfaces—with no API key configured by default. Document an explicit loopback bind for local use.

• Integrity gap — Windows images aren’t signed. The publishing workflow (.github/workflows/backend_build_windows.yml:246) pushes images without cosign signing, despite
the guide claiming a verification policy. Signature-enforcing installations cannot use these artifacts. Add signing and a matching policy, and correct the
documentation.

There is one merge conflict, in core/gallery/backends.go:437. Keep master’s implementation:

Also it would make sense to add some Windows tests for basic things like starting LocalAI and doing some requests on a mock backend using the existing infrastructure for doing that as far as possible.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from 0d12a6e to 6602a1c Compare September 26, 2026 20:10
@LionelColaso

LionelColaso commented Sep 27, 2026 •

Copy link
Copy Markdown
Author

@richiejp thanks for the thorough review. All four points are addressed on the rebased head, and the conflict note is confirmed:

Merge conflict (core/gallery/backends.go:437) - confirmed, master's implementation was kept. The rebase swallowed our ocifile:// skip into master's uri.LooksLikeRegistryOCI() (uri.go), so the digest lookup sites (backends.go, upgrade.go) now exclude ocifile:// natively and our local !uri.LooksLikeOCIFile() guard became redundant. Behavior is preserved and covered by the ocifile:// digest-lookup specs, which pass on the current head.

P1 - Backend survives shutdown - implemented with a Windows Job Object. In createProcess the spawned PowerShell wrapper is assigned to a job configured with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE (pkg/model/windows_job_windows.go), before run.ps1 spawns the backend, so the binary inherits the membership. All three stop paths (deleteProcess, the graceful-shutdown handler, stopLoadProcess) now terminate the job after process.Stop() kills the wrapper, reaping the whole tree; kill-on-close also covers the case where local-ai.exe itself dies abruptly. It degrades to a logged warning if the host already nested the process in a non-breakaway job. A Windows-only unit test (TestWindowsJobObjectKillsProcessTree) parks a wrapper until it is assigned to the job, then spawns a child, and asserts terminating the job reaps both - it passes locally.

P2 - Misleading firewall advice - fixed. The docs now state the real :8080 all-interfaces default, recommend ./local-ai.exe --address 127.0.0.1:8080 --api-key your-key-here for local use, and advise denying the firewall prompt for loopback-only setups.

P3 - Windows images aren't signed - fixed: backend_build_windows.yml's publish job now installs cosign and, after each crane push, resolves the digest and signs it keyless (--recursive --new-bundle-format --registry-referrers-mode=oci-1-1, id-token: write, COSIGN_EXPERIMENTAL=1), matching backend_merge.yml. I deferred the gallery verification: block deliberately: it is gallery-level, so enforcing it now would also block the unsigned linux/cuda/rocm variants - the 4.3 blog post says the blocks are intentionally unpopulated until a gallery covers every published variant. The docs now say exactly that instead of overclaiming.

Windows tests - a job-object unit test (TestWindowsJobObjectKillsProcessTree) and a standalone Windows-host smoke suite (tests/e2e/windows) that builds and boots the real local-ai.exe with the mock-backend laid out exactly as a gallery install ships it - run.sh for discovery plus run.ps1 to launch - covering both halves of the Windows-native path:

  • a chat completion served by the backend started via run.ps1
  • a hard-kill of local-ai.exe asserting the wrapper + backend tree is reaped by the job object (kill-on-close), with the tree identified by parentage from the server PID so unrelated processes on a dev box cannot false-positive

Wired as make test-windows-smoke and a windows-latest job in tests-e2e.yml (self-skips on non-Windows hosts). Both specs pass locally on a Windows host (2/2) and on the tests-windows-smoke CI job (windows-latest); pkg/model stays at its known 152-pass baseline there.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 7 times, most recently from b4fd355 to 6d92e74 Compare October 3, 2026 11:54
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 3 times, most recently from 56f8395 to 27f5ea2 Compare October 10, 2026 06:16
Comment thread pkg/model/process.go Outdated
Comment thread pkg/model/process_runtime.go Outdated
@mudler

mudler commented Oct 10, 2026

Copy link
Copy Markdown
Owner

Just small code nits on my side, but otherwise the direction looks good. Unfortunately, I do not have windows to test this, so we will have to see via CI and user reports how this behaves in real hardware.

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 2 times, most recently from 9a6c223 to 80deff1 Compare October 10, 2026 15:24
Comment thread pkg/model/process.go Outdated
Comment thread pkg/model/process.go Outdated
Comment thread pkg/model/process.go Outdated

@mudler mudler left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we are almost there :) thanks for the patience @LionelColaso. Last review pass and then should be good to go

@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch 4 times, most recently from 8320fb6 to 7bc3106 Compare October 10, 2026 17:21
…ckend)

Native Windows support end to end: LocalAI releases a windows/amd64 server
binary and the llama-cpp backend is built natively for windows/amd64 under
MSYS2 UCRT64 and packaged as an OCI image tar that LocalAI installs and
runs as a native process - no docker daemon or WSL required on the host.

Server side:
- goreleaser: add windows (amd64/arm64) to the release targets
- Makefile: download the win64 protoc zip and rename protoc.exe to protoc,
  resolve code-gen plugins via --plugin instead of PATH, force SHELL=sh and
  name the binary local-ai.exe on Windows, ignore protoc.exe
- build-test.yaml: add a native windows-latest build gate that installs GNU
  make via choco, adds Git for Windows' usr/bin to PATH and builds with
  CGO_ENABLED=0
- pkg/downloader: close the write handle before removing or renaming the
  partial so Windows file locks do not break resume and error paths; guard
  the POSIX-permission and symlink tests on non-Windows
- tests: Windows guards and path fixes for core/gallery, video_internal,
  loader and the testcontainers database setup

Backend side:
- scripts/build/llama-cpp-windows.sh: builds gRPC from source (pinned
  v1.59.0, with mingw-w64 fixes for c-ares, boringssl and zlib), then the
  three llama.cpp variants (cpu-all with GGML_CPU_ALL_VARIANTS + Vulkan,
  rpc, and the AVX-off fallback), bundles the mingw runtime DLLs and ships
  an OCI tar via local-ai util create-oci-image. Re-runnable and
  auto-dispatchs into MSYS2 when launched from Git for Windows' bash.
  make backends/llama-cpp-windows hands the build to
  scripts/build/llama-cpp-windows.ps1 on Windows: it locates MSYS2 (or, after
  asking for confirmation, installs it via winget and pacman-installs the
  mingw-w64-ucrt toolchain) and runs the sh script under MSYS2 UCRT64 bash.
- backend/cpp/llama-cpp/run.ps1: PowerShell launcher (mirrors run.sh) that
  pkg/model starts on Windows. The launcher runs through the Windows
  PowerShell resolved via exec.LookPath, because os.StartProcess resolves
  argv0 against the backend workDir on a Windows host.
- backend/index.yaml: windows/amd64 backend entry and variants.
- pkg/system/capabilities.go: windows engine preference rules so the
  gallery picks the native build on Windows hosts.
- .github/backend-matrix.yml + backend_build_windows.yml: windows matrix
  entries and a reusable windows build workflow; backend.yml and
  backend_pr.yml wire the windows backend jobs (build on PR, publish on
  master).
- core/gallery: backends install skips the OCI-registry digest lookup for
  ocifile:// streams (a local tarball is not a registry reference), so the
  local-build install step no longer prints a confusing "could not parse
  reference" digest warning.
- docs: getting-started/windows.md plus related page updates.

JOBS in the build script honors an override so memory-limited hosts can
build with reduced parallelism.

Process-tree cleanup (Windows job objects):
- pkg/model/process_tree.go declares a processTree interface; the Windows
  implementation (process_tree_windows.go, kill-on-close job object) and the
  no-op for other platforms (process_tree_other.go) are selected at build
  time, keeping Windows-specific code out of the shared process runtime.
- The teardown is wired into the stop paths so the backend tree cannot
  outlive a deliberate unload.

Launcher resolution:
- pkg/model/process_launcher_windows.go / process_launcher_other.go resolve
  the executable and argv per platform: other platforms spawn the
  gallery-contract run.sh stub as-is, Windows substitutes the bundled
  run.ps1 through the system PowerShell. The shared startProcess calls
  resolveLauncher, so no runtime.GOOS branch remains in generic code.

Windows smoke test:
- tests/e2e/windows: standalone ginkgo suite (skips itself on non-Windows)
  that builds and boots the real local-ai.exe (LOCAL_AI_EXE override for
  pre-built binaries) with the mock-backend laid out as the gallery ships
  it - run.sh for discovery plus run.ps1 to launch. Two specs:
  a chat completion through the backend launched via run.ps1, and a
  hard-kill of local-ai.exe with an assertion that the wrapper + backend
  tree is reaped via the job object (kill-on-close), identifying the tree
  by parentage from the server PID rather than command-line text so
  unrelated processes on a developer machine cannot false-positive.
- Makefile: test-windows-smoke target (protogen-go + react-ui + ginkgo on
  the suite).
- tests-e2e.yml: windows-latest smoke job mirroring the ubuntu e2e job's
  proto setup (win64 protoc + plugins), then builds both binaries with
  CGO_ENABLED=0 and runs the suite.

Assisted-by: opencode:big-pickle
Signed-off-by: Lionel Colaso <lionelcolaso@outlook.com>
@LionelColaso
LionelColaso force-pushed the windows-backends-llama-cpp branch from 7bc3106 to 12cbf4e Compare October 10, 2026 17:26

@mudler-agent mudler-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two findings from the Windows support review. Rechecked against the current head; the affected code is unchanged.

git checkout -q -B build "$LLAMA_VERSION"
git submodule update --init --recursive --depth 1 --single-branch
cd "$ROOT/backend/cpp/llama-cpp"
bash prepare.sh

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Create the gRPC staging directory before calling prepare.sh

prepare.sh copies upstream server files into llama.cpp/tools/grpc-server/ under set -e, but this script never creates that directory. The preceding git clean -qfd also removes staging left by an earlier run. A fresh backend build therefore exits during source preparation, before compiling the llama.cpp variants.

I reproduced the failure with the unchanged helper and a minimal upstream server fixture: cp: cannot create regular file 'llama.cpp/tools/grpc-server/': Not a directory. The existing backend Makefile creates the directory before invoking the helper. Please add mkdir -p llama.cpp/tools/grpc-server here before bash prepare.sh.

Comment thread pkg/model/process.go
// tracker so that stopping it reaps the whole tree. The mechanism is
// platform specific; see the processTree interface.
if pid, err := strconv.Atoi(grpcControlProcess.CurrentPID()); err == nil {
runtime.trackProcessTree(pid)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Assign the Windows job before the launcher can spawn children

grpcControlProcess.Run() starts PowerShell before this job assignment. If LocalAI is descheduled after starting the wrapper, PowerShell can launch the backend first. Assigning an already-running parent to a job does not retroactively enroll its existing children, so a later unload or abrupt LocalAI exit can leave the backend alive with its memory/GPU allocation and listening socket.

windows_job_test.go explicitly gates child creation until after assignment, but the production launcher has no equivalent synchronization. The llama.cpp parent-death watcher is a no-op on Windows, so it does not cover this case.

Please enforce the ordering, for example by creating the launcher suspended, assigning it to the job, then resuming it, or by adding a startup handshake before child creation. A regression test should exercise that production ordering. This finding is based on source analysis; I have not reproduced the race on a Windows host.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Native windows version?

6 participants