Repository navigation
feat: add native Windows support (server + windows/amd64 llama-cpp backend) - #11429
LionelColaso wants to merge 1 commit into
Conversation
6caf372 to
b347f46
Compare
b347f46 to
c16c463
Compare
b6bbe35 to
4e16c60
Compare
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
@mudler the Windows launcher concerns are resolved on the rebased head: the unused run.cmd is removed, the native run.exe path is now explained, and backends/llama-cpp-windows is serialized in .NOTPARALLEL. I reviewed the range-diff from the prior head; git diff --check, the launcher build, pkg/system, and all 48 backend-filter tests pass. Good from my side once the newly restarted Windows/backend CI completes.
319917b to
b269dde
Compare
|
@mudler the requested launcher change is addressed on the rebased head: the compiled |
b269dde to
7705d38
Compare
e9f247d to
547ff9d
Compare
9c1fe20 to
788603f
Compare
e353b89 to
c334688
Compare
c334688 to
4c2f267
Compare
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
Found a runtime blocker: pkg/model/process.go sets process.WithName("powershell.exe") (a bare filename) together with process.WithWorkDir(workDir). On Windows, Go's syscall.StartProcess resolves the application name relative to attr.Dir when both are set — so it looks for <workDir>\powershell.exe instead of searching PATH, and CreateProcess with an explicit lpApplicationName does no PATH lookup. Every Windows backend launch will fail with ERROR_FILE_NOT_FOUND.
Fix: resolve the full path before calling WithName, e.g. exec.LookPath("powershell.exe") (which searches PATH and finds System32), or hardcode C:\\Windows\\System32\\WindowsPowerShell\\v1.0\\powershell.exe.
The rest of the PR looks sound — capability detection ordering, docs, CI workflows. Once this is fixed, it should be good to merge.
4c2f267 to
bd37efd
Compare
bd37efd to
78bcf2d
Compare
|
I would just resolve the merge conflict and enable auot-merge, but there appear to be a few issues
Also it would make sense to add some Windows tests for basic things like starting LocalAI and doing some requests on a mock backend using the existing infrastructure for doing that as far as possible. |
ce305b8 to
0d12a6e
Compare
0d12a6e to
6602a1c
Compare
|
@richiejp thanks for the thorough review. All four points are addressed on the rebased head, and the conflict note is confirmed: Merge conflict ( P1 - Backend survives shutdown - implemented with a Windows Job Object. In P2 - Misleading firewall advice - fixed. The docs now state the real P3 - Windows images aren't signed - fixed: Windows tests - a job-object unit test (
Wired as |
b4fd355 to
6d92e74
Compare
56f8395 to
27f5ea2
Compare
|
Just small code nits on my side, but otherwise the direction looks good. Unfortunately, I do not have windows to test this, so we will have to see via CI and user reports how this behaves in real hardware. |
9a6c223 to
80deff1
Compare
mudler
left a comment
There was a problem hiding this comment.
we are almost there :) thanks for the patience @LionelColaso. Last review pass and then should be good to go
8320fb6 to
7bc3106
Compare
…ckend) Native Windows support end to end: LocalAI releases a windows/amd64 server binary and the llama-cpp backend is built natively for windows/amd64 under MSYS2 UCRT64 and packaged as an OCI image tar that LocalAI installs and runs as a native process - no docker daemon or WSL required on the host. Server side: - goreleaser: add windows (amd64/arm64) to the release targets - Makefile: download the win64 protoc zip and rename protoc.exe to protoc, resolve code-gen plugins via --plugin instead of PATH, force SHELL=sh and name the binary local-ai.exe on Windows, ignore protoc.exe - build-test.yaml: add a native windows-latest build gate that installs GNU make via choco, adds Git for Windows' usr/bin to PATH and builds with CGO_ENABLED=0 - pkg/downloader: close the write handle before removing or renaming the partial so Windows file locks do not break resume and error paths; guard the POSIX-permission and symlink tests on non-Windows - tests: Windows guards and path fixes for core/gallery, video_internal, loader and the testcontainers database setup Backend side: - scripts/build/llama-cpp-windows.sh: builds gRPC from source (pinned v1.59.0, with mingw-w64 fixes for c-ares, boringssl and zlib), then the three llama.cpp variants (cpu-all with GGML_CPU_ALL_VARIANTS + Vulkan, rpc, and the AVX-off fallback), bundles the mingw runtime DLLs and ships an OCI tar via local-ai util create-oci-image. Re-runnable and auto-dispatchs into MSYS2 when launched from Git for Windows' bash. make backends/llama-cpp-windows hands the build to scripts/build/llama-cpp-windows.ps1 on Windows: it locates MSYS2 (or, after asking for confirmation, installs it via winget and pacman-installs the mingw-w64-ucrt toolchain) and runs the sh script under MSYS2 UCRT64 bash. - backend/cpp/llama-cpp/run.ps1: PowerShell launcher (mirrors run.sh) that pkg/model starts on Windows. The launcher runs through the Windows PowerShell resolved via exec.LookPath, because os.StartProcess resolves argv0 against the backend workDir on a Windows host. - backend/index.yaml: windows/amd64 backend entry and variants. - pkg/system/capabilities.go: windows engine preference rules so the gallery picks the native build on Windows hosts. - .github/backend-matrix.yml + backend_build_windows.yml: windows matrix entries and a reusable windows build workflow; backend.yml and backend_pr.yml wire the windows backend jobs (build on PR, publish on master). - core/gallery: backends install skips the OCI-registry digest lookup for ocifile:// streams (a local tarball is not a registry reference), so the local-build install step no longer prints a confusing "could not parse reference" digest warning. - docs: getting-started/windows.md plus related page updates. JOBS in the build script honors an override so memory-limited hosts can build with reduced parallelism. Process-tree cleanup (Windows job objects): - pkg/model/process_tree.go declares a processTree interface; the Windows implementation (process_tree_windows.go, kill-on-close job object) and the no-op for other platforms (process_tree_other.go) are selected at build time, keeping Windows-specific code out of the shared process runtime. - The teardown is wired into the stop paths so the backend tree cannot outlive a deliberate unload. Launcher resolution: - pkg/model/process_launcher_windows.go / process_launcher_other.go resolve the executable and argv per platform: other platforms spawn the gallery-contract run.sh stub as-is, Windows substitutes the bundled run.ps1 through the system PowerShell. The shared startProcess calls resolveLauncher, so no runtime.GOOS branch remains in generic code. Windows smoke test: - tests/e2e/windows: standalone ginkgo suite (skips itself on non-Windows) that builds and boots the real local-ai.exe (LOCAL_AI_EXE override for pre-built binaries) with the mock-backend laid out as the gallery ships it - run.sh for discovery plus run.ps1 to launch. Two specs: a chat completion through the backend launched via run.ps1, and a hard-kill of local-ai.exe with an assertion that the wrapper + backend tree is reaped via the job object (kill-on-close), identifying the tree by parentage from the server PID rather than command-line text so unrelated processes on a developer machine cannot false-positive. - Makefile: test-windows-smoke target (protogen-go + react-ui + ginkgo on the suite). - tests-e2e.yml: windows-latest smoke job mirroring the ubuntu e2e job's proto setup (win64 protoc + plugins), then builds both binaries with CGO_ENABLED=0 and runs the suite. Assisted-by: opencode:big-pickle Signed-off-by: Lionel Colaso <lionelcolaso@outlook.com>
7bc3106 to
12cbf4e
Compare
mudler-agent
left a comment
There was a problem hiding this comment.
Two findings from the Windows support review. Rechecked against the current head; the affected code is unchanged.
| git checkout -q -B build "$LLAMA_VERSION" | ||
| git submodule update --init --recursive --depth 1 --single-branch | ||
| cd "$ROOT/backend/cpp/llama-cpp" | ||
| bash prepare.sh |
There was a problem hiding this comment.
[P1] Create the gRPC staging directory before calling prepare.sh
prepare.sh copies upstream server files into llama.cpp/tools/grpc-server/ under set -e, but this script never creates that directory. The preceding git clean -qfd also removes staging left by an earlier run. A fresh backend build therefore exits during source preparation, before compiling the llama.cpp variants.
I reproduced the failure with the unchanged helper and a minimal upstream server fixture: cp: cannot create regular file 'llama.cpp/tools/grpc-server/': Not a directory. The existing backend Makefile creates the directory before invoking the helper. Please add mkdir -p llama.cpp/tools/grpc-server here before bash prepare.sh.
| // tracker so that stopping it reaps the whole tree. The mechanism is | ||
| // platform specific; see the processTree interface. | ||
| if pid, err := strconv.Atoi(grpcControlProcess.CurrentPID()); err == nil { | ||
| runtime.trackProcessTree(pid) |
There was a problem hiding this comment.
[P2] Assign the Windows job before the launcher can spawn children
grpcControlProcess.Run() starts PowerShell before this job assignment. If LocalAI is descheduled after starting the wrapper, PowerShell can launch the backend first. Assigning an already-running parent to a job does not retroactively enroll its existing children, so a later unload or abrupt LocalAI exit can leave the backend alive with its memory/GPU allocation and listening socket.
windows_job_test.go explicitly gates child creation until after assignment, but the production launcher has no equivalent synchronization. The llama.cpp parent-death watcher is a no-op on Windows, so it does not cover this case.
Please enforce the ordering, for example by creating the launcher suspended, assigning it to the job, then resuming it, or by adding a startup handshake before child creation. A regression test should exercise that production ordering. This finding is based on source analysis; I have not reproduced the race on a Windows host.
Description
Adds native Windows support end to end: LocalAI now releases a
windows/amd64server binary, and thellama-cppbackend is built natively forwindows/amd64under MSYS2 UCRT64 and packaged as an OCI image tar that LocalAI installs and runs as a native process - no docker daemon or WSL required on the host.Server side
windows(amd64/arm64) to the release targetsprotoczip and renameprotoc.exetoprotoc, resolve code-gen plugins via--plugininstead ofPATH, forceSHELL=shand name the binarylocal-ai.exeon Windows, ignoreprotoc.exewindows-latestbuild gate that installs GNU make via Chocolatey, adds Git for Windows'usr/bintoPATHand builds withCGO_ENABLED=0core/gallery,video_internal,loaderand the testcontainers database setupBackend side
GGML_CPU_ALL_VARIANTS+ Vulkan, rpc, and the AVX-off fallback), bundles the mingw runtime DLLs and ships an OCI tar vialocal-ai util create-oci-image. Re-runnable and auto-dispatches into MSYS2 when launched from Git for Windows' bash.JOBSoverride supported for memory-limited hosts.backends installskips the OCI-registry digest lookup forocifile://streams (a local tarball is not a registry reference), so the local-build install step no longer prints a confusingcould not parse referencedigest warningrun.sh) thatpkg/model/process.gostarts on Windowswindows/amd64backend entry and variantsbackend.ymlandbackend_pr.ymlwire the windows backend jobs (build on PR, publish on master)Notes for Reviewers
run.ps1launcher:llama-cpp-cpu-all.exe(all ggml CPU variants, auto-detects a Vulkan device at runtime and falls back to CPU),llama-cpp-grpc.exe(gRPC-RPC build, selected whenLLAMACPP_GRPC_SERVERSis set) andllama-cpp-fallback.exe(static, AVX-off fallback). The mingw runtime DLLs are bundled in the image, so no MSYS2 install is needed on the host.JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE(pkg/model/windows_job_windows.go), so model unloads, graceful shutdown, and an abrupt local-ai.exe exit all reap the wrapper + backend tree; it degrades to a logged warning when the host already nests the process in a non-breakaway job.tests/e2e/windows): builds and boots the real local-ai.exe with the mock-backend laid out exactly as a gallery install ships it (run.shfor discovery,run.ps1to launch) and asserts a chat completion via run.ps1 plus job-object tree reaping after hard-killing local-ai.exe. Self-skips on non-Windows hosts; wired asmake test-windows-smokeand awindows-latestjob intests-e2e.yml. 2/2 specs pass locally on a Windows host and on thetests-windows-smokeCI job.backend_build_windows.yml's publish job now signs the windows images keyless with cosign (same flow asbackend_merge.yml); the galleryverification:block stays unpopulated until it can cover every published variant.Assisted-by:trailer per .agents/ai-coding-assistants.md.Signed commits
Closes #2368