Repository navigation
Make SIMD work by default in the interpreter (#181) - #205
Conversation
ea2591e to
a3e02a0
Compare
|
The spec-vector execution failed because the differential test found 113 mismatches, all in f32x4/f64x2 extract_lane and replace_lane: the existing Vector code went through a Java float/double for the lane value, and on JDK 21 that canonicalises NaN payloads (0xffffffff came back as 0x7fc00000) This bug was pre-existing and was actually exposed via the Differential Test - fixed as well. |
andreaTP
left a comment
There was a problem hiding this comment.
Hi @boev and thanks a lot for taking a stab at this big lift, much appreciated!
All in all I see this going definitely in the right direction, given the measurement you have done seems like it's not even worth keeping the implementation based on the Vector API, which is not extremely surprising, having to convert every time to/from 2 longs is probably eating up most of the time.
Before taking a direction I'd like to have a look at numbers from real-world modules with the compiler.
No need to have a clean implementation/PR for now, but hacking it up so that we can take an informed decision based on numbers.
Would you be up in doing the experiment?
|
Reporting progress as I go deeper on this. I did a local spike on top of this branch: compiler support for all v128 opcodes with two interchangeable backends — a scalar one (v128 as two Runtime in ms/op. First column is the library compiled without
Interpreter: where clang actually vectorised, the Compiler: everything with < 1 % SIMD is unchanged; the SIMD-heavy kernel runs 5–17× slower from the So for this PR: let's drop the Vector API implementation and keep the scalar one. Same speed in the interpreter, needed anyway for Java 11/17 and native-image, and ~3 000 lines less plus no multi-release jar / One takeaway worth stating explicitly: with the compiler, the plain i32 wasm code is handed to C2 as ordinary scalar Java bytecode, which it compiles straight to native and can auto-vectorise itself; explicit v128 instructions take that opportunity away and give it something it handles badly — boxed vectors with the Vector API, or emulated lanes with the scalar code. Today the fastest way to run SIMD-heavy code on endive's compiler is to not compile it with Next I'll look at the compiler side separately: whether C2 can keep a hand-written Vector API ChaCha20 kernel allocation-free at all, and how close a properly emitted scalar backend gets to the plain binary. More numbers as they come. |
|
Update after digging into the compiler side. The compiler numbers above were not a JVM limitation but a shape problem in my spike: the Vector API's own methods are Hand-written ChaCha20 over 64 KiB in Java, ms/op (JDK 25 / JDK 21):
For reference the compiler runs the plain (no Where that leaves us:
and, I need to sleep on this data 😅 |
|
Let's go scalar-only in this PR -> I think this is good choice. I will prepare it and ping you - great testing! |
|
Deal! Thanks 👍 |
8fbf514 to
966a5b4
Compare
|
@andreaTP. hi. I think this is now done as you requested - vector API gone. Markdown simply says: no module, no maven dependency. Works on all Javas. (did not test 27 yet, but has to work there too.) |
|
Thanks @boev I'll have a closer look sometime soon! |
966a5b4 to
f285261
Compare
The runtime gets a scalar v128 implementation (V128Ops), so every v128 instruction runs on every supported JDK without the incubator Vector API. All simd_*.wast spec tests run in the default build. The simd module is removed; the docs carry a migration note. Refs bytecodealliance#181
f285261 to
f557c9b
Compare
andreaTP
left a comment
There was a problem hiding this comment.
This LGTM @boev thanks a lot for the help!
As a follow up we need to craft a clear plan to get SIMD in the compiler too, so far, the only way to get them to perform was to directly emit the needed code without access to any helper(no static methods invocations etc.).
Summary
This MR is directly related to Issue 181 and implemented SIMD instructions into the runtime, making them available in every Java - no flags, no dependency, no module added.
Before
With this MR:
artifact, same on every JDK; behaviour changes at run time, not at build time.
Open, deliberately:
JDK 25.0.4.1, AMD Ryzen 7 7700X, 2026-09-20. Milliseconds per call, median of 1,000 calls (5 JVMs x 200); spread is min..max of the five per-JVM medians. Vector/scalar below 1.00 means the Vector path is faster.

Refs #181