Skip to content

Make SIMD work by default in the interpreter (#181) - #205

Merged
andreaTP merged 1 commit into
bytecodealliance:mainfrom
boev:simd-scalar-differential
Oct 6, 2026
Merged

andreaTP merged 1 commit into
bytecodealliance:mainfrom
boev:simd-scalar-differential

Conversation

@boev

@boev boev commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This MR is directly related to Issue 181 and implemented SIMD instructions into the runtime, making them available in every Java - no flags, no dependency, no module added.

Before

  • SIMD was an optional add-on: a separate module that only existed when building on Java 21+, needed the incubator Vector API and a JVM flag, and users had to opt in by hand. On Java 11 or 17, or without the flag, SIMD instructions simply failed.
  • The Vector API changes between JDK releases, so the build generated different source variants per JDK to keep up.
  • The test suite only ran SIMD tests inside that special build, so SIMD was never tested on Java 11 or 17, and the whole thing was held together by a templating plugin.
  • Six correctness bugs sat in the Vector code unnoticed, because nothing compared it to a reference.

With this MR:

  • SIMD just works: built into the core runtime, on by default, on every supported Java, no flags, no extra dependency. A plain-Java implementation of all 236 instructions is the baseline.
  • The Vector API version is kept as an optional fast path (it is questionable if it really is faster, see one of open points below) inside the same jar, activated only when the JVM actually has the module. One
    artifact, same on every JDK; behaviour changes at run time, not at build time.
  • The JDK naming differences are resolved once at startup, so no more per-JDK source generation.
  • The spec test suite runs SIMD on every CI cell, and on Java 21+ also runs a second time through the Vector path. The templating is gone.
  • A differential test drives every instruction through both implementations on a lot of inputs. Tests runs tens of thousands every build and I ran a one-off with many millions. It found the six bugs; they are fixed - all in the existing Vector code: f64x2/i64x2 comparison masks, i64x2 signed compare operand order, f64x2 add and div, f64x2.promote_low_f32x4, f64x2.convert_low_i32x4_s/u, i64x2.extmul_high_i32x4_u.
  • The simd module is removed. Its Vector code moved into the runtime (git shows it as a rename), and the docs name the two lines a user of run.endive:simd has to delete. -> compatibility for people locking run.endive:simd can be added, but I deliberately decided against it as it introduces complexity and adds no obvious advantage.

Open, deliberately:

  • Compiler support is a later step.
  • The Vector path is not faster in the interpreter on my measurements; Issue 181 asked for, so it stays, but we might wanna consider removing it for simplicity. It would really simplify build, pom configs, tests. Here is a deeper comparison of times on my machine:

JDK 25.0.4.1, AMD Ryzen 7 7700X, 2026-09-20. Milliseconds per call, median of 1,000 calls (5 JVMs x 200); spread is min..max of the five per-JVM medians. Vector/scalar below 1.00 means the Vector path is faster.
image

Refs #181

@boev
boev requested a review from andreaTP as a code owner September 20, 2026 12:10
@boev
boev force-pushed the simd-scalar-differential branch from ea2591e to a3e02a0 Compare September 20, 2026 15:53
@boev

boev commented Sep 20, 2026

Copy link
Copy Markdown
Contributor Author

The spec-vector execution failed because the differential test found 113 mismatches, all in f32x4/f64x2 extract_lane and replace_lane: the existing Vector code went through a Java float/double for the lane value, and on JDK 21 that canonicalises NaN payloads (0xffffffff came back as 0x7fc00000)

This bug was pre-existing and was actually exposed via the Differential Test - fixed as well.

@andreaTP andreaTP left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @boev and thanks a lot for taking a stab at this big lift, much appreciated!

All in all I see this going definitely in the right direction, given the measurement you have done seems like it's not even worth keeping the implementation based on the Vector API, which is not extremely surprising, having to convert every time to/from 2 longs is probably eating up most of the time.

Before taking a direction I'd like to have a look at numbers from real-world modules with the compiler.
No need to have a clean implementation/PR for now, but hacking it up so that we can take an informed decision based on numbers.
Would you be up in doing the experiment?

Comment thread runtime-tests/src/test/java/run/endive/testing/V128DifferentialTest.java Outdated
Comment thread runtime/src/main/java21/run/endive/runtime/internal/V128Dispatch.java Outdated
Comment thread runtime/src/main/java21/run/endive/runtime/internal/V128Dispatch.java Outdated
@andreaTP

Copy link
Copy Markdown
Contributor

Reporting progress as I go deeper on this. I did a local spike on top of this branch: compiler support for all v128 opcodes with two interchangeable backends — a scalar one (v128 as two longs, reusing the V128Ops lane code) and a Vector API one (v128 as a LongVector) — then OpenSSL and libsodium compiled to wasm32-wasi twice, with and without -msimd128, run through the interpreter and the compiler (JDK 25, 64 KiB inputs, JMH). Both backends pass the SIMD spec suite through the compiler and all configurations produce identical output.

Runtime in ms/op. First column is the library compiled without -msimd128 (no v128 instructions in the binary, what people run today); the other two are the -msimd128 binary executed with each v128 implementation.

runtime, ms/op wasm without -msimd128 wasm with -msimd128, scalar v128 impl wasm with -msimd128, Vector API v128 impl
interpreter
sodium chacha20 30.4 20.3 20.9
sodium chacha20poly1305 38.7 29.0 27.2
openssl chacha20poly1305 60.4 41.1 33.5
other ops (SHA-2, AES, BLAKE2b, X25519, Poly1305) within ±10 % across all three
compiler
sodium chacha20 0.22 3.75 2.39
sodium chacha20poly1305 0.29 3.99 2.44
openssl chacha20poly1305 0.54 2.97 2.61
other ops within noise across all three

Interpreter: where clang actually vectorised, the -msimd128 binary runs 25–40 % faster, simply because there are fewer instructions to dispatch; scalar vs Vector API implementation is a wash, ±10 % both ways — your conclusion holds on real modules.

Compiler: everything with < 1 % SIMD is unchanged; the SIMD-heavy kernel runs 5–17× slower from the -msimd128 binary than from the plain one, with either implementation. The Vector API arithmetic is native, but C2 boxes every LongVector held in a JVM local or crossing a loop back-edge (~48 bytes per v128 op); the scalar representation avoids the allocation but pays ~10× instruction inflation.

So for this PR: let's drop the Vector API implementation and keep the scalar one. Same speed in the interpreter, needed anyway for Java 11/17 and native-image, and ~3 000 lines less plus no multi-release jar / --add-modules. Please keep the differential-test approach even with a single implementation, e.g. against a generated .wat corpus checked against a reference.

One takeaway worth stating explicitly: with the compiler, the plain i32 wasm code is handed to C2 as ordinary scalar Java bytecode, which it compiles straight to native and can auto-vectorise itself; explicit v128 instructions take that opportunity away and give it something it handles badly — boxed vectors with the Vector API, or emulated lanes with the scalar code. Today the fastest way to run SIMD-heavy code on endive's compiler is to not compile it with -msimd128.

Next I'll look at the compiler side separately: whether C2 can keep a hand-written Vector API ChaCha20 kernel allocation-free at all, and how close a properly emitted scalar backend gets to the plain binary. More numbers as they come.

@andreaTP

Copy link
Copy Markdown
Contributor

Update after digging into the compiler side.

The compiler numbers above were not a JVM limitation but a shape problem in my spike: the Vector API's own methods are @ForceInline and bypass C2's inlining budget, while any wrapper around them is an ordinary call subject to NodeCountInliningCutoff. A kernel of a few dozen v128 ops is already over that budget, so wrapper calls at the tail of the method stay real calls, their arguments get boxed and the wrapper runs the Java fallback path (the int[] allocations in the profile). Same result with same-class private statics; no flag fixes it.

Hand-written ChaCha20 over 64 KiB in Java, ms/op (JDK 25 / JDK 21):

JDK 25 JDK 21 alloc
plain scalar Java 0.168 0.158 0
idiomatic Vector API 0.061 0.097 0
same ops through helper methods (what the spike emitted) 1.48 1.89 11 MB/op
Vector API calls emitted directly, lane-typed locals, flat byte[] memory, constant shuffles 0.088 0.108 0
same with 4 blocks in flight (16 vector locals) 0.052 0.049 0

For reference the compiler runs the plain (no -msimd128) libsodium ChaCha20 in 0.22 ms. So the Vector API is 2–4× faster than what we produce today, but only when the generated bytecode has the shape javac produces for idiomatic code: the calls inline in the method body, species/shuffles/constants as static final fields, nothing routed through helper classes. JDK 25 accepts a completely naive emission (canonical LongVector, reinterpret on every op); JDK 21 needs lane-typed values and single-source shuffles to stay allocation-free.

Where that leaves us:

  • Interpreter: the scalar implementation is the optimal solution. Same speed as the Vector API on real modules, no incubator module, no multi-release jar, works on every JDK. Let's go scalar-only in this PR.
  • Compiler: it needs bespoke bytecode generation for v128 — emitting the Vector API sequences inline per opcode, with the two-long scalar code as the fallback for the ops the API cannot express and for JDKs without the module. That is a separate piece of work; the spike branch has the emitter plumbing and the experiments, and I'll open an issue for it.

and, I need to sleep on this data 😅

@boev

boev commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Let's go scalar-only in this PR -> I think this is good choice. I will prepare it and ping you - great testing!

@andreaTP

Copy link
Copy Markdown
Contributor

Deal! Thanks 👍

@boev
boev force-pushed the simd-scalar-differential branch 2 times, most recently from 8fbf514 to 966a5b4 Compare September 24, 2026 17:39
@boev

boev commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author

@andreaTP. hi. I think this is now done as you requested - vector API gone.

Markdown simply says:
SIMD support is built into Endive.

no module, no maven dependency. Works on all Javas. (did not test 27 yet, but has to work there too.)

@andreaTP

Copy link
Copy Markdown
Contributor

Thanks @boev I'll have a closer look sometime soon!

@andreaTP andreaTP left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot @boev ! This is a big PR and it needs a bit of care, but I really like where we are going, thanks helping out!

Comment thread docs/docs/advanced/simd.md Outdated
Comment thread simd/src/test/java/run/endive/simd/BasicSimdTest.java
Comment thread runtime-tests/src/test/resources/wast/v128_corpus.wast Outdated
Comment thread runtime/src/main/java/run/endive/runtime/internal/V128Ops.java
Comment thread runtime/src/main/java/run/endive/runtime/InterpreterMachine.java Outdated
Comment thread runtime/pom.xml Outdated
Comment thread test-gen-lib/src/main/java/run/endive/testgen/TestGen.java Outdated
Comment thread test-gen-lib/src/main/java/run/endive/testgen/V128CorpusGenerator.java Outdated
@boev
boev force-pushed the simd-scalar-differential branch from 966a5b4 to f285261 Compare September 30, 2026 22:27
The runtime gets a scalar v128 implementation (V128Ops), so every v128
instruction runs on every supported JDK without the incubator Vector
API. All simd_*.wast spec tests run in the default build. The simd
module is removed; the docs carry a migration note.

Refs bytecodealliance#181
@boev
boev force-pushed the simd-scalar-differential branch from f285261 to f557c9b Compare October 5, 2026 18:11

@andreaTP andreaTP left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This LGTM @boev thanks a lot for the help!

As a follow up we need to craft a clear plan to get SIMD in the compiler too, so far, the only way to get them to perform was to directly emit the needed code without access to any helper(no static methods invocations etc.).

@andreaTP
andreaTP merged commit 6457200 into bytecodealliance:main Oct 6, 2026
25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants