Repository navigation
Redline: a watchdog thread is started for every call into native code #201
Description
Activity
If I'm understanding correctly this is to allow native code to also detect
Thread.interrupt(), correct?Rather than polling periodically what about compiling the output code to check interrupt on other non-temporal boundaries, like on function entry or return, on branch backedges, etc? That's how JRuby currently compiles Ruby's thread interrupt-checking, with all interrupt checks guarded by SwitchPoint so they're free until invalidated. Interrupt-checking from JRuby's Java code is done explicitly at those boundaries.
Down side of SwitchPoint is that JVM will be forced to invalidate the compiled code on invalidation. I believe @forax has some other ideas about low-cost interrupt checking, like most using Panama to ping a memory page and using a page fault to signal the interrupt.
Being invoked, here is the different ways to stop a thread.
@forax Your slides don't include a benchmark for the MutableCallSite version (I expect SwitchPoint would be similar) I'm assuming it's basically free but did you ever compare it with the Arena version?
Nevermind, I figured out how to run your JMH benchmarks. Here's the result I got with everything uncommented.
Benchmark Mode Cnt Score Error Units ThreadStopLoopArrayAccessBench.no_stop avgt 5 11.637 ± 0.106 us/op ThreadStopLoopArrayAccessBench.stop_arena avgt 5 11.878 ± 0.716 us/op ThreadStopLoopArrayAccessBench.stop_interrupt avgt 5 51.693 ± 7.732 us/op ThreadStopLoopArrayAccessBench.stop_opaque avgt 5 12.080 ± 0.118 us/op ThreadStopLoopArrayAccessBench.stop_reentrant_lock avgt 5 469.446 ± 1.987 us/op ThreadStopLoopArrayAccessBench.stop_synchronized avgt 5 383.378 ± 1.480 us/op ThreadStopLoopArrayAccessBench.stop_volatile avgt 5 55.444 ± 31.285 us/op ThreadStopLoopBench.no_stop avgt 5 26.443 ± 0.997 us/op ThreadStopLoopBench.stop_arena avgt 5 25.967 ± 0.452 us/op ThreadStopLoopBench.stop_callsite avgt 5 27.451 ± 8.746 us/op ThreadStopLoopBench.stop_interrupt avgt 5 43.547 ± 6.186 us/op ThreadStopLoopBench.stop_opaque avgt 5 28.102 ± 5.464 us/op ThreadStopLoopBench.stop_reentrant_lock avgt 5 477.997 ± 30.385 us/op ThreadStopLoopBench.stop_synchronized avgt 5 392.615 ± 17.043 us/op ThreadStopLoopBench.stop_volatile avgt 5 44.064 ± 9.784 us/opTake this with a grain of salt as it was run on my MacBook Air M4.
Was the callsite benchmark commented out because it interferes with other benchmarks (due to JIT deopt effects?)
Thanks @headius and @forax for the pointers!
Before settling on the external thread I have attempted various approaches, and tested those on relevant real world Wasm payloads (SQLite, QuickJs, CPython etc.).
What I ended with is:- instrumenting the code, often, means adding instructions on a hot hot path
- we don't have a strong requisite for being "precise" when stopping(so far)
The benchmark that I run convinced me the external Thread was the cheapest and more portable solution.
External thread still means checking the thread interrupt status or some volatile memory, correct? I think the examples provided by @forax show that other options can be faster.
When you can rely on Panama being available, the arena strategy seems like the clear winner. Otherwise, the mutable call site safepoint is just as cheap while having impacts on jit compiled code. Then it's a trade-off between how likely interrupts are and how long it takes to recover from jit invalidation.
External thread still means checking the thread interrupt status or some volatile memory, correct?
That's correct, at the moment we are using a
MemorySegmenton Panama, the check is happening at very long intervals.At the moment, what I can commit to, is the fix of the issue ( #210 ); when someone wants to experiment and provide feedback I'm really curious to see the results and discuss the right way forward!
JffiNativeMachine.call()and the PanamaNativeMachine.call()create, start and then interrupt a new watchdogThreadon every call into native code (it pollscaller.isInterrupted()every 1 ms to raise the interrupt flag):A host function that calls back into an export (e.g.
cabi_reallocto return a value) pays it again for the nested call. Measured on quickjs4j (Endive 1.1.0, javy module): 16 threads started per JS evaluation, 58 per Camel message, versus 0 on bytecode. It makes an evaluation ~2.4x slower than the bytecode build-time compiler single-threaded, and ~6x slower with 8 threads, where thread creation serializes.Removing the watchdog from the 1.1.0 jffi runner brings the native path back in line (numbers in the quickjs4j investigation).
Suggested fix: one shared daemon watchdog (or none at all, checking the caller's interrupt flag from the host-call / memory-grow safepoints), never a thread per call.
🤖 Generated with Claude Code