A patch release that moves to the v0.10.4 native libraries and fixes LLM and install issues.
Bug fixes
- The XNNPACK backend no longer frees the constants of operators it does not pack, so models with PReLU layers stop returning NaN or wrong values. It also no longer keeps an unpacked copy of every weight with a bias, so Kokoro TTS on XNNPACK peaks at about 1.3 GB instead of being killed on iPhone at the 3.3 GB limit.
- Opting out of a backend now also removes it from what ships: the
core-*artifacts carry the runtime only, Android excludes opted-out backends from the APK, and turning a backend off after an install takes effect (#1466 by @msluszniak). - LLM prompts longer than the widest input the graph accepts are prefilled in chunks at that bound instead of failing with
NotSupported, which hit Gemma 4 on MLX and XNNPACK (#1491 by @msluszniak). - The LLM runner loads models with
mmap. A multimodal model is no longer copied onto the heap, where a 2.9 GB Gemma 4 MLX export was killed during load on iPhone, and text models are no longer locked into memory, which made the memory reported to JavaScript about 6x the legacy API's (#1492 by @msluszniak). - Chat responses no longer end with a stop token the model declares besides
eos_token, such as Qwen's<|endoftext|>(#1486 by @msluszniak). - The native library download works on Windows (#1494 by @msluszniak).
API changes
LLMRunner.generateno longer acceptsecho; the prompt is never echoed (#1491).
Full changelog: v0.10.2...v0.10.3