
Security News
Ruby's Bundler 4.0.18 Extends Cooldown to bundle lock and bundle cache
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.
@ondeinference/react-native
Advanced tools
On-device LLM inference for React Native. Run Qwen 2.5 models locally with Metal on iOS, CPU on Android. No cloud, no API key.
On-device LLM inference for React Native — Metal on iOS, CPU on Android.
Rust SDK · Swift SDK · Flutter SDK · Website
Run Qwen 2.5 models on the phone. No server, no API key, nothing leaves the device.
The model downloads from HuggingFace on first load, then runs locally. ~941 MB for the 1.5B variant. Metal gives you ~15 tok/s on an iPhone 15; Android runs on CPU, slower but works.
npx expo install @ondeinference/react-native
import { OndeChatEngine, userMessage } from "@ondeinference/react-native";
// Picks the right model for the device:
// iOS → Qwen 2.5 1.5B (~941 MB, Metal)
// Android → Qwen 2.5 1.5B (~941 MB, CPU)
const seconds = await OndeChatEngine.loadDefaultModel(
"You are a helpful assistant."
);
const reply = await OndeChatEngine.sendMessage("Hello!");
console.log(reply.text);
// One-shot — doesn't touch conversation history
const expanded = await OndeChatEngine.generate(
[userMessage("Expand: a cat in space")],
{ temperature: 0.0 }
);
await OndeChatEngine.unloadModel();
| Platform | Backend | Default model |
|---|---|---|
| iOS | Metal | Qwen 2.5 1.5B (~941 MB) |
| Android | CPU | Qwen 2.5 1.5B (~941 MB) |
| Method | Returns | What it does |
|---|---|---|
loadDefaultModel(systemPrompt?, sampling?) | Promise<number> | Load the platform default. Returns load time in seconds. |
loadModel(config, systemPrompt?, sampling?) | Promise<number> | Load a specific GGUF model. |
unloadModel() | Promise<string | null> | Drop the model, free memory. Returns the model name. |
isLoaded() | boolean | Is anything loaded right now? |
info() | Promise<EngineInfo> | Status, model name, memory, history length. |
sendMessage(message) | Promise<InferenceResult> | Chat turn. Appends to history automatically. |
generate(messages, sampling?) | Promise<InferenceResult> | One-shot. History stays untouched. |
setSystemPrompt(prompt) | void | Replace the system prompt. |
clearSystemPrompt() | void | Remove it. |
setSampling(config) | void | Swap sampling params. |
history() | Promise<ChatMessage[]> | Full conversation so far. |
clearHistory() | number | Wipe it. Returns how many messages were removed. |
pushHistory(message) | void | Inject a message without running inference. |
import {
defaultModelConfig, // platform-aware (1.5B on mobile, 3B on desktop)
qwen251_5bConfig, // force 1.5B (~941 MB)
qwen253bConfig, // force 3B (~1.93 GB)
defaultSamplingConfig, // temp=0.7, top_p=0.95, max_tokens=512
deterministicSamplingConfig, // temp=0.0
mobileSamplingConfig, // temp=0.7, max_tokens=128
systemMessage,
userMessage,
assistantMessage,
} from "@ondeinference/react-native";
There's a working chat app in example/:
cd example
npm install
npx expo run:ios
Single file, ~290 lines. Shows loading, chat, status, history management, and error handling.
You need Rust and the right cross-compilation targets.
# iOS
rustup target add aarch64-apple-ios aarch64-apple-ios-sim
./scripts/build-rust.sh ios
# Android (set ANDROID_NDK_HOME first)
rustup target add aarch64-linux-android armv7-linux-androideabi x86_64-linux-android i686-linux-android
./scripts/build-rust.sh android
The script compiles the Rust FFI bridge in rust/, then copies the static lib (iOS) or shared libs (Android) into the right places under ios/ and android/.
TypeScript → Expo Module (Swift / Kotlin) → Rust C FFI → onde crate → mistral.rs
@_silgen_name (iOS) ↓
JNI external (Android) Metal / CPU
The native module talks to Rust through extern "C" functions. Complex types cross the boundary as JSON strings — the TypeScript layer handles camelCase ↔ snake_case conversion. A global tokio::Runtime (created once) runs the async inference.
Dual-licensed under MIT and Apache 2.0. Pick whichever works for you.
© 2026 Splitfire AB
© 2026 Onde Inference (Splitfire AB).
FAQs
On-device LLM inference for React Native. Run Qwen 2.5 models locally with Metal on iOS, CPU on Android. No cloud, no API key.
The npm package @ondeinference/react-native receives a total of 10 weekly downloads. As such, @ondeinference/react-native popularity was classified as not popular.
We found that @ondeinference/react-native demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 3 open source maintainers collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Security News
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.

Security News
During a UK cyber test, a Mythos 5 agent used sockpuppets, social engineering, and prompt injection to try to get a maintainer to merge malware.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.