🎩 You're Invited:Meet the Socket team at Black Hat in Las Vegas, August 3-6.RSVP
Sign In

@ondeinference/react-native

Package Overview
Dependencies
Maintainers
3
Versions
14
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@ondeinference/react-native

On-device LLM inference for React Native. Run Qwen 2.5 models locally with Metal on iOS, CPU on Android. No cloud, no API key.

Source
npmnpm
Version
1.0.0
Version published
Weekly downloads
10
-84.37%
Maintainers
3
Weekly downloads
 
Created
Source

Onde Inference

Onde Inference

On-device LLM inference for React Native — Metal on iOS, CPU on Android.

npm crates.io Swift Package Index pub.dev Website

Rust SDK · Swift SDK · Flutter SDK · Website

Run Qwen 2.5 models on the phone. No server, no API key, nothing leaves the device.

The model downloads from HuggingFace on first load, then runs locally. ~941 MB for the 1.5B variant. Metal gives you ~15 tok/s on an iPhone 15; Android runs on CPU, slower but works.

Installation

npx expo install @ondeinference/react-native

Quick start

import { OndeChatEngine, userMessage } from "@ondeinference/react-native";

// Picks the right model for the device:
//   iOS     → Qwen 2.5 1.5B (~941 MB, Metal)
//   Android → Qwen 2.5 1.5B (~941 MB, CPU)
const seconds = await OndeChatEngine.loadDefaultModel(
  "You are a helpful assistant."
);

const reply = await OndeChatEngine.sendMessage("Hello!");
console.log(reply.text);

// One-shot — doesn't touch conversation history
const expanded = await OndeChatEngine.generate(
  [userMessage("Expand: a cat in space")],
  { temperature: 0.0 }
);

await OndeChatEngine.unloadModel();

Platforms

PlatformBackendDefault model
iOSMetalQwen 2.5 1.5B (~941 MB)
AndroidCPUQwen 2.5 1.5B (~941 MB)

API

OndeChatEngine

MethodReturnsWhat it does
loadDefaultModel(systemPrompt?, sampling?)Promise<number>Load the platform default. Returns load time in seconds.
loadModel(config, systemPrompt?, sampling?)Promise<number>Load a specific GGUF model.
unloadModel()Promise<string | null>Drop the model, free memory. Returns the model name.
isLoaded()booleanIs anything loaded right now?
info()Promise<EngineInfo>Status, model name, memory, history length.
sendMessage(message)Promise<InferenceResult>Chat turn. Appends to history automatically.
generate(messages, sampling?)Promise<InferenceResult>One-shot. History stays untouched.
setSystemPrompt(prompt)voidReplace the system prompt.
clearSystemPrompt()voidRemove it.
setSampling(config)voidSwap sampling params.
history()Promise<ChatMessage[]>Full conversation so far.
clearHistory()numberWipe it. Returns how many messages were removed.
pushHistory(message)voidInject a message without running inference.

Helpers

import {
  defaultModelConfig,     // platform-aware (1.5B on mobile, 3B on desktop)
  qwen251_5bConfig,       // force 1.5B (~941 MB)
  qwen253bConfig,         // force 3B (~1.93 GB)
  defaultSamplingConfig,  // temp=0.7, top_p=0.95, max_tokens=512
  deterministicSamplingConfig,  // temp=0.0
  mobileSamplingConfig,   // temp=0.7, max_tokens=128
  systemMessage,
  userMessage,
  assistantMessage,
} from "@ondeinference/react-native";

Example app

There's a working chat app in example/:

cd example
npm install
npx expo run:ios

Single file, ~290 lines. Shows loading, chat, status, history management, and error handling.

Building from source

You need Rust and the right cross-compilation targets.

# iOS
rustup target add aarch64-apple-ios aarch64-apple-ios-sim
./scripts/build-rust.sh ios

# Android (set ANDROID_NDK_HOME first)
rustup target add aarch64-linux-android armv7-linux-androideabi x86_64-linux-android i686-linux-android
./scripts/build-rust.sh android

The script compiles the Rust FFI bridge in rust/, then copies the static lib (iOS) or shared libs (Android) into the right places under ios/ and android/.

How it works

TypeScript  →  Expo Module (Swift / Kotlin)  →  Rust C FFI  →  onde crate  →  mistral.rs
                @_silgen_name (iOS)                               ↓
                JNI external (Android)                     Metal / CPU

The native module talks to Rust through extern "C" functions. Complex types cross the boundary as JSON strings — the TypeScript layer handles camelCase ↔ snake_case conversion. A global tokio::Runtime (created once) runs the async inference.

License

Dual-licensed under MIT and Apache 2.0. Pick whichever works for you.

© 2026 Splitfire AB

© 2026 Onde Inference (Splitfire AB).

Keywords

react-native

FAQs

Package last updated on 26 Apr 2026

Did you know?

Socket

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

Related posts