@jawcode-dev/ai
Unified LLM API with automatic model discovery, provider configuration, token and cost tracking, and simple context persistence and hand-off to other models mid-session.
Note: This library only includes models that support tool calling (function calling), as this is essential for agentic workflows.
Table of Contents
Supported Providers
- OpenAI
- OpenAI code provider (ChatGPT Plus/Pro subscription, requires OAuth, see below)
- Anthropic
- Google
- Vertex AI (Gemini via Vertex AI)
- Mistral
- Groq
- Cerebras
- Together
- Moonshot (requires
MOONSHOT_API_KEY)
- Qianfan (requires
QIANFAN_API_KEY)
- NVIDIA (requires
NVIDIA_API_KEY)
- NanoGPT (requires
NANO_GPT_API_KEY)
- Hugging Face Inference
- xAI
- Venice (requires
VENICE_API_KEY)
- OpenRouter
- Kilo Gateway (supports OAuth
/login kilo or KILO_API_KEY)
- LiteLLM (requires
LITELLM_API_KEY)
- zAI (requires
ZAI_API_KEY)
- MiniMax Coding Plan (requires
MINIMAX_CODE_API_KEY or MINIMAX_CODE_CN_API_KEY)
- Xiaomi MiMo (requires
XIAOMI_API_KEY)
- ZenMux (requires
ZENMUX_API_KEY)
- Qwen Portal (supports
QWEN_OAUTH_TOKEN or QWEN_PORTAL_API_KEY)
- Cloudflare AI Gateway (requires
CLOUDFLARE_AI_GATEWAY_API_KEY and provider-specific gateway base URL)
- Ollama (local OpenAI-compatible runtime; optional
OLLAMA_API_KEY)
- Ollama Cloud (hosted native Ollama API; requires
OLLAMA_CLOUD_API_KEY)
- llama.cpp (local OpenAI and Anthropic compatible inference server)
- vLLM (OpenAI-compatible server;
VLLM_API_KEY for secured deployments)
- GitHub Copilot (requires OAuth, see below)
- Google Gemini CLI (requires OAuth, see below)
- Antigravity (requires OAuth, see below)
- Any OpenAI-compatible API: LM Studio, custom proxies, etc.
Installation
npm install @jawcode-dev/ai
Quick Start
import { z, getModel, stream, complete, Context, Tool } from "@jawcode-dev/ai";
const model = getModel("openai", "gpt-4o-mini");
const tools: Tool[] = [
{
name: "get_time",
description: "Get the current time",
parameters: z.object({
timezone: z
.string()
.optional()
.describe("Optional timezone (e.g., America/New_York)"),
}),
},
];
const context: Context = {
systemPrompt: ["You are a helpful assistant."],
messages: [{ role: "user", content: "What time is it?" }],
tools,
};
const s = stream(model, context);
for await (const event of s) {
switch (event.type) {
case "start":
console.log(`Starting with ${event.partial.model}`);
break;
case "text_start":
console.log("\n[Text started]");
break;
case "text_delta":
process.stdout.write(event.delta);
break;
case "text_end":
console.log("\n[Text ended]");
break;
case "thinking_start":
console.log("[Model is thinking...]");
break;
case "thinking_delta":
process.stdout.write(event.delta);
break;
case "thinking_end":
console.log("[Thinking complete]");
break;
case "toolcall_start":
console.log(`\n[Tool call started: index ${event.contentIndex}]`);
break;
case "toolcall_delta":
const partialCall = event.partial.content[event.contentIndex];
if (partialCall.type === "toolCall") {
console.log(`[Streaming args for ${partialCall.name}]`);
}
break;
case "toolcall_end":
console.log(`\nTool called: ${event.toolCall.name}`);
console.log(`Arguments: ${JSON.stringify(event.toolCall.arguments)}`);
break;
case "done":
console.log(`\nFinished: ${event.reason}`);
break;
case "error":
console.error(`Error: ${event.error}`);
break;
}
}
const finalMessage = await s.result();
context.messages.push(finalMessage);
const toolCalls = finalMessage.content.filter((b) => b.type === "toolCall");
for (const call of toolCalls) {
const result =
call.name === "get_time"
? new Date().toLocaleString("en-US", {
timeZone: call.arguments.timezone || "UTC",
dateStyle: "full",
timeStyle: "long",
})
: "Unknown tool";
context.messages.push({
role: "toolResult",
toolCallId: call.id,
toolName: call.name,
content: [{ type: "text", text: result }],
isError: false,
timestamp: Date.now(),
});
}
if (toolCalls.length > 0) {
const continuation = await complete(model, context);
context.messages.push(continuation);
console.log("After tool execution:", continuation.content);
}
console.log(`Total tokens: ${finalMessage.usage.input} in, ${finalMessage.usage.output} out`);
console.log(`Cost: $${finalMessage.usage.cost.total.toFixed(4)}`);
const response = await complete(model, context);
for (const block of response.content) {
if (block.type === "text") {
console.log(block.text);
} else if (block.type === "toolCall") {
console.log(`Tool: ${block.name}(${JSON.stringify(block.arguments)})`);
}
}
Tools
Tools enable LLMs to interact with external systems. This library uses Zod schemas for type-safe tool definitions with automatic validation. Schemas are converted to JSON Schema for providers as needed.
Defining Tools
import { z, Tool } from "@jawcode-dev/ai";
const weatherTool: Tool = {
name: "get_weather",
description: "Get current weather for a location",
parameters: z.object({
location: z.string().describe("City name or coordinates"),
units: z.enum(["celsius", "fahrenheit"]).default("celsius"),
}),
};
const bookMeetingTool: Tool = {
name: "book_meeting",
description: "Schedule a meeting",
parameters: z.object({
title: z.string().min(1),
startTime: z.string().describe("ISO 8601 date-time"),
endTime: z.string().describe("ISO 8601 date-time"),
attendees: z.array(z.email()).min(1),
}),
};
Handling Tool Calls
Tool results use content blocks and can include both text and images:
import * as fs from "node:fs";
const context: Context = {
messages: [{ role: "user", content: "What is the weather in London?" }],
tools: [weatherTool],
};
const response = await complete(model, context);
for (const block of response.content) {
if (block.type === "toolCall") {
const result = await executeWeatherApi(block.arguments);
context.messages.push({
role: "toolResult",
toolCallId: block.id,
toolName: block.name,
content: [{ type: "text", text: JSON.stringify(result) }],
isError: false,
timestamp: Date.now(),
});
}
}
const imageBuffer = fs.readFileSync("chart.png");
context.messages.push({
role: "toolResult",
toolCallId: "tool_xyz",
toolName: "generate_chart",
content: [
{ type: "text", text: "Generated chart showing temperature trends" },
{ type: "image", data: imageBuffer.toBase64(), mimeType: "image/png" },
],
isError: false,
timestamp: Date.now(),
});
Streaming Tool Calls with Partial JSON
During streaming, tool call arguments are progressively parsed as they arrive. This enables real-time UI updates before the complete arguments are available:
const s = stream(model, context);
for await (const event of s) {
if (event.type === "toolcall_delta") {
const toolCall = event.partial.content[event.contentIndex];
if (toolCall.type === "toolCall" && toolCall.arguments) {
if (toolCall.name === "write_file" && toolCall.arguments.path) {
console.log(`Writing to: ${toolCall.arguments.path}`);
if (toolCall.arguments.content) {
console.log(`Content preview: ${toolCall.arguments.content.substring(0, 100)}...`);
}
}
}
}
if (event.type === "toolcall_end") {
const toolCall = event.toolCall;
console.log(`Tool completed: ${toolCall.name}`, toolCall.arguments);
}
}
Important notes about partial tool arguments:
- During
toolcall_delta events, arguments contains the best-effort parse of partial JSON
- Fields may be missing or incomplete - always check for existence before use
- String values may be truncated mid-word
- Arrays may be incomplete
- Nested objects may be partially populated
- At minimum,
arguments will be an empty object {}, never undefined
- The Google provider does not support function call streaming. Instead, you will receive a single
toolcall_delta event with the full arguments.
Validating Tool Arguments
When using agentLoop, tool arguments are automatically validated against your Zod parameter schemas before execution. If validation fails, the error is returned to the model as a tool result, allowing it to retry.
When implementing your own tool execution loop with stream() or complete(), use validateToolCall to validate arguments before passing them to your tools:
import { stream, validateToolCall, Tool } from "@jawcode-dev/ai";
const tools: Tool[] = [weatherTool, calculatorTool];
const s = stream(model, { messages, tools });
for await (const event of s) {
if (event.type === "toolcall_end") {
const toolCall = event.toolCall;
try {
const validatedArgs = validateToolCall(tools, toolCall);
const result = await executeMyTool(toolCall.name, validatedArgs);
} catch (error) {
context.messages.push({
role: "toolResult",
toolCallId: toolCall.id,
toolName: toolCall.name,
content: [{ type: "text", text: error.message }],
isError: true,
timestamp: Date.now(),
});
}
}
}
Complete Event Reference
All streaming events emitted during assistant message generation:
start | Stream begins | partial: Initial assistant message structure |
text_start | Text block starts | contentIndex: Position in content array |
text_delta | Text chunk received | delta: New text, contentIndex: Position |
text_end | Text block complete | content: Full text, contentIndex: Position |
thinking_start | Thinking block starts | contentIndex: Position in content array |
thinking_delta | Thinking chunk received | delta: New text, contentIndex: Position |
thinking_end | Thinking block complete | content: Full thinking, contentIndex: Position |
toolcall_start | Tool call begins | contentIndex: Position in content array |
toolcall_delta | Tool arguments streaming | delta: JSON chunk, partial.content[contentIndex].arguments: Partial parsed args |
toolcall_end | Tool call complete | toolCall: Complete validated tool call with id, name, arguments |
done | Stream complete | reason: Stop reason ("stop", "length", "toolUse"), message: Final assistant message |
error | Error occurred | reason: Error type ("error" or "aborted"), error: AssistantMessage with partial content |
Image Input
Models with vision capabilities can process images. You can check if a model supports images via the input property. If you pass images to a non-vision model, they are silently ignored.
import * as fs from "node:fs";
import { getModel, complete } from "@jawcode-dev/ai";
const model = getModel("openai", "gpt-4o-mini");
if (model.input.includes("image")) {
console.log("Model supports vision");
}
const imageBuffer = fs.readFileSync("image.png");
const base64Image = imageBuffer.toBase64();
const response = await complete(model, {
messages: [
{
role: "user",
content: [
{ type: "text", text: "What is in this image?" },
{ type: "image", data: base64Image, mimeType: "image/png" },
],
},
],
});
for (const block of response.content) {
if (block.type === "text") {
console.log(block.text);
}
}
Thinking/Reasoning
Many models support thinking/reasoning capabilities where they can show their internal thought process. You can check if a model supports reasoning via the reasoning property. If you pass reasoning options to a non-reasoning model, they are silently ignored.
Unified Interface (streamSimple/completeSimple)
import { getModel, streamSimple, completeSimple } from "@jawcode-dev/ai";
const model = getModel("anthropic", "anthropic-model-sonnet-4-20250514");
if (model.reasoning) {
console.log("Model supports reasoning/thinking");
}
const response = await completeSimple(
model,
{
messages: [{ role: "user", content: "Solve: 2x + 5 = 13" }],
},
{
reasoning: "medium",
}
);
for (const block of response.content) {
if (block.type === "thinking") {
console.log("Thinking:", block.thinking);
} else if (block.type === "text") {
console.log("Response:", block.text);
}
}
Provider-Specific Options (stream/complete)
For fine-grained control, use the provider-specific options:
import { getModel, complete } from "@jawcode-dev/ai";
const openaiModel = getModel("openai", "gpt-5-mini");
await complete(openaiModel, context, {
reasoningEffort: "medium",
reasoningSummary: "detailed",
});
const anthropicModel = getModel("anthropic", "anthropic-model-sonnet-4-20250514");
await complete(anthropicModel, context, {
thinkingEnabled: true,
thinkingBudgetTokens: 8192,
});
const googleModel = getModel("google", "gemini-2.5-flash");
await complete(googleModel, context, {
thinking: {
enabled: true,
budgetTokens: 8192,
},
});
Streaming Thinking Content
When streaming, thinking content is delivered through specific events:
const s = streamSimple(model, context, { reasoning: "high" });
for await (const event of s) {
switch (event.type) {
case "thinking_start":
console.log("[Model started thinking]");
break;
case "thinking_delta":
process.stdout.write(event.delta);
break;
case "thinking_end":
console.log("\n[Thinking complete]");
break;
}
}
Stop Reasons
Every AssistantMessage includes a stopReason field that indicates how the generation ended:
"stop" - Normal completion, the model finished its response
"length" - Output hit the maximum token limit
"toolUse" - Model is calling tools and expects tool results
"error" - An error occurred during generation
"aborted" - Request was cancelled via abort signal
Error Handling
When a request ends with an error (including aborts and tool call validation errors), the streaming API emits an error event:
for await (const event of stream) {
if (event.type === "error") {
console.error(`Error (${event.reason}):`, event.error.errorMessage);
console.log("Partial content:", event.error.content);
}
}
const message = await stream.result();
if (message.stopReason === "error" || message.stopReason === "aborted") {
console.error("Request failed:", message.errorMessage);
}
Aborting Requests
The abort signal allows you to cancel in-progress requests. Aborted requests have stopReason === 'aborted':
import { getModel, stream } from "@jawcode-dev/ai";
const model = getModel("openai", "gpt-4o-mini");
const signal = AbortSignal.timeout(2000);
const s = stream(
model,
{
messages: [{ role: "user", content: "Write a long story" }],
},
{
signal,
}
);
for await (const event of s) {
if (event.type === "text_delta") {
process.stdout.write(event.delta);
} else if (event.type === "error") {
console.log(`${event.reason === "aborted" ? "Aborted" : "Error"}:`, event.error.errorMessage);
}
}
const response = await s.result();
if (response.stopReason === "aborted") {
console.log("Request was aborted:", response.errorMessage);
console.log("Partial content received:", response.content);
console.log("Tokens used:", response.usage);
}
Continuing After Abort
Aborted messages can be added to the conversation context and continued in subsequent requests:
const context = {
messages: [{ role: "user", content: "Explain quantum computing in detail" }],
};
const controller1 = new AbortController();
setTimeout(() => controller1.abort(), 2000);
const partial = await complete(model, context, { signal: controller1.signal });
context.messages.push(partial);
context.messages.push({ role: "user", content: "Please continue" });
const continuation = await complete(model, context);
Common Stream Options
All providers accept the base StreamOptions (in addition to provider-specific options):
apiKey: Override the provider API key
headers: Extra request headers merged on top of model-defined headers
sessionId: Provider-specific session identifier (prompt caching/routing)
signal: Abort in-flight requests
onPayload: Callback invoked with the provider request payload just before sending
Example:
const response = await complete(model, context, {
apiKey: "sk-live",
headers: { "X-Debug-Trace": "true" },
onPayload: (payload) => {
console.log("request payload", payload);
},
});
APIs, Models, and Providers
The library implements 4 API interfaces, each with its own streaming function and options:
anthropic-messages: Anthropic's Messages API (streamAnthropic, AnthropicOptions)
google-generative-ai: Google's Generative AI API (streamGoogle, GoogleOptions)
openai-completions: OpenAI's Chat Completions API (streamOpenAICompletions, OpenAICompletionsOptions)
openai-responses: OpenAI's Responses API (streamOpenAIResponses, OpenAIResponsesOptions)
Providers and Models
A provider offers models through a specific API. For example:
- Anthropic models use the
anthropic-messages API
- Google models use the
google-generative-ai API
- OpenAI models use the
openai-responses API
- Mistral, xAI, Cerebras, Groq, etc. models use the
openai-completions API (OpenAI-compatible)
Querying Providers and Models
import { getProviders, getModels, getModel } from "@jawcode-dev/ai";
const providers = getProviders();
console.log(providers);
const anthropicModels = getModels("anthropic");
for (const model of anthropicModels) {
console.log(`${model.id}: ${model.name}`);
console.log(` API: ${model.api}`);
console.log(` Context: ${model.contextWindow} tokens`);
console.log(` Vision: ${model.input.includes("image")}`);
console.log(` Reasoning: ${model.reasoning}`);
}
const model = getModel("openai", "gpt-4o-mini");
console.log(`Using ${model.name} via ${model.api} API`);
Custom Models
You can create custom models for local inference servers or custom endpoints.
For local Ollama, OLLAMA_API_KEY is optional and mainly needed for authenticated/self-hosted gateways. ollama remains the local OpenAI-compatible runtime integration.
import { Model, stream } from "@jawcode-dev/ai";
const ollamaModel: Model<"openai-completions"> = {
id: "llama-3.1-8b",
name: "Llama 3.1 8B (Ollama)",
api: "openai-completions",
provider: "ollama",
baseUrl: "http://localhost:11434/v1",
reasoning: false,
input: ["text"],
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
contextWindow: 128000,
maxTokens: 32000,
};
const localResponse = await stream(ollamaModel, context, {
apiKey: process.env.OLLAMA_API_KEY,
});
const ollamaCloudModel: Model<"ollama-chat"> = {
id: "gpt-oss:120b",
name: "GPT OSS 120B (Ollama Cloud)",
api: "ollama-chat",
provider: "ollama-cloud",
baseUrl: "https://ollama.com",
reasoning: true,
input: ["text", "image"],
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
contextWindow: 262144,
maxTokens: 8192,
};
const cloudResponse = await stream(ollamaCloudModel, context, {
apiKey: process.env.OLLAMA_CLOUD_API_KEY,
});
const litellmModel: Model<"openai-completions"> = {
id: "gpt-4o",
name: "GPT-4o (via LiteLLM)",
api: "openai-completions",
provider: "litellm",
baseUrl: "http://localhost:4000/v1",
reasoning: false,
input: ["text", "image"],
cost: { input: 2.5, output: 10, cacheRead: 0, cacheWrite: 0 },
contextWindow: 128000,
maxTokens: 16384,
compat: {
supportsStore: false,
},
};
const proxyModel: Model<"anthropic-messages"> = {
id: "anthropic-model-sonnet-4",
name: "Anthropic model Sonnet 4 (Proxied)",
api: "anthropic-messages",
provider: "custom-proxy",
baseUrl: "https://proxy.example.com/v1",
reasoning: true,
input: ["text", "image"],
cost: { input: 3, output: 15, cacheRead: 0.3, cacheWrite: 3.75 },
contextWindow: 200000,
maxTokens: 8192,
headers: {
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36",
"X-Custom-Auth": "bearer-token-here",
},
};
OpenAI Compatibility Settings
The openai-completions API is implemented by many providers with minor differences. By default, the library auto-detects compatibility settings based on baseUrl for known providers (Cerebras, xAI, Mistral, Chutes, etc.). For custom proxies or unknown endpoints, you can override these settings via the compat field:
interface OpenAICompat {
supportsStore?: boolean;
supportsDeveloperRole?: boolean;
supportsReasoningEffort?: boolean;
maxTokensField?: "max_completion_tokens" | "max_tokens";
extraBody?: Record<string, unknown>;
}
If compat is not set, the library falls back to URL-based detection. If compat is partially set, unspecified fields use the detected defaults. This is useful for:
- LiteLLM proxies: May not support
store field
- Custom inference servers: May use non-standard field names
- Self-hosted endpoints: May have different feature support
Type Safety
Models are typed by their API, ensuring type-safe options:
const anthropic-model = getModel("anthropic", "anthropic-model-sonnet-4-20250514");
await stream(anthropic-model, context, {
thinkingEnabled: true,
thinkingBudgetTokens: 2048,
});
Cross-Provider Handoffs
The library supports seamless handoffs between different LLM providers within the same conversation. This allows you to switch models mid-conversation while preserving context, including thinking blocks, tool calls, and tool results.
How It Works
When messages from one provider are sent to a different provider, the library automatically transforms them for compatibility:
- User and tool result messages are passed through unchanged
- Assistant messages from the same provider/API are preserved as-is
- Assistant messages from different providers have their thinking blocks converted to text with
<thinking> tags
- Tool calls and regular text are preserved unchanged
Example: Multi-Provider Conversation
import { getModel, complete, Context } from "@jawcode-dev/ai";
const anthropic-model = getModel("anthropic", "anthropic-model-sonnet-4-20250514");
const context: Context = {
messages: [],
};
context.messages.push({ role: "user", content: "What is 25 * 18?" });
const anthropic-modelResponse = await complete(anthropic-model, context, {
thinkingEnabled: true,
});
context.messages.push(anthropic-modelResponse);
const gpt5 = getModel("openai", "gpt-5-mini");
context.messages.push({ role: "user", content: "Is that calculation correct?" });
const gptResponse = await complete(gpt5, context);
context.messages.push(gptResponse);
const gemini = getModel("google", "gemini-2.5-flash");
context.messages.push({ role: "user", content: "What was the original question?" });
const geminiResponse = await complete(gemini, context);
Provider Compatibility
All providers can handle messages from other providers, including:
- Text content
- Tool calls and tool results (including images in tool results)
- Thinking/reasoning blocks (transformed to tagged text for cross-provider compatibility)
- Aborted messages with partial content
This enables flexible workflows where you can:
- Start with a fast model for initial responses
- Switch to a more capable model for complex reasoning
- Use specialized models for specific tasks
- Maintain conversation continuity across provider outages
Context Serialization
The Context object can be easily serialized and deserialized using standard JSON methods, making it simple to persist conversations, implement chat history, or transfer contexts between services:
import { Context, getModel, complete } from "@jawcode-dev/ai";
const context: Context = {
systemPrompt: ["You are a helpful assistant."],
messages: [{ role: "user", content: "What is TypeScript?" }],
};
const model = getModel("openai", "gpt-4o-mini");
const response = await complete(model, context);
context.messages.push(response);
const serialized = JSON.stringify(context);
console.log("Serialized context size:", serialized.length, "bytes");
localStorage.setItem("conversation", serialized);
const restored: Context = JSON.parse(localStorage.getItem("conversation")!);
restored.messages.push({ role: "user", content: "Tell me more about its type system" });
const newModel = getModel("anthropic", "anthropic-model-haiku-4-5-20251001");
const continuation = await complete(newModel, restored);
Note: If the context contains images (encoded as base64 as shown in the Image Input section), those will also be serialized.
Browser Usage
The library supports browser environments. You must pass the API key explicitly since environment variables are not available in browsers:
import { getModel, complete } from "@jawcode-dev/ai";
const model = getModel("anthropic", "anthropic-model-haiku-4-5-20251001");
const response = await complete(
model,
{
messages: [{ role: "user", content: "Hello!" }],
},
{
apiKey: "your-api-key",
}
);
Security Warning: Exposing API keys in frontend code is dangerous. Anyone can extract and abuse your keys. Only use this approach for internal tools or demos. For production applications, use a backend proxy that keeps your API keys secure.
Environment Variables (Node.js only)
In Node.js environments, you can set environment variables to avoid passing API keys:
| OpenAI | OPENAI_API_KEY |
| Anthropic | ANTHROPIC_API_KEY or ANTHROPIC_OAUTH_TOKEN (or ANTHROPIC_FOUNDRY_API_KEY when ANTHROPIC_MODEL_CODE_USE_FOUNDRY=true) |
| Google | GEMINI_API_KEY |
| Vertex AI | GOOGLE_CLOUD_PROJECT (or GCLOUD_PROJECT) + GOOGLE_CLOUD_LOCATION + ADC |
| Mistral | MISTRAL_API_KEY |
| Groq | GROQ_API_KEY |
| Cerebras | CEREBRAS_API_KEY |
| Together | TOGETHER_API_KEY |
| Qianfan | QIANFAN_API_KEY |
| Hugging Face | HUGGINGFACE_HUB_TOKEN or HF_TOKEN |
| Synthetic | SYNTHETIC_API_KEY |
| NVIDIA | NVIDIA_API_KEY |
| NanoGPT | NANO_GPT_API_KEY |
| Venice | VENICE_API_KEY |
| Moonshot | MOONSHOT_API_KEY |
| xAI | XAI_API_KEY |
| OpenRouter | OPENROUTER_API_KEY |
| LiteLLM | LITELLM_API_KEY |
| Ollama | OLLAMA_API_KEY (optional for local deployments) |
| Ollama Cloud | OLLAMA_CLOUD_API_KEY |
| Qwen Portal | QWEN_OAUTH_TOKEN or QWEN_PORTAL_API_KEY |
| zAI | ZAI_API_KEY |
| MiniMax Code | MINIMAX_CODE_API_KEY (international) or MINIMAX_CODE_CN_API_KEY (China) |
| Xiaomi MiMo | XIAOMI_API_KEY |
| ZenMux | ZENMUX_API_KEY |
| vLLM | VLLM_API_KEY |
| Cloudflare AI Gateway | CLOUDFLARE_AI_GATEWAY_API_KEY |
| GitHub Copilot | COPILOT_GITHUB_TOKEN or GH_TOKEN or GITHUB_TOKEN |
For Cloudflare AI Gateway models, use provider base URL format
https://gateway.ai.cloudflare.com/v1/<account>/<gateway>/anthropic.
For Anthropic Foundry routing, set ANTHROPIC_MODEL_CODE_USE_FOUNDRY=true plus:
FOUNDRY_BASE_URL, ANTHROPIC_FOUNDRY_API_KEY, optional ANTHROPIC_CUSTOM_HEADERS,
and optional mTLS material (ANTHROPIC_MODEL_CODE_CLIENT_CERT, ANTHROPIC_MODEL_CODE_CLIENT_KEY, NODE_EXTRA_CA_CERTS).
Provider endpoint defaults for the current OpenAI-compatible integrations:
- Together:
https://api.together.xyz/v1
- Moonshot:
https://api.moonshot.ai/v1
- Qianfan:
https://qianfan.baidubce.com/v2
- NVIDIA:
https://integrate.api.nvidia.com/v1
- NanoGPT:
https://nano-gpt.com/api/v1
- Hugging Face Inference:
https://router.huggingface.co/v1
- Venice:
https://api.venice.ai/api/v1
- Xiaomi MiMo:
https://api.xiaomimimo.com/anthropic
- ZenMux (OpenAI):
https://zenmux.ai/api/v1
- ZenMux (Anthropic models):
https://zenmux.ai/api/anthropic
- vLLM:
http://127.0.0.1:8000/v1
- Ollama: local OpenAI-compatible runtime (
http://127.0.0.1:11434/v1)
- Ollama Cloud: native Ollama API host (
https://ollama.com/api, configured here as base URL https://ollama.com)
- LiteLLM:
http://localhost:4000/v1
- Cloudflare AI Gateway:
https://gateway.ai.cloudflare.com/v1/<account>/<gateway>/anthropic
- Qwen Portal:
https://portal.qwen.ai/v1
When set, the library automatically uses these keys:
const model = getModel("openai", "gpt-4o-mini");
const response = await complete(model, context);
const response = await complete(model, context, {
apiKey: "sk-different-key",
});
Checking Environment Variables
import { getEnvApiKey } from "@jawcode-dev/ai";
const key = getEnvApiKey("openai");
OAuth Providers
Several providers support OAuth authentication (some also support static API keys):
- Anthropic (Anthropic model Pro/Max subscription)
- OpenAI code provider (ChatGPT Plus/Pro subscription, access to GPT-5.x OpenAI code models)
- GitHub Copilot (Copilot subscription)
- Google Gemini CLI (Gemini 2.0/2.5 via Google Cloud Code Assist; free tier or paid subscription)
- Antigravity (Free Gemini 3, Anthropic model, GPT-OSS via Google Cloud)
- Qwen Portal (Qwen OAuth token or API key)
- xAI (Grok OAuth login via xAI account)
For paid Cloud Code Assist subscriptions, set GOOGLE_CLOUD_PROJECT or GOOGLE_CLOUD_PROJECT_ID to your project ID.
Vertex AI (ADC)
Vertex AI models use Application Default Credentials (ADC):
- Local development: Run
gcloud auth application-default login
- CI/Production: Set
GOOGLE_APPLICATION_CREDENTIALS to point to a service account JSON key file
Also set GOOGLE_CLOUD_PROJECT (or GCLOUD_PROJECT) and GOOGLE_CLOUD_LOCATION. You can also pass project/location in the call options.
Example:
gcloud auth application-default login
export GOOGLE_CLOUD_PROJECT="my-project"
export GOOGLE_CLOUD_LOCATION="us-central1"
export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
import { getModel, complete } from "@jawcode-dev/ai";
(async () => {
const model = getModel("google-vertex", "gemini-2.5-flash");
const response = await complete(model, {
messages: [{ role: "user", content: "Hello from Vertex AI" }],
});
for (const block of response.content) {
if (block.type === "text") console.log(block.text);
}
})().catch(console.error);
Official docs: Application Default Credentials
CLI Login
The quickest way to authenticate:
bunx @jawcode-dev/ai login
bunx @jawcode-dev/ai login anthropic
bunx @jawcode-dev/ai login vllm
bunx @jawcode-dev/ai login xai
bunx @jawcode-dev/ai list
Credentials are saved to agent.db in the agent directory. /login qianfan opens the Qianfan console and stores the pasted API key; /login xai opens xAI/Grok OAuth login and stores refreshable OAuth credentials.
login supports OAuth providers (Anthropic, OpenAI code provider, GitHub Copilot, Gemini CLI, Antigravity, xAI) and API-key onboarding flows.
For the current API-key onboarding flows, the library covers Together, Moonshot, Qianfan, NVIDIA, NanoGPT, Hugging Face, Venice, Xiaomi, vLLM, LiteLLM, Cloudflare AI Gateway, Qwen Portal, and Ollama Cloud. Ollama remains the local runtime integration; set OLLAMA_API_KEY only when your local or self-hosted deployment enforces bearer auth.
Programmatic OAuth
The library provides login and token refresh functions. Credential storage is the caller's responsibility.
import {
loginAnthropic,
loginOpenAIOpenAI code,
loginGitHubCopilot,
loginGeminiCli,
loginAntigravity,
loginCloudflareAiGateway,
loginHuggingface,
loginLiteLLM,
loginMoonshot,
loginNvidia,
loginNanoGPT,
loginQianfan,
loginQwenPortal,
loginTogether,
loginVenice,
loginVllm,
loginXiaomi,
refreshOAuthToken,
getOAuthApiKey,
type OAuthProvider,
type OAuthCredentials,
} from "@jawcode-dev/ai";
loginOpenAIOpenAI code accepts an optional originator value used in the OAuth flow:
await loginOpenAIOpenAI code({
onAuth: ({ url }) => console.log(url),
originator: "my-cli",
});
Login Flow Example
import { loginGitHubCopilot } from "@jawcode-dev/ai";
import * as fs from "node:fs";
const credentials = await loginGitHubCopilot({
onAuth: (url, instructions) => {
console.log(`Open: ${url}`);
if (instructions) console.log(instructions);
},
onPrompt: async (prompt) => {
return await getUserInput(prompt.message);
},
onProgress: (message) => console.log(message),
});
const auth = { "github-copilot": { type: "oauth", ...credentials } };
fs.writeFileSync("credentials.json", JSON.stringify(auth, null, 2));
Using OAuth Tokens
Use getOAuthApiKey() to get an API key, automatically refreshing if expired:
import { getModel, complete, getOAuthApiKey } from "@jawcode-dev/ai";
import * as fs from "node:fs";
const auth = JSON.parse(fs.readFileSync("credentials.json", "utf-8"));
const result = await getOAuthApiKey("github-copilot", auth);
if (!result) throw new Error("Not logged in");
auth["github-copilot"] = { type: "oauth", ...result.newCredentials };
fs.writeFileSync("credentials.json", JSON.stringify(auth, null, 2));
const model = getModel("github-copilot", "gpt-4o");
const response = await complete(
model,
{
messages: [{ role: "user", content: "Hello!" }],
},
{ apiKey: result.apiKey }
);
Provider Notes
OpenAI code provider: Requires a ChatGPT Plus or Pro subscription. Provides access to GPT-5.x OpenAI code models with extended context windows and reasoning capabilities. The library automatically handles session-based prompt caching when sessionId is provided in stream options.
GitHub Copilot: If you get "The requested model is not supported" error, enable the model manually in VS Code: open Copilot Chat, click the model selector, select the model (warning icon), and click "Enable".
Google Gemini CLI / Antigravity: These use Google Cloud OAuth. The apiKey returned by getOAuthApiKey() is a JSON string containing both the token and project ID, which the library handles automatically.
License
MIT