New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

@aimarket/warden

Package Overview
Dependencies
Maintainers
1
Versions
5
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

@aimarket/warden

WARDEN — MCP security firewall. Vets an MCP server's tool definitions, threat feed, origin and pinning before any tool runs. Zero runtime dependencies.

Source
npmnpm
Version
0.3.0
Version published
Weekly downloads
51
-16.39%
Maintainers
1
Weekly downloads
 
Created
Source

WARDEN — MCP security firewall

CI npm version Zero runtime dependencies 95 tests passing Node >= 20 License: MIT

🌐 English · Русский · Español · Français · 中文 · Glossary

An MCP server tells your agent what its tools do. The agent believes it — that sentence is the attack surface. A tool description is prompt text delivered by a third party straight into your model's context, and a schema field named api_key is a request for your secrets phrased as an API.

WARDEN vets a server before any of its tools reach the model, and returns a verdict you can record: allow/block, a 0..1 score, the findings that produced it, a per-tool partition, and the exact rule table that was in force.

npm install @aimarket/warden

Zero runtime dependencies. The only import in the whole package is node:crypto. It is the firewall out of ARGUS, extracted so you can put it in front of your own MCP host without adopting an agent.

Quick start

import { Warden, ThreatFeed, silentLogger } from "@aimarket/warden";

const threatFeed = new ThreatFeed({ feedPublicKey: process.env.FEED_PUBKEY });
await threatFeed.load(process.env.FEED_URL); // omit → built-in deny-list only, no network

const pins = new Map();
const warden = Warden.create({
  policy: {
    blockAtSeverity: "high",
    sensitiveToolPatterns: ["*delete*", "*transfer*", "*key*"],
    allowUnknownServers: false, // fail-closed: only servers you declared
    pinToolDefs: true,
  },
  threatFeed,
  store: {
    getPin: async (id) => pins.get(id),
    putPin: async (p) => void pins.set(p.serverId, p),
  },
  log: silentLogger(), // or your own logger
});

const verdict = await warden.vet(server, await client.listTools());

if (!verdict.allow) throw new Error(`blocked by ${verdict.decidedBy}`);
const usable = verdict.allowedTools; // a poisoned tool can be quarantined alone
await warden.approve(server, tools); // pin what the user accepted

vet() performs no network I/O. The only request WARDEN ever makes is the threat-feed fetch you asked for by passing a URL to load().

The gate chain

flowchart LR
  T["tool defs<br/>from the server"] --> S["static scan<br/>25 rules"]
  S --> F["threat feed<br/>11 built-ins + signed"]
  F --> O["origin<br/>declared vs catalog"]
  O --> P["pinning<br/>drift vs approval"]
  P --> V["verdict<br/>allow · score · findings<br/>allowedTools / blockedTools"]
GateWhat it decidesNetworkFatal?
static-scanInjection, exfiltration, credential requests and hidden-Unicode/base64 tells in description and inputSchema — 25 rules, v2, of which 18 can block and 7 are advisory-onlynoneno
threat-feedKnown-bad server identity or tool, from 11 built-in records plus an optional signed feedonly the feed fetchyes, for a server-scoped critical
originWhether the operator declared this server or it arrived from a remote catalognoneyes, under allowUnknownServers: false
pinningWhether the tool defs still match what the user approvednoneyes, under pinToolDefs: true

The composite score is the product of gate contributions, so one bad gate drags the whole server down rather than being averaged away. Severity and blocking are separate axes: an advisory finding is reported and never blocks and never costs a tool, at any blockAtSeverity — because "how much attention does this deserve" and "is this a defect at all" are different questions, and encoding the second as a low severity made it blocking again for anyone who tightened the threshold.

The verdict is meant to be recorded

{
  allow: false,
  score: 0,
  decidedBy: "threat-feed",
  findings: [{ gate, severity, code: "THREAT_TOOL_MATCH", message, tool, advisory? }],
  allowedTools: ["add"],
  blockedTools: ["sweeper"],
  rulesets: { staticScan: { version: "2", digest: "sha256-gWC14PR4…" } }
}

rulesets is not decoration. The same server scores differently under a later rule table, and without the version and a digest over the rules there is no way to tell that apart from the server having changed. A stored scan without them is not reproducible.

Signed threat feed

WARDEN will not read an unsigned remote feed. The contract is deliberately boring:

GET <your feed url>
{ "records": [ {pattern, severity, code, reason, source, scope}, … ],
  "timestamp": 1786205907380,   // epoch ms, integer — required
  "signature": "f588d5a4…"      // Ed25519 (hex) over the RFC 8785 canonical
}                               // form of {records, timestamp}

Three properties are checked, and any failure keeps the built-in floor rather than degrading to no protection:

  • authenticity — Ed25519 against the key you pinned in advance (feedPublicKey);
  • freshness — the signed timestamp must be inside maxAgeMs (24 h by default), so whoever serves the URL cannot replay a months-old snapshot and silently erase every record added since. A signature says who wrote a document, never when you were handed it;
  • determinism — RFC 8785 canonical bytes, so publisher and verifier agree regardless of JSON key order.

MOMUS is a reference publisher of this contract (/warden/threat-feed) if you want something to point load() at.

Also in the box

  • EgressGuard — an outbound allowlist to wrap any request a tool makes. A tool reaching a host you never listed is the classic phone-home tell. *.example.com matches subdomains; an empty allowlist blocks everything rather than allowing everything.
  • isSensitiveTool / classifyTools — glob classification of tools that must require per-call approval. Sensitive tools stay advertised; they just cannot run unattended.
  • canonicalize / parseJsonStrict — a strict RFC 8785 (JCS) implementation, also exported as @aimarket/warden/jcs so another implementation can be byte-checked against it. Integers only beyond MAX_SAFE_JSON_INTEGER, refusal (not escaping) on lone surrogates, and a reason code on every refusal.

Documentation

The gate chainEvery rule tier, every finding code, how the composite score is built, and how to add a gate
The signed threat feedThe wire contract, the three checks, and how to publish a feed WARDEN will accept
Integration guideWiring WARDEN into your own MCP host, policy choices, and what to record

What this is not

  • Not a sandbox. These are in-process JS decisions. OS-level confinement of the MCP child process (seccomp/Landlock, sandbox-exec) is not here.
  • Not a model. No LLM is called anywhere in the chain. That is why vet() is fast, offline and deterministic — and why the static scan is regex-shaped and will miss a paraphrase no rule covers.
  • Not a reputation service. An earlier version had a gate that asked a trust oracle for a score it had no data to compute, then reported the oracle as unreachable without having sent a request. It was removed, and test/no-phantom-gate.test.ts fails if any gate ever claims unreachability again.
  • Not a substitute for reading the tool defs. 11 built-in threat records is a floor, not a catalog.

Development

npm install && npm run build && npm test   # 95 tests

test/packaging.test.ts is what keeps the headline honest: it fails if a runtime dependency appears, if any source file imports outside the package, or if the entry point stops exporting the enforcement surface.

Used by ARGUS (the reference host), MOMUS (the publisher side), and the AICOM MCP-security course.

MIT © AICOM (alexar76)

Keywords

mcp

FAQs

Package last updated on 24 Aug 2026

Related posts