🎩 You're Invited:Meet the Socket team at Black Hat in Las Vegas, August 3-6.RSVP
Sign In

podium-mcp

Package Overview
Dependencies
Maintainers
1
Versions
4
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

podium-mcp

Mobile + canvas automation MCP for AI agents — one stdio server, 60 tools: iOS (simulator + real) & Android control, native UI automation, Maestro E2E flows, evidenced oracle-ladder assertions, React Native/Metro debugging, WebView DOM + network capture,

latest
Source
npmnpm
Version
0.5.0
Version published
Weekly downloads
64
-30.43%
Maintainers
1
Weekly downloads
 
Created
Source

podium-mcp

One baton. Every instrument.

A single MCP stdio endpoint with 60 tools for iOS (simulator + real) and Android device control, native UI automation, end-to-end flows, trustworthy assertions, React Native debugging, WebView DOM + network inspection, and a no-vision canvas/WebGL brain for Pixi/Konva/Fabric/Phaser/Three/Babylon (validated live in WebKit) — plus an experimental engine bridge for instrumented Unity/GL builds (AltTester) — one connection instead of half a dozen servers.

npm Glama mcp.so CI License: MIT

tools 60 tests 1069 tokens ~5x cheaper canvas 6 live engines Node ≥22 TypeScript strict MCP stdio PRs welcome


podium-mcp Canvas Brain — no-vision inspect, resolve, and tap on a live Pixi.js game, ~5x cheaper than screenshot loops

Canvas Brain in action: canvas_inspect detects Pixi objects → canvas_resolve scores candidates → canvas_tap executes — all without vision. 81% cheaper than screenshot loops.

A podium is where a maestro stands — one place to conduct the whole orchestra. This MCP server unifies eight capability sets behind a single stdio endpoint:

  • Device & app management — iOS simulators (simctl), real iPhones (devicectl), and Android (adb) behind one platform-tagged device model.
  • Native UI inspection & gestures — route through idb/mobilecli with a Maestro fallback (no per-gesture JVM spin-up).
  • End-to-end flows & batch automation — declarative Maestro flows, ordered action batches, and an engineer→QA flow exporter.
  • Trustworthy assertions — an oracle ladder (WebView-DOM › native a11y › Maestro) that returns falsifiable, evidenced verdicts and fails closed.
  • WebView DOM + network — resolve WKWebView DOM to tap coordinates, evaluate JS, drive navigation, and capture in-page HTTP traffic as JSON/HAR.
  • React Native debugging — Metro console logs, network requests, and in-app state over CDP, plus host/simulator crash reports.
  • Real devices — Android emulator/device via adb (gestures + uiautomator hierarchy); real iOS via devicectl lifecycle + an opt-in WebDriverAgent backend.
  • Canvas & game-engine automation, no vision — a canvas/WebGL brain drives Pixi/Konva/Fabric/Phaser/Three/Babylon UIs as addressable objects (validated live in WebKit). An experimental engine bridge drives Unity/GL via an AltTester-instrumented build (or a window.__podiumEngine WebGL bridge) — code-complete + mock-tested, not yet run against a live Unity build.

Rather than wiring several MCP servers into every client config, podium-mcp exposes everything behind one connection, with a shared execFile layer (no shell), consistent structured errors, automatic retry around Maestro's iOS-driver flakiness, and a single health-check tool to confirm what's available on the host.

Table of contents

Why

Driving a React Native app end-to-end usually means juggling several MCP servers — one for device/app control, one for UI flows, one for Metro/debugger logs, another for WebView inspection — each with its own config entry, quirks, and failure modes. podium-mcp collapses that into one server with:

  • a single execFile-based command runner (no shell — arguments are passed verbatim),
  • consistent structured errors (a tool never crashes the server),
  • automatic retry around Maestro's known iOS-driver flakiness,
  • graceful degradation when a toolchain (e.g. adb) is absent,
  • evidenced verdicts so an agent knows when a flow actually worked.

v0.5.0 highlights

v0.5.0 merges the deterministic-resolver and auto-schema work into one release: every element on screen gets a stable semantic ID, forms can be auto-analyzed into a fillable schema, and run_intent can drive that schema with zero per-step model inference — while staying a strict, backward-compatible superset of the existing free-form flow.

  • Semantic IDs — every native element gets a semanticId, either read from an explicit attribute or auto-generated as {role}_{label}_{position}. The resolver checks it first: an exact semantic-ID match is a confidence: 1.0 pick with zero reasoning, so a weak model and a strong one target the identical element. See docs/semantic-ids.md.
  • Deterministic resolver — a strict priority ladder (semantic ID → CSS/ XPath selector → partial semantic ID → label/synonym scoring) that only falls back to fuzzy inference when nothing deterministic matches, and stays fail-closed (AMBIGUOUS, top-3 candidates) whenever it can't decide. (src/lib/resolver.ts, src/lib/selector.ts.)
  • Auto-schema generationgenerateSchema() reads a live accessibility tree and auto-detects the form container, classifies every field (text/choice/checkbox/switch), and clusters related buttons (e.g. a row of mood buttons) into named choice groups — no manual schema authoring. See docs/schema-generation.md.
  • Schema-driven flows — pass a generated schema + field values to run_intent and it fills the form deterministically: confidence: 1.0, tokensUsed: 0, ~85–87% fewer tokens per step than free-form on Opus/ Sonnet, and the only supported mode on Haiku. See docs/schema-driven-flows.md.
  • Full model support, model-aware cost tracking — Opus, Sonnet, and Haiku are all ✅ (see the model support matrix below); run_intent detects the tier and returns a modelUsed + tokenEstimate on every call. Measured savings vs. the v0.4.x baseline: ~40% (Opus), ~50% (Sonnet), ~40% (Haiku). See docs/model-support.md.
  • Fully backward compatibleschema is opt-in; omitting it drives the existing free-form steps[] path unchanged.

Benchmarks

Podium is built on two choices that make it fast and cheap: it drives UIs as structured data — never screenshots — and routes gestures through a native backend with no per-action JVM spin-up.

Token economics — no-vision is ~5× cheaper

A screenshot-driven agent sends an image to a vision model on every step. Podium returns a compact structured element list instead. On an equivalent 8-step mobile flow (1179×2556 screenshots vs ~20-element lists):

ApproachPer step8-step flow
Screenshot / vision loop~2,070 tokens16,557 tokens
Podium — no-vision, structured~390 tokens3,117 tokens
Savings5.3×−13,440 tokens (−81%)
vision loop  ████████████████████████████████  16,557 tokens
Podium       ██████  3,117 tokens   (5.3× cheaper, −81%)

The gap compounds with every step — a 30-step session runs roughly 62k vs 12k input tokens. On top of per-step cost, the full 60-tool schema travels with every request (~3,612 tokens, ~71/tool); Podium keeps tool descriptions lean so the tool block never dominates the context window.

For canvas / WebGL UIs the advantage is structural, not just cheaper: the Canvas Brain addresses objects by name and text, where a screenshot-only agent must re-analyze pixels on every frame.

Speed — native-first gesture backend

Gestures route through idb / mobilecli instead of spinning up Maestro's JVM per action (measured on a live iPhone 16 Pro simulator):

OperationMaestro (per-call JVM)Podium nativeSpeedup
tap_on~14.7 s~0.6 s~24×
inspect_screen~8.9 s~0.9 s~10×

One connection, not six

All 60 tools — device & app control, UI automation, declarative Maestro flows, evidenced assertions, WebView DOM + network capture, React Native / Metro debugging, and no-vision canvas/WebGL automation (plus an experimental engine bridge for instrumented Unity/GL) — sit behind a single stdio endpoint, replacing the usual stack of half a dozen separate MCP servers.

Token figures are heuristic estimates (~4 chars/token; Anthropic's ~750 px/token image formula) — reproduce with npm run token-bench, or swap in the Anthropic count_tokens API for exact counts. Speed figures were measured on a live iPhone 16 Pro simulator (npm run benchmark).

Model support matrix

Podium is model-agnostic: the same intent PLAN drives every tier. run_intent detects the current tier (execution context → PODIUM_MODEL env → step-count heuristic → Sonnet default) and returns a modelUsed + tokenEstimate so callers get model-aware cost tracking with zero extra model round-trips.

ModelSupportCapabilitiesFree-formSchema-driven
Opusfree-form · complex-reasoning · multi-step · schema-driven✅ 150–200 tok/step✅ 20–30 tok/step
Sonnetfree-form · schema-driven · standard-forms✅ 80–120 tok/step✅ 10–15 tok/step
Haikuschema-driven · deterministic-only— (unsupported)✅ 5–10 tok/step

See docs/model-support.md for per-tier capability details, token-usage tables, and a relative cost comparison.

Token efficiency (measured savings)

Schema-driven mode replaces free-form inference with a deterministic field→element map, cutting per-step tokens sharply. Versus the v0.4.x baseline:

ModelSavings vs v0.4.xSchema-driven vs free-form
Opus~40%~85%
Sonnet~50%~87%
Haiku~40%n/a (no free-form baseline)

Estimate any flow programmatically:

import { detectCurrentModel, estimateTokens, recommendMode } from "podium-mcp/lib/intent";

const model = detectCurrentModel();                 // e.g. { model: "sonnet", ... }
estimateTokens("schema-driven", 5, model);          // → { estimated: 60, range: { min: 50, max: 75 }, savings: 87, supported: true }
recommendMode(model, "complex");                     // → { recommended: "schema-driven", reason: "…" }

Model-aware best practices for callers

  • Haiku supports schema-driven only — always pass a UISchema; free-form inference is unavailable (estimateTokens reports supported: false).
  • Sonnet — prefer schema-driven for every complexity: ~87% cheaper than free-form with equal reliability.
  • Opus — use free-form for genuinely complex forms (best accuracy); schema-driven for simple/medium forms to save ~85% with no accuracy loss.
  • Pin the tier explicitly with PODIUM_MODEL=opus|sonnet|haiku when the host does not surface a model id; otherwise Podium infers it from plan size and defaults to Sonnet.

Requirements

  • macOS with Xcode command-line tools (xcrun, simctl)
  • Node.js ≥ 22 (uses native fetch and WebSocket; .npmrc sets engine-strict=true)
  • mobilecli — bundled automatically as an npm dependency; the default native gesture + WebView backend (no separate install)
  • (optional) idb (idb + idb_companion) — preferred native gesture backend when both are present; auto-detected
  • (optional) Maestro on PATH (or at ~/.maestro/bin) — the run_flow engine and the gesture fallback path
  • (optional) a running Metro bundler for the metro_* debugging tools
  • (optional) Android SDK + adb — adb paths are detection-only and degrade gracefully when absent

Platform scope (v0.3.0): podium automates iOS simulators, real iPhones (devicectl lifecycle + opt-in WebDriverAgent), and Android emulators/devices (adb gestures + uiautomator hierarchy). device_list tags each target with its platform and the backend is selected per target. When a toolchain (e.g. adb) is absent, those paths degrade to an informative result instead of failing.

Install

No manual config — one-time marketplace setup, then install:

/plugin marketplace add github:hoainho/podium-mcp
/plugin install podium-mcp@podium

The plugin auto-starts the MCP server (all 60 tools) and ships five skills:

SkillInvokeWhat it does
Device info/podium-mcp:device-info <UDID> [<BUNDLE_ID>]Health check, screen size, orientation, app list
E2E flow/podium-mcp:e2e <UDID> <BUNDLE_ID> [path or description]Run or author a Maestro flow
Bug repro/podium-mcp:bug-repro <UDID> <BUNDLE_ID> <description>Video + logs + crash evidence capture
RN debug/podium-mcp:rn-debug [UDID] [logs|apps|crash|all]Metro logs, connected apps, crash reports
Canvas brain/podium-mcp:canvas <UDID> <intent>Inspect / resolve / tap canvas-WebGL UIs, no vision

npx (zero install)

{
  "mcpServers": {
    "podium": { "command": "npx", "args": ["-y", "podium-mcp"] }
  }
}

Manual (from source)

git clone git@github.com:hoainho/podium-mcp.git
cd podium-mcp
npm install
npm run build

Usage

Register the built server with any MCP client. Claude Code (.mcp.json):

{
  "mcpServers": {
    "podium": {
      "type": "stdio",
      "command": "node",
      "args": ["/absolute/path/to/podium-mcp/dist/index.js"]
    }
  }
}

Quick manual smoke test over raw stdio (lists the 60 registered tools):

printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"smoke","version":"0"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' | node dist/index.js

Always call podium_health first to confirm which toolchain is available on the host.

Quick start (order of use)

  • podium_health — confirm xcrun / maestro / native backend availability.
  • device_list — pick a booted simulator udid.
  • Read stateapp_list, app_state, screen_size, orientation_get.
  • Drive the deviceapp_launch, then tap_on / input_text / swipe / press_key, plus set_location and orientation_set. Batch several with run_steps.
  • Author & verifyinspect_screen to discover elements, run_flow for declarative checks, then assert_visible / validate_flow for an evidenced verdict.
  • Inspect WebViewswebview_inspect → tap coordinates, webview_eval, webview_navigate, webview_network.
  • Capture & debugscreenshot / record_startrecord_stop; metro_logs / metro_network / metro_state; crash_list / crash_get.

The 60 tools

Every tool returns structured JSON and never throws — failures come back as MCP tool errors. See docs/tool-catalog.md for the authoritative per-parameter reference.

Platform support (v0.3.0): the gesture / inspect / lifecycle tools below run on iOS simulators, real iPhones (devicectl + opt-in WebDriverAgent via PODIUM_WDA_URL), and Android (emulator/device via adb; hierarchy from uiautomator). device_list tags each device with its platform and the backend is selected per target.

Game engine — Unity / GL via AltTester, no vision · experimental (4)

ToolKey paramsBacking engineBehavior
engine_inspectudid, by?, valueAltTester (TCP) / WebGL CDP bridgeLists engine objects (by name/path/component/text) with absolute screen coords — no screenshots
engine_tapudid, by?, valueAltTester / CDPResolves the object and taps its screen coordinates
engine_swipeudid, fromX/Y, toX/Y, durationMs?AltTester / CDPSwipe inside the engine view
engine_calludid, by?, value, component, method, parameters?AltTester / CDPInvokes a C# component method by reflection (the engine analog of a DOM event handler)

Status: experimental. The wire shapes are unit-tested against mocks; the AltTester path has not yet been validated against a live Unity build (engine-smoke skips until an instrumented build is provided), and Unity-WebGL needs the app to expose window.__podiumEngine. Engine tools require an AltTester-instrumented build (dev/staging) or that WebGL bridge; on a non-instrumented build they fail closed with an actionable error — never a vision fallback. For canvas/WebGL apps using a JS framework, the canvas brain below is the validated path.

Canvas brain — Pixi/Konva/Fabric/Phaser/Three/Babylon, no vision (3)

ToolKey paramsBacking engineBehavior
canvas_inspectudid, by?, value?, webviewId?injected scene-graph bridge (CDP eval)Lists canvas objects with tap-ready CSS-px coords — no screenshots
canvas_resolveudid, intent, webviewId?bridge + semantic resolverMaps a fuzzy intent ("close", "✕") to a ranked, evidenced target; fail-closed confidentEnough
canvas_tapudid, intent, bundleId?, webviewId?resolver + native tapResolves + taps the confident match at absolute screen coords (else fails closed)

Validated live: all six frameworks pass a Playwright-WebKit (≈ WKWebView) suite at DPR 1 + 3 (npm run test:canvas, 19 of its 21 tests — the other 2 are the Rive web-runner's live closed-loop check). Canvas tools require an inspectable WKWebView hosting a supported framework with its root reachable (commonly on window, or Pixi's __PIXI_APP__). No framework / no inspectable WebView → fails closed with an actionable error — never a vision fallback.

Rive — drive & analyze Rive canvases, no vision (8)

ToolKey paramsBacking engineBehavior
rive_inspectudid, instance?, webviewId?injected __podiumRive bridge (CDP eval)Auto-schema of the canvas: SM inputs + view-model tree + text runs + artboards + events — no screenshots
rive_analyzeudid?, instance?, webviewId?, src?bridge (live) + @rive-app/canvas-advanced WASM (static src)Read-only complete element inventory + per-type counts + triggerable rollup + states, merged from a live device and/or a static .riv src, plus an advice field (score/grade/findings from the read-only best-practice advisor)
rive_fireudid, name, stateMachine?, instance?bridgeFire a trigger — SM input or view-model trigger — by name//-path
rive_setudid, name, value, stateMachine?, instance?bridgeSet a boolean/number — SM input or view-model property — by name//-path
rive_set_textudid, name, value, instance?bridgeSet a view-model string (or text run) by name//-path
rive_eventsudid, instance?bridgeDrain buffered Rive events (the verify oracle)
rive_runudid, steps, instance?bridgeDrive a whole ordered flow (fire/set/setText) + verify, in one call
rive_e2e_inittargetDir, src?, udid?, webviewId?, url?, appId?, webview?, threshold?, flowName?rive_analyze + scaffolderScaffolds a coverage-gated Rive E2E kit into a target repo — rive-e2e/ config, frozen contract, generated auto-smoke suite, web/mobile/drift runners, and a CI workflow (the podium init rive-e2e capability)

Rive canvases have no display-object tree to tap by coordinate — you drive them by their runtime (named, typed inputs / view-model paths), never by pixels. rive_analyze's static source runs in Node with no device. Full guide: docs/rive.md.

Rive E2E kit + advisor. rive_e2e_init (or the podium init rive-e2e CLI) turns that same inventory into a measured coverage number — coverage = |driven ∩ verified| / |addressable surface|, gated in CI — by mechanically driving every contract knob through the same bridge and verifying the effect (event drain / read-back). The advisor is a separate, read-only lint pass over the identical inventory: a codified best-practice catalog producing a score/grade/findings, also usable standalone (no device, no MCP server) via import { adviseRive } from "podium-mcp/analyzer". Quickstart, the coverage identity, and honest boundaries (web/WebView-only live drive+verify, static-vs-live gaps): docs/rive-e2e.md.

Diagnostics (1)

ToolKey paramsBacking engineBehavior
podium_token_reportsteps?, screenshotWidth?, screenshotHeight?, elementsPerStep?, toolCount?token estimatorsNo-vision vs screenshot/vision-loop input tokens, the savings ratio, and the per-request tool-definition overhead

Health & toolchain (1)

ToolKey paramsBacking engineBehavior
podium_healthwhich probesNever fails; reports toolchain { xcrun, maestro, adb }, native backend, and platforms: [ios-sim, ios-real, android]

Device & simulator (6)

ToolKey paramsBacking engineBehavior
device_listsimctl list -j + adb devicesMerged iOS inventory; adb absent → android: { available: false } (detection-only)
device_bootudidsimctl bootIdempotent — already-booted → alreadyBooted: true; waits up to 30 s
screen_sizeudidsimctl io screenshot + sips{ widthPx, heightPx } (real pixels)
orientation_getudidnative query → screenshot heuristic{ orientation, basis } (exact when native)
set_locationudid, latitude, longitudesimctl location setCodifies the QA geo-spinner fix
open_urludid, urlsimctl openurlDeep links + https://

Apps (6)

ToolKey paramsBacking engineBehavior
app_installudid, path (.app/.zip)simctl installStructured tool error
app_launchudid, bundleIdsimctl launchExplicit 30 s timeout (cold RN launches no longer mis-report failure)
app_terminateudid, bundleIdsimctl terminateStructured tool error
app_uninstalludid, bundleIdsimctl uninstallStructured tool error
app_listudidsimctl listapps + plutil{ count, apps: [{ bundleId, name, type }] }
app_stateudid, bundleIdsimctl listapps + launchctl{ installed, running }exact bundle-id match

Capture (3)

ToolKey paramsBacking engineBehavior
screenshotudid, saveTo?simctl io screenshotReturns path + byteSize (no base64 bloat)
record_startudid, saveTo? (.mp4)detached simctl io recordVideo{ ok, path, pid }; timestamped path + duration watchdog (PODIUM_MAX_RECORDING_MS); one per udid
record_stopudidSIGINT recorder + flush{ ok, path, sizeBytes }

UI inspection & gestures (8)

ToolKey paramsBacking engineBehavior
inspect_screenudid, compact?native flat AX list → maestro hierarchycompact:true (default) returns only meaningful nodes
tap_onudid, bundleId, text|id|x+y, double?, long?native tap → Maestro fallbacktext/id resolved via the element list; reports backend
input_textudid, bundleId, text, submit?native → Maestro fallbackreports backend
swipeudid, bundleId, direction, start/end?native → Maestro fallback%/pixel overrides resolved vs logical screen size
press_keyudid, bundleId, keynative → Maestro fallbackback/power/tab are Android-only
orientation_setudid, bundleId, valuenative → Maestro fallbackPORTRAIT / LANDSCAPE_LEFT / LANDSCAPE_RIGHT / UPSIDE_DOWN
tap_with_fallbackudid, x, y, maxRetries?, offsetStep?native tap + before/after oracleFor WebGL/Canvas overlays; no blind walk (offsetStep opt-in)
notification_bar_clearudid, bundleId?native tap + oracleDismisses the RN debug notification bar

Flows & batch automation (5)

ToolKey paramsBacking engineBehavior
run_stepsudid, bundleId, steps[]native backend (idb/mobilecli)Ordered action batch in one call; per-step results
run_flowudid + exactly one of yaml/files/dir(+tags), env?maestro testExactly-one-of validated before exec; per-step pass/fail
run_intentudid, bundleId?, goal?, steps[] OR schema+paramsresolver + oracle ladderFree-form (fuzzy target resolution, per-step verify) or schema-driven (0-inference form fill, confidence: 1.0) intent execution in one closed loop
export_flowsteps[], output pathflow generatorExports a run_steps batch to a reusable Maestro flow (engineer→QA bridge)
cheat_sheetbundled assets/maestro-cheat-sheet.yamlFully offline Maestro syntax reference

Assertions & verdicts — the oracle ladder (5)

ToolKey paramsBacking engineBehavior
assert_visibleudid, text|id, …oracle ladder (WebView-DOM › a11y › Maestro)Evidenced pass/fail; reports which oracle proved it
assert_textudid, textoracle ladderby-text shorthand for assert_visible
assert_not_visibleudid, text|idoracle ladderFails closed — if absence can't be verified, it fails
wait_for_elementudid, text|id, timeoutMs?oracle ladder (polling)Polls until visible or times out
validate_flowudid, flow + assertionsoracle ladder + flow runTrustworthy, falsifiable verdict on whether a just-built flow works

WebView DOM & network (4)

ToolKey paramsBacking engineBehavior
webview_inspectudid, selector?, webviewId?, max?mobilecli (CDP)Resolves a CSS selector to DOM elements with absolute tapX/tapY
webview_evaludid, expression, webviewId?mobilecli (CDP)Runs JS in the page context; gated by PODIUM_DISABLE_WEBVIEW_EVAL=1
webview_navigateudid, action (goto/back/forward/reload), url?mobilecli (CDP)Drives WebView navigation
webview_networkudid, durationMs?, format (json/har)?, saveTo?, redact?, includeResources?CDP + in-page fetch/XHR shim + Resource TimingCaptures in-WebView HTTP traffic; exports redacted JSON or HAR 1.2

React Native debugging — Metro CDP (4)

ToolKey paramsBacking engineBehavior
metro_appsport? (8081)GET http://localhost:<port>/jsonDifferentiated errors (timeout vs not-running vs other)
metro_logswsUrl?/port?, durationMs?, maxLogs?WebSocket + CDP Runtime.enableAuto-discovers first app when URL omitted
metro_networkwsUrl?/port?, durationMs?, maxEntries?CDP Network.enableRequests (url/method/status/mimeType/ts)
metro_stateexpression?/wsUrl?/port?, timeoutMs?CDP Runtime.evaluateReads in-app state (default: globally-exposed Redux store)

Crash diagnostics (2)

ToolKey paramsBacking engineBehavior
crash_listprocessName?, sinceHours?, udid?host + sim DiagnosticReportsNewest-first; tagged source: host | simulator
crash_getid, udid?samePath-traversal-safe (basename only); truncates honestly

The oracle ladder — trustworthy assertions

"It works" is operationalized as a falsifiable, evidenced verdict — never "looks ok". Assertions and validate_flow resolve visibility through a three-rung ladder, using the strongest available signal:

  • WebView DOM — when an inspectable WKWebView is present, query the real DOM.
  • Native accessibility — the native AX element set (via idb/mobilecli).
  • MaestroassertVisible/assertNotVisible as the fallback.

assert_not_visible fails closed: if absence can't be positively verified (e.g. a WebView is unreadable), it reports failure rather than a false pass. Every verdict names the oracle that produced it, so an agent can weight its confidence.

Native-first gesture backend

Imperative gestures (tap_on, input_text, swipe, press_key, orientation_set, run_steps) and inspect_screen route through the fastest available backend, probed once and cached (with a short negative-cache TTL so a backend that starts after launch is picked up):

  • idb — when both idb and idb_companion are installed (native, fastest).
  • mobilecli — the bundled npm dependency (prebuilt Go binary). Default; no install.
  • Maestro fallback — when no native backend resolves, or for actions it can't express (double/long-press, UPSIDE_DOWN). The gesture generates a minimal flow with launchApp: { stopApp: false }, foregrounding the app without restarting so state is preserved.

Each result reports the backend it used. Set PODIUM_DISABLE_NATIVE=1 to force Maestro. Eliminating the per-gesture JVM spin-up cut tap_on ~14.7 s → ~0.6 s and inspect_screen ~8.9 s → ~0.9 s on an iPhone 16 Pro simulator. Run npm run benchmark for a full pass/fail sweep.

Model profiles (PODIUM_PROFILE)

Weak / free models (DeepSeek, GLM, MiniMax, MiMo) fail most on tool choice among ~60 tools. Set PODIUM_PROFILE=lite to expose only the intent-level essentials (≤ 8 tools: inspect_screen, run_intent, assert_visible/assert_text, tap_on, podium_health) so a weak model states its whole intent via run_intent and the server drives every step deterministically. Default full keeps all tools. See docs/model-profiles.md.

Maestro flakiness retry: when the fallback runs, its iOS driver intermittently fails with Failed to connect to 127.0.0.1:<port>. Flows retry up to 2× with 2 s / 5 s backoff and report the retries count; a persistent failure returns the raw output with remediation hints.

WebView & RN network introspection

Two distinct network layers, two tools:

  • metro_network captures requests on the RN/Hermes target via the CDP Network domain — the right tool for a native RN app's own fetch.
  • webview_network captures traffic inside a WKWebView: it injects a fetch/XHR recorder (rich — method/status/headers/body for calls after capture starts) and reads the browser's Performance Resource Timing buffer (includeResources, default on) — every request since navigation, including pre-capture ones (URL/timing/size). The merge yields a near-complete request list, exported as redacted JSON or HAR 1.2.

For an RN shell that hosts its UI in a WebView, the app's API calls run in the web layer — so metro_network sees nothing and webview_network is the tool to reach for. WebView tools require WKWebView.isInspectable = true (default in debug/staging builds; off in production); when none is found they return an actionable error.

Enabling an inspectable WebView

Every WebView-backed tool — webview_*, canvas_*, and rive_* (incl. the rive-e2e mobile runner) — drives the page through a CDP/JS bridge, which requires the host app to opt the WebView into inspection. This is an app-side setting; podium-mcp cannot flip it for you.

iOS (WKWebView, iOS 16.4+)

Apple gates WKWebView inspection behind an explicit flag, default false. Set it only in debug/staging builds:

import WebKit

let webView = WKWebView(frame: .zero, configuration: config)
#if DEBUG
if #available(iOS 16.4, *) {
    webView.isInspectable = true
}
#endif

Then, on the Mac the simulator/device is attached to:

  • Safari → Settings → Advanced → enable "Show features for web developers" (older Safari: Preferences → Advanced → "Show Develop menu in menu bar").
  • Safari menu bar → Develop → <Simulator or device name><page title> — opens the Web Inspector against that exact WebView, with full DOM/console/network access. The same <device> → <page> list is what confirms an app is actually inspectable before pointing rive_*/webview_*/canvas_* tools at it.

React Native (react-native-webview)

react-native-webview forwards the native flag — set it on the <WebView> in debug/staging builds only:

import { WebView } from "react-native-webview";

<WebView
  source={{ uri: "https://your-app.example" }}
  webviewDebuggingEnabled={__DEV__} // iOS: isInspectable · Android: setWebContentsDebuggingEnabled
/>

webviewDebuggingEnabled needs a recent react-native-webview (check your installed version's changelog — the prop was added once RN's WKWebView wrapper picked up isInspectable support); on older versions set the platform flags directly in native code as shown above/below.

The ios_webkit_debug_proxy (CDP) path

For CDP-based automation (what Podium's webview_*/canvas_*/rive_* tools use under the hood) rather than manual Web Inspector, ios_webkit_debug_proxy bridges a device/simulator's private WebKit remote-debugging protocol to standard CDP over a local TCP socket:

brew install ios-webkit-debug-proxy
ios_webkit_debug_proxy -c null:9221   # auto-discovers attached devices/sims

It auto-discovers every attached, inspectable device/simulator and exposes each one's inspectable WebViews as a CDP target on http://localhost:9221/json (one port per proxy instance; add -c udid:port pairs for multiple devices). Podium's own WebView resolution (resolveWebview) talks this same protocol — ios_webkit_debug_proxy is the tool to reach for when you want to confirm what's inspectable before pointing an MCP tool at it, or to drive CDP directly outside of Podium.

Android WebView

if (BuildConfig.DEBUG) {
    WebView.setWebContentsDebuggingEnabled(true);
}

Then on the host machine: open Chrome → chrome://inspect#devices → the device's WebView(s) appear under "Remote Target" once adb sees the device and debugging is enabled — click "inspect" for full DevTools.

⚠️ Security — debug/staging only, never production

isInspectable = true / setWebContentsDebuggingEnabled(true) expose the WebView's full JS runtime, DOM, and network traffic to anyone who can attach a debugger to the process (physically, or via a compromised dev toolchain). Never ship this flag enabled in a production/App Store or Play Store build — gate it behind #if DEBUG / BuildConfig.DEBUG / __DEV__, exactly as shown above. All of webview_*, canvas_*, and rive_* — including the rive-e2e mobile runner — require this flag to be on to function at all; on a non-inspectable WebView they fail closed with an actionable error rather than falling back to vision.

Documented limits (by design, not bugs)

  • Canvas/WebGL needs a cooperating JS framework — the canvas brain automates Pixi/Konva/Fabric/Phaser/Three/Babylon UIs by selector when the app exposes its scene-graph root (validated live). A raw/custom WebGL canvas, an opaque/production build, or Unity without an AltTester / window.__podiumEngine bridge is not selector-addressable — fall back to tap_with_fallback with screenshot-derived coordinates, or instrument the build.
  • WebView tools are dev/QA only — production App Store builds typically set isInspectable = false; tools return an actionable error and fall back to coordinate taps.
  • WebView content-process memory is unreadable from the app sandbox (platform limit) — use indirect signals (memory warnings, process terminations).
  • Maestro text: matcher is full-string regex (IGNORE_CASE) — partial strings don't match; copy hierarchy text verbatim or anchor with .*.
  • Android requires adb on PATH — gestures / inspect / screenshot work once adb is present; when it's absent every Android path degrades to a structured "adb not found" result.
  • orientation_get is a screenshot-aspect heuristic when no native backend is present — iOS simulators expose no direct orientation query.
  • record_start/record_stop keep state in-process — serialize start → … → stop on one connection; one active recording per udid (a watchdog finalizes one that's never stopped).

Architecture

src/
  index.ts          # MCP server entry — registers every tool group, warms caches
  lib/
    exec.ts         # execFile-based runner (NO shell) + timeout/timedOut flag
    result.ts       # shared ok/error MCP content helpers
    simctl.ts       # xcrun simctl wrappers + device-list TTL cache
    native.ts       # gesture/inspect backend: idb → mobilecli → null (re-probe TTL)
    idb.ts          # idb gesture/inspect adapter
    gesture.ts      # unified native→Maestro executors (shared by screen + steps)
    oracle.ts       # the oracle ladder: WebView-DOM › a11y › Maestro
    maestro.ts      # Maestro engine: flow runner, idb retry, hierarchy
    export-maestro.ts # run_steps → reusable Maestro flow
    har.ts          # HAR 1.2 export for webview_network
    webview.ts      # mobilecli CDP — WebView list/inspect/eval/navigate/network
    metro.ts        # Metro CDP — app discovery, logs, network, state
    crash.ts        # DiagnosticReports crash listing/reading
    recording.ts    # detached screen recording lifecycle + watchdog (platform-aware)
    device-target.ts # DeviceTarget model + PlatformDriver registry (v0.3.0)
    drivers/        # per-platform lifecycle: ios-sim, android, ios-real
    adb.ts          # Android adb driver (list/install/launch/screenshot/wm size)
    adb-backend.ts  # adb gesture/inspect (input + uiautomator → AX elements)
    iosreal.ts      # real iOS via devicectl (list/install/launch) + capture
    wda.ts          # opt-in WebDriverAgent backend (/source + tap/swipe/keys)
    engine.ts       # no-vision engine client (AltTester + WebGL-in-WebView)
    engine-transport.ts # WebSocket transport for the AltTester bridge
    canvas-types.ts # Canvas Brain shared contract (CanvasObject, selectors)
    canvas-adapters.ts  # in-page bridge: detect + walk Pixi/Konva/Fabric/Phaser/Three/Babylon
    canvas-resolver.ts  # semantic "close brain": intent → ranked, evidenced target
    canvas-a11y.ts  # Flutter/ARIA fallback reader → CanvasObject (scaffolding, not wired — #9)
    canvas-vision.ts # opt-in vision fallback scaffolding (not wired — #9)
    token-report.ts # token estimators + no-vision vs vision-loop comparison
  tools/            # one file per group:
                    #   health, device, screen, steps, flow, assert, validate,
                    #   webview, debug, engine, canvas, token
assets/             # bundled offline Maestro cheat sheet + demo.gif
scripts/            # benchmark.ts, compare-mcps.ts, token-bench.mjs
e2e/                # smoke suites (smoke / full-smoke / webview-network-live / android-smoke / engine-smoke)
test/canvas-e2e/    # live Playwright-WebKit canvas bridge suite (6 frameworks)
docs/               # tool catalog, e2e transcript, roadmap, token-economics

Development & testing

npm run build       # tsc
npm run typecheck   # tsc --noEmit
npm test            # vitest run — 1048 unit/integration tests (exec/network mocked, no sim needed; no browser launch)
npm run test:canvas # live canvas + Rive-web-runner suite in Playwright WebKit — 21 tests (run `npx playwright install webkit` first)
npm run benchmark   # spawn a fresh server over stdio and sweep the tool suite
node e2e/smoke.e2e.mjs        # real E2E against a booted simulator (macOS + Xcode)
node e2e/full-smoke.e2e.mjs   # drives the iOS-sim tool handlers (happy + structured-error paths)
node e2e/android-smoke.e2e.mjs # Android emulator/device smoke (story A3)
node e2e/engine-smoke.e2e.mjs  # AltTester engine smoke; skips without an instrumented build (story C4)

1048 unit/integration tests across 81 files (fully mocked, no browser launch), plus 21 live canvas-bridge/Rive-web-runner tests in Playwright WebKit (1069 total), all passing — including the Rive E2E kit's device-free core (contract/coverage/smoke/drift, plus the web runner's live closed-loop check against a real WebKit bridge in test:canvas), the v0.3.0 device-target registry, the Android adb driver + uiautomator parser, the AltTester engine client + WebGL bridge, the devicectl/WDA real-iOS parsers, the Rive dual-source analyzer (live bridge + static WASM), plus the v0.2.0 oracle ladder, recording watchdog, gesture-parity, HAR export, WebView, and Metro paths.

Standards: TypeScript strict, no as any / @ts-ignore, no shell execution (all commands via lib/exec.ts), tools return structured errors instead of throwing. See CONTRIBUTING.md for the "add a new tool" checklist.

E2E on CI: the E2E (simulator) workflow boots a real iOS simulator on a macOS runner and runs the smoke suites nightly + on demand (not a PR gate — simulator runs are slow). full-smoke.e2e.mjs asserts the happy path where a target exists and the real structured-error path where a dependency is absent (a debug isInspectable app for WebView; a connected RN app for metro_*).

Roadmap & contributing

podium-mcp is production-ready for iOS/Android UI automation and no-vision canvas/WebGL (Pixi/Konva/Fabric/Phaser/Three/Babylon — validated live). The frontier, where a contributor can make a real dent, lives in open issues:

High-impacthelp wanted

  • #1 — validate the AltTester/Unity engine path against a live instrumented Unity build (the biggest gap to real Unity automation).
  • #2real-device WKWebView e2e for the canvas brain (today validated in Playwright WebKit).
  • #3Unity-WebGL adapter: auto-detect + a drop-in window.__podiumEngine bridge.

Good first issuesgood first issue

  • #4 — more canvas adapters (PlayCanvas, Cocos Creator, p5.js).
  • #5 — expose canvas_hittest / canvas_object_rect tools.
  • #7 — exact token counts via the Anthropic count_tokens API.
  • #6 — address Konva Group/Container targets.

Adding a tool follows one checklist in CONTRIBUTING.md: TypeScript strict, no shell, structured-errors-never-throw, a vitest test, and a row in the tool catalog. PRs welcome.

Releasing

server.json is the official MCP Registry manifest. Pushing a v* tag runs Publish to npm then Publish to MCP Registry (GitHub OIDC for the io.github.hoainho/* namespace — no long-lived token). Both workflows run typecheck → build → test as a gate first; the registry publish only succeeds once the matching npm version is live, and versions are immutable.

Prompt playbook & references

Rive Analyzer — reference site

../rive-analyzer is a small standalone website that consumes this package's browser-safe podium-mcp/analyzer subpath (mergeRiveReport + adviseRive). Drop a .riv file in and it renders the same inventory + scored A–F best-practice advisor the rive_analyze MCP tool returns, entirely client-side (no server, no upload).

It exists as a live consumer/proof of the podium-mcp/analyzer export map — the same reason rive-analyzer/src/analyzer-bridge.ts imports only mergeRiveReport/adviseRive/types and never the Node-only analyzeRiveFile, so a browser bundler is forced to tree-shake the two apart. It's a companion to, not a replacement for, the older nano-step/rive-playground — that tool already covers broad .riv parsing and codegen; this site's only differentiator is the scored advisor. See rive-analyzer/README.md for the honest v1 boundaries (live-instance-only, no static source yet) and deploy status.

Design ideas

  • One podium, one connection. A single server fronts every mobile capability so an agent configures one endpoint and discovers all 60 tools at once.
  • Safe by construction. Every external command runs through an execFile layer with an explicit argument array — never a shell string.
  • Never crash the conductor. Tools return structured results and errors instead of throwing; one bad call can't take the server down.
  • Degrade, don't fail. A missing toolchain (e.g. Android's adb) yields an informative result rather than a hard error.
  • Prove it, don't guess. Assertions return evidenced verdicts via the oracle ladder and fail closed when they can't verify.

Contributing

Contributions welcome — see CONTRIBUTING.md and the Code of Conduct. Use the issue templates for bugs and feature requests.

Security

Please report vulnerabilities privately per SECURITY.md — do not open a public issue. SECURITY.md also documents the webview_eval / run_flow trust boundary and the PII-in-transcript caveat.

License

MIT © 2026 hoainho

Keywords

mcp

FAQs

Package last updated on 13 Jul 2026

Did you know?

Socket

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

Related posts