
Product
Introducing Socket Scanning for VS Code Marketplace Extensions
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.
@ashlr/lexicon
Advanced tools
Personal lexicon for voice-to-agents. Fixes the words STT gets wrong (Ashler -> Ashlr.AI) before your agent sees the prompt.
A personal lexicon for voice-to-agents. One YAML file of the words speech-to-text gets wrong, applied everywhere your voice lands: MCP, Claude Code, the browser, macOS. The package is @ashlr/lexicon; the command is lexicon.
![]()
lexicon.ashlr.ai is the live demo, the benchmarks and the install commands.
You said: "tell Ashlr.AI to deploy the Kubernetes auth service"
STT heard: "tell Ashler to deploy the Cooper Nettie's off service"
Agent received: "tell Ashlr.AI to deploy the Kubernetes auth service"
Try it in your browser. No install: the page runs this repo's real matcher on your text, in your browser. (The Dictate button uses your browser's own speech recognizer, which in Chrome sends audio to Google.)
Method and full tables are in docs/BENCHMARK.md. Reproduce with npm run bench:audio.
| corpus | proper nouns recovered, raw STT | after lexicon | clean prose wrongly changed |
|---|---|---|---|
| real audio, whisper.cpp base.en (330 clips) | 41.9% | 86.4% | 0 of 72 |
| real audio, whisper.cpp small.en with prompt hints | 76.0% | 95.7% | 0 of 72 |
| synthetic STT errors (398 sentences, 70 terms) | 5.1% | 96.5% | 0 of 95 |
Latency is about 0.3 ms per sentence. The real-audio rows use macOS text-to-speech read into whisper.cpp, so they are cleaner than a phone microphone.
The last column counts ordinary prose only. Each corpus also contains sentences deliberately built to trip the matcher (a bare "llama" next to an Ollama term, sound-alikes, code spans), marked expected-hard; with those included the false-positive rate is 15% (18 of 120) synthetic and 20% (18 of 90) on audio. Both numbers, and every failing case, are in docs/BENCHMARK.md.
curl -fsSL https://ashlrai.github.io/lexicon/install.sh | sh # CLI + the setup wizard
brew install ashlrai/tap/lexicon # or Homebrew (macOS, Linux)
npm i -g @ashlr/lexicon # or npm (Node 20+)
Then open Claude Code and say a sentence with your company name in it. Done.
The install script runs lexicon setup for you (LEXICON_NO_SETUP=1 skips it); after a Homebrew or npm install, run it yourself. It is six steps: seed the lexicon with your name and company, install starter packs, harvest the current repo, register the MCP server and hooks in every agent client it detects, install the local API as a login service, and export to your dictation app. Every step is optional and safe to rerun, and lexicon setup --dry-run prints the whole plan without writing anything. The walkthrough is in docs/QUICKSTART.md.
Or skip the wizard and add one term by hand. The first argument is the canonical spelling, the rest are what STT actually produces:
lexicon add Ashlr.AI Ashler Ashlar "Ashler AI" --phonetic ASH-ler
lexicon normalize "tell Ashler to ship it"
# tell Ashlr.AI to ship it
Claude Code plugin, if you would rather not install a CLI at all. No Node install step, no build:
claude plugin marketplace add ashlrai/lexicon
claude plugin install lexicon@ashlrai
lexicon doctor checks the install. There is no telemetry and all state is local files: the CLI, hooks, MCP server, local API and extension make no request beyond loopback. The one outbound request in the codebase is lexicon voice fetching a whisper model on first use. The install script, npm and Homebrew fetch the package itself. See SECURITY.md.
Speech-to-text is about 95% accurate on ordinary English and much worse on invented names. In the benchmark above, raw whisper.cpp base.en transcribed 117 of 279 dictated proper nouns correctly. "Ashlr.AI" becomes "Ashler", "Kubernetes" becomes "Cooper Nettie's", "SaaS" becomes "sauce", "auth" becomes "off". Those are exactly the words an agent needs to get right.
Dictation apps (Wispr Flow, Superwhisper, Aqua) each keep their own dictionary and none of them share it. Agents (Claude Code /voice, ChatGPT voice, Codex, local Whisper) run their own recognizer with no user vocabulary at all. This is the portable layer in between: corrections happen after STT and before the model, wherever the text passes through.
This is not a dictation app. It sits between whatever dictation you already use and whatever agent you talk to. The research behind that call, including the kill criteria, is in docs/RESEARCH.md.
setup_lexicon, lexicon_doctor, install_client, trust_project, import_dictionary and suggest_terms mean "set up my lexicon" works without a terminal. The tools that change your machine preview first: setup_lexicon and install_client return a plan and write nothing until the agent passes apply: true, trust_project shows the file's terms before pinning it, and import_dictionary takes dryRun.SessionStart and UserPromptSubmit hooks, a lexicon skill and a /lexicon command. Installs from this repo's marketplace with no build step.lexicon add to lexicon voice.normalize() is a pure function: text plus lexicon in, corrected text and a replacement list out.| Surface | How | Docs |
|---|---|---|
| Claude Code | Plugin, or MCP server plus two hooks that correct the prompt before the model reads it | CLIENTS.md |
| Codex, Cursor, Windsurf, Gemini CLI, VS Code, Claude Desktop | lexicon install <client> --apply registers the MCP server | CLIENTS.md |
| Any MCP client | stdio server, nineteen tools | MCP.md |
| ChatGPT, Claude.ai, Grok, Gemini, Perplexity, Poe, Copilot | Browser extension: rewrites the composer when you press send | EXTENSION.md |
| Any macOS app, any dictation tool | LexiconBar menu bar app: rewrites dictated text in the focused field through Accessibility, with an undo bubble | MACOS-APP.md |
| Shortcuts, Raycast, scripts, your own app | lexicon serve: loopback HTTP API on 127.0.0.1:41733 behind a bearer token | LOCAL-API.md |
| Dictation without a dictation app | lexicon voice: ffmpeg records, whisper.cpp transcribes with your canonicals as prompt hints, the lexicon corrects | VOICE.md |
| Any text field, any OS | lexicon daemon --once --paste on a hotkey | DAEMON.md |
| Wispr Flow, Superwhisper, macOS Text Replacement, espanso, Deepgram, Azure, Google | Export into their own dictionaries and biasing parameters | EXPORTS.md |
| Your own STT pipeline | npm i @ashlr/lexicon, call normalize() between transcription and the model | LIBRARY.md |
Three tiers over token windows: exact alias first, then double-metaphone phonetic, then Damerau-Levenshtein fuzzy above a confidence floor. Exact hits win the span; matches never overlap. A stoplist of about 3400 common English words, per-term never lists, and (with the default skipCode) code spans, URLs, emails, paths and glued identifiers are all off limits. That is why zero clean sentences changed in the benchmark. Every replacement reports its reason and confidence.
lexicon normalize --diff "deploy to head sner with cooper netties"
# stderr: "head sner" -> "Hetzner" (alias, 1.00)
# "cooper netties" -> "Kubernetes" (phonetic, 0.85)
# stdout: deploy to Hetzner with Kubernetes
The rules in full, including every guard, are in docs/MATCHING.md.
Start here
| Page | What it covers |
|---|---|
| QUICKSTART.md | Five minutes from nothing to corrections in Claude Code, with what each setup step writes |
| CLIENTS.md | Installing into Claude Code (plugin, hooks, headless) and every other agent client |
| PACKS.md | The four starter packs, how install and remove behave, how the aliases were chosen |
Reference
| Page | What it covers |
|---|---|
| CLI.md | Every command and flag, generated from --help |
| MCP.md | The MCP server: nineteen tools, two resources, two prompts |
| LEXICON-FILE.md | File locations, the term schema, settings, never |
| MATCHING.md | The three matching tiers and every guard against a false positive |
| EXPORTS.md | Fifteen export formats and seven importers |
| LIBRARY.md | Using normalize() and the store functions from your own code |
| TRUST.md | Why a project .lexicon.yaml is off until you approve it |
Surfaces
| Page | What it covers |
|---|---|
| EXTENSION.md | The browser extension for ChatGPT, Claude, Grok, Gemini, Perplexity, Poe and Copilot |
| MACOS-APP.md | LexiconBar, the macOS menu bar app and its Accessibility rewrite |
| LOCAL-API.md | The loopback HTTP API, its routes and its token |
| VOICE.md | Local push-to-talk with ffmpeg and whisper.cpp |
| DAEMON.md | The clipboard daemon and hotkey recipes for macOS, Linux and Windows |
Growing and measuring
| Page | What it covers |
|---|---|
| GROWING.md | Harvesting a repo, learning from corrections, stats, reviewing terms |
| SUGGEST.md | What lexicon suggest mines from your voice history, and how it scores |
| BENCHMARK.md | The accuracy benchmark: corpora, metrics, results and the fix log |
| RESEARCH.md | Why this layer exists, the market read, and the kill criteria |
Internals
| Page | What it covers |
|---|---|
| ARCHITECTURE.md | Module map and the design decisions behind it |
| CONTRACT.md | The per-module API contract every change is written against |
| AGENT-NATIVE.md | The agent-as-UI design: which tool an agent calls when |
| DOGFOOD.md, DOGFOOD-AGENT-NATIVE.md | Two live runs against the real Claude Code CLI, and the bugs they found |
| RELEASING.md | Cutting a release: npm, GitHub assets, the Homebrew bump |
| LANDING.md | The landing page at lexicon.ashlr.ai: what it claims, and how to deploy it |
Also at the root: CONTRIBUTING.md, SECURITY.md, CODE_OF_CONDUCT.md, CHANGELOG.md.
Every GitHub release attaches the browser extension for Chrome/Edge/Brave and for Firefox, LexiconBar.app.zip for macOS, the npm tarball for offline installs, and SHA256SUMS. The Homebrew formula lives in ashlrai/homebrew-tap; npm i -g github:ashlrai/lexicon#v0.4.0 installs a tag straight from GitHub and builds on install.
Non-goals: this is not a dictation app, and there are no hosted accounts and no sync service. It is a file.
Good first issues are labelled and scoped: a new starter pack, an exporter, an importer, a harvester source. CONTRIBUTING.md has the setup, the test layout and a recipe for each.
Found a name it gets wrong? Open a misheard term issue.
MIT. Copyright 2026 Ashlr.AI.
FAQs
Personal lexicon for voice-to-agents. Fixes the words STT gets wrong (Ashler -> Ashlr.AI) before your agent sees the prompt.
The npm package @ashlr/lexicon receives a total of 67 weekly downloads. As such, @ashlr/lexicon popularity was classified as not popular.
We found that @ashlr/lexicon demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Product
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.

Research
/Security News
Socket uncovered two malicious VS Code themes in a GlassWorm-linked cluster with thousands of installs across VS Code Marketplace and Open VSX.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.