Parséman (PAR-zə-mahn)
Write parsers as TypeScript functions. Ship them like hand-written parsers.
Parséman is a TypeScript parser-combinator library with an optional compiler/macro path that turns your grammar into flat JavaScript. Use the same grammar interpreted in tests and REPLs, macro-compiled at build time in production, or compiled on demand at runtime with compile().
Use Parséman when you want:
- normal TypeScript instead of grammar files
- parser-combinator ergonomics without parser-combinator slowness
- CST/AST nodes with spans and trivia
- error recovery for editor tooling
- incremental re-parsing
- fast parsers for DSLs, config languages, formatters, linters, and language servers
📖 Full documentation: matthew-dean.github.io/parseman
Why Parséman?
Most parser tools make you choose between ergonomics and performance.
Parser combinators are pleasant to write, but often slow. Parser generators can be fast, but usually involve grammar files, generated code, and extra tooling. Hand-written parsers are fast, but expensive to design and maintain.
Parséman aims for the useful middle: write your parser as ordinary TypeScript, then compile it into code that behaves more like a hand-written parser.
Install
npm install parseman
Parseman is pre-1.0. Minor versions may include breaking changes; check the
changelog before upgrading.
Quick start
import { literal, sequence, choice, regex, transform, parse } from 'parseman'
const method = choice(literal('GET'), literal('POST'), literal('PUT'), literal('DELETE'))
const target = regex(/[^\s]+/)
const version = regex(/1\.[01]/)
const requestLine = transform(
sequence(method, literal(' '), target, literal(' HTTP/'), version),
([verb, , path, , ver]) => ({ verb, path, version: `HTTP/${ver}` })
)
parse(requestLine, 'GET /api/v1 HTTP/1.1')
Three modes, one grammar
The same combinator code runs three ways, with identical results:
- Interpreter — zero setup, works anywhere (tests, REPLs, dynamic grammars).
- Macro build — a bundler plugin evaluates your grammar at build time and replaces it with inline JS. Zero runtime cost; the
parseman import disappears from the bundle.
compile() — the same optimizer, run on demand at runtime.
import { literal, sequence, choice } from 'parseman' with { type: 'macro' }
See The three modes for the full story.
What's in the box
- Combinators —
literal, regex, sequence, choice, many, sepBy, token, not, and more.
- Whitespace & trivia — grammar-defined filler skipping, with per-chunk kind capture.
- Recursive rules —
rules() for mutually recursive grammars; fully macro-compilable.
- CST / AST nodes —
node() captures terminals, named field() values, and trivia for you, with unwrap for AST/value wrappers, collapse for grammar-local CST wrappers, and cstBuildHost({ collapse }) for public CST policies.
- Incremental re-parsing —
parseDoc re-parses just the edited subtree on each keystroke.
- Error recovery —
recover, expect, and a { recover: true } channel keep parsing broken input and report every error.
- Context-sensitive parsing —
withCtx / guard without mutating shared state.
Full API in the reference.
Compared to other parser tools
Wondering how Parséman compares to Peggy, Chevrotain, Lezer, tree-sitter, Parsimmon, Nearley, or hand-written parsers?
See the full comparison: How Parséman compares
Real grammar example: GraphQL
Parséman includes a GraphQL grammar used in the benchmark suite. It parses executable GraphQL documents (queries, mutations, subscriptions, fragments, directives, variables, all value types) into typed AST nodes — not just syntax-validating them.
This is a real-world example of Parséman on a non-trivial, spec-shaped language, not a toy grammar.
Benchmarks
Parséman includes benchmarks against several JavaScript/TypeScript parser libraries across JSON, CSV, and GraphQL fixtures. Benchmarks are not universal truth tablets — results depend on grammar shape, input size, runtime, and what each parser is asked to produce. The benchmark suite is included so results can be inspected and reproduced (see Reproducing the numbers).
When parsing to JS values — objects, row arrays, AST nodes — Parséman's macro build is the fastest general-purpose JS parser we benchmark, beating Peggy, Parsimmon, Chevrotain, Nearley, and Jison at every grammar and size. The only thing that edges it out is a purpose-built native like JSON.parse; for anything that doesn't have a built-in, Parséman is the one to beat.
For syntax tree building, the compiled CST path (macro build) beats Lezer on the JSON CST fixture too — while producing a richer object tree with spans and trivia. For incremental re-parse, Parséman's parseDoc stores parent-relative spans so in-place value edits are ~110× faster than a full reparse and ~20× ahead of Lezer; structural edits (inserting/removing a whole list element) reuse the collection's untouched tail and land within a few × of Lezer's buffer reuse rather than at full-reparse cost. Full breakdown in the benchmarks guide.
Measured on Apple M4 Pro. Bars show µs per parse — shorter is faster. Refresh: pnpm bench:svg (benchmarks chart parsers and updates assets/bench-*.svg).
Compared parsers: Parséman, Peggy, Parsimmon, Chevrotain, Nearley, and Jison (plus JSON.parse on JSON). Each implements the same parsing work on the bench fixtures — building JS values / row arrays / GraphQL AST nodes, not syntax-only validation. Peggy grammars in bench/*.pegjs are the reference; Nearley JSON uses kach/nearley examples/json.ne; other Nearley and Jison grammars are ports of those Peggy files (bench/vendor/).




Parséman has three modes — interpreter (zero setup, works anywhere), macro build (compiled by the bundler plugin at build time, zero runtime cost), and compile() (optional runtime JIT). Most production use lands on one of the first two. The initialization section only shows parsers with a nonzero setup cost: compile() costs 75–650 µs depending on grammar size; Chevrotain always costs 840–1,400 µs. Parsers not listed there start for free. (Init numbers are pinned on the charts — they're noisy run-to-run; warm-parse bars are the meaningful comparison.)
On JSON, CSV, and GraphQL, Parséman macro beats every other library at every fixture size in the charts above — e.g. GraphQL large 142 µs vs Peggy 339 µs, JSON large 141 µs vs Peggy 466 µs / Chevrotain 250 µs. Native JSON.parse is the one thing faster on JSON (54.6 µs large). Even the zero-setup interpreter is highly competitive: on JSON and CSV it's the fastest option after the macro build — ahead of Peggy and every other generator — with Peggy edging it out only on GraphQL. On the CST chart, macro build beats Lezer at every size (172 µs vs 619 µs large, parse-only). Full write-up and how to refresh the charts: benchmarks guide. Grammar-level speed levers: performance guide; library-level codegen: PERF_IDEAS.md.
Developing
pnpm install
pnpm test
pnpm typecheck
pnpm build
pnpm bench
pnpm bench:parseman
pnpm bench:svg
pnpm bench:compile-grammars
pnpm docs:dev
License
MIT © Matthew Dean