tabnas and other parsers

Compare the grammar representation, parsing algorithm, and tooling each project provides. Follow the project links for its own documentation and supported features.

Tool Kind Algorithm Grammar lives in Code generation Extending someone else's grammar
tabnas Runtime rule machine Deterministic LL, no backtracking Data: ABNF, JSON, or API None Add rules to an existing grammar
ANTLR Parser generator ALL(*) .g4 file Generates source Edit the grammar, regenerate
Peggy (PEG.js) Parser generator PEG, ordered choice + backtracking .peggy file Generates source Fork the grammar
Chevrotain Parser toolkit LL(k), configurable lookahead TypeScript methods None Subclass the parser
nearley Parser generator Earley: any CFG, ambiguity allowed .ne file Generates source Edit the grammar, recompile
tree-sitter Parser generator GLR, incremental grammar.js Generates C inherits / injection, then regenerate

Grammar representation

ANTLR and Peggy generate parser code from grammar definitions. tabnas executes a rule table at runtime, so plugins can change the grammar without generating parser source.

tabnas keeps the grammar as data the engine walks at runtime. That makes extension the normal case rather than a fork: you add rules and token alternates to a grammar that already exists, which is how jsonc is built on json, and jsonic on that. It also makes the grammar inspectable while it runs: you can print it back as ABNF or draw it. Why tabnas has the longer version.

Chevrotain also runs without code generation. Its grammar uses a JavaScript DSL; tabnas exposes a rule table that can be defined as data or through an API.

An agent can produce ABNF or a declarative tabnas rule table. Neither representation establishes that the grammar is correct: validate it and test valid and invalid inputs. The agent guide describes those checks.

When to use something else

Your grammar is ambiguous
tabnas is deterministic and does not backtrack. If you need every valid parse of an ambiguous input, nearley's Earley parser is built for that; PEG's ordered choice with backtracking handles a different set of awkward cases.
You're building editor tooling
tree-sitter reparses incrementally on every keystroke and recovers from errors mid-file. tabnas does neither, and isn't trying to.
You need a mature grammar for a big language
ANTLR's grammars-v4 has hundreds of battle-tested grammars. tabnas has 35 packages and one maintainer.
You're working in a language the engine doesn't run in
The engine is implemented in TypeScript and Go. ANTLR generates for Java, C#, Python, Go, JavaScript, C++, Swift, and more.
You just need to parse JSON
Use your platform's JSON parser. Really.

What tabnas asks of you

  • You write the lexer. tabnas draws a hard line between lexing and parsing and is mostly about the parsing half. It gives you tools for the lexing part, but you assemble it.
  • Left recursion is a rewrite, not native support. The ABNF compiler rewrites left-recursive rules (Paull's algorithm), so the resulting tree is flat rather than left-nested, and a purely left-recursive rule with no base branch is an error.
  • It's pre-1.0. Minor versions can break things, and grammar packages track the engine release they were built against.

If something on this page is wrong or unfair to one of these projects, please say so: getting a comparison wrong is worse than not having one.