Why tabnas #
tabnas runs a grammar as a table of rules and token alternates. Plugins extend those tables to add syntax to an existing language. You can define a table directly or compile an ABNF grammar into it.
The same representation supports inspection and debugging: print a grammar as ABNF, draw its rules, or inspect the parser state. Agents can emit a declarative grammar, validate it, and test sample inputs before using it.
The parsing algorithm#
The engine loops over tokens and uses an explicit stack to track rule depth. Each rule has open and close states, with an ordered list of alternates for each state. An alternate matches tokens and can run a condition or action, push a child rule, or repeat a rule.
Actions build the value returned by a parse. The rule table controls which actions run and when. See how it works for the algorithm and ABNF grammars for the compiler’s handling of left recursion.
Project history#
Richard Rodger began with a more permissive JSON parser, jsonic, built using PEG.js. The format supported configuration in the Seneca microservices framework, including unquoted keys and comments.
A rewrite replaced generated parser code with recursive descent, then with token lookup tables and an explicit stack. Entering an object or array pushes a rule instance; completing it pops that instance.
Removing entries from the tables produced a strict JSON parser. Adding entries extended it to jsonic. That relationship became the plugin model: start with a grammar, then add or modify rules for another format.
How tabnas works #
To parse {"foo":1}, first split the source into tokens. This is
lexing. tabnas tries registered lexers in order; you can also register
custom lexers for your language.
Lexing produces a token stream:
{"foo":1}becomes:
{ foo : 1 }
OB KEY CN VAL CBwhere each token has a code name, such as OB, “Open Brace”, for {.
Many parsing systems mix lexing (clumps of characters) and parsing (semantic units), forcing both into the grammar with equal status. For example ABNF format, used in most Internet RFCs, often looks like:
FWS = ([*WSP CRLF] 1*WSP) / obs-FWS
; Folding white space
ctext = %d33-39 / ; Printable US-ASCII
%d42-91 / ; characters not including
%d93-126 / ; "(", ")", or "\"
obs-ctext
ccontent = ctext / quoted-pair / comment
comment = "(" *([FWS] ccontent) [FWS] ")"
CFWS = (1*([FWS] comment) [FWS]) / FWSThis is from RFC 5322, section 3.2.2
tabnas separates lexing from parsing. Built-in lexers, custom functions, and regular expressions recognise tokens. Grammar rules then operate on those tokens.
Your first small language#
Build a grammar that adds numbers:
1+2+3 # and the result should be 6The builtin lexer already handles numbers and fixed strings, so you can assume these tokens:
NR - number
PL - plus characterThat means the parser has to handle this series of tokens:
NR PL NR PL NREach number NR adds to the running total. A plus character after a
number means the parse keeps going; otherwise it is done.
If you’re familiar with ABNF (Augmented Backus-Naur Form) from Internet RFCs, here is the grammar:
val = add
add = NR [ PL add ]
NR = <number>
PL = "+"The val rule contains a child rule, add. That rule matches a number
NR, optionally followed by a plus token PL and another add.
The ABNF plugin can compile this grammar. The next steps define its rule table directly in TypeScript so you can see how the engine executes it. The Go API supports the same grammar.
Create a tabnas parser and define its grammar. You can do
this with declarative JSON (apart from actions). First, set some
options. Define a custom token PL that represents the +
character, and start with the grammar rule val. That means the top level of the grammar tree is a val,
representing the final value of the additions.
const { Tabnas } = require('@tabnas/parser')
// Create a new parser.
const tn = new Tabnas()
// Define the grammar.
tn.grammar({
options: {
// Define a new token named #PL, a "+" character.
fixed: { token: { '#PL': '+' } },
// Start parsing at the 'val' rule.
rule: { start: 'val' },
},
...Now define the rules. Start with val, which sets up the
accumulator for the addition, and expects to have add as a child rule:
...
rule: {
// The 'val' rule holds the running total.
// Each rule instance has a 'node' representing its value.
val: {
// Define the "opening" phase of the rule.
open: [
// This is an "alternate", it matches any tokens.
{
// "push" down into an 'add' rule.
p: 'add',
// An "action" - set the counter to 0.
a: (r) => { r.node = 0 }
}
],
... The open array is a list of alternates. The parser tries to match
each one’s tokens and then performs its actions. Here the parse is just
starting, in the “open” state of the val rule, so the alternate looks
at no tokens and immediately “pushes” the add rule onto the rule
stack. That means the add rule is the next rule to run.
The action at this point sets the result value to 0, since
nothing has been added yet. Every rule gets an instance, and each rule
instance has a node value, where you store the value of that rule.
Add each number to the total:
...
// The 'add' rule performs the addition.
add: {
open: [
{
// Match a number - #NR is a built-in token for numbers.
s: '#NR',
// Add the number to the total.
a: (r) => {
r.parent.node += // The parent is the 'val'.
r.o[0].val // Get the value of the first opening token.
}
}
],
...Most alternates try to match a sequence of tokens. In this case the
first alternate of the add rule looks for NR, a number. Token
names are prefixed with # to make them easier to see.
If a number matches, the action adds it to the running total set up
in the val rule. It takes the parent’s node (r.parent.node) and
adds the value of the first matched token (r.o[0].val). The rule
is in the “open” state, so r.o holds the sequence of matched tokens. The val field of a token holds its
evaluated value. For NR, this is a number.
The rule has no children to push. It transitions from open to close:
...
add:
...
close: [
// If there is a "+" following the number, keep going.
{
s: '#PL', // This is our "+" token, #PL
r: 'add' // "Repeat" the 'add' rule
},
// Else end the rule.
{}
]
...The first alternate tries to match a plus token (PL). If it
does match, the parser “repeats” the rule: it runs the add rule again. A new
instance starts in the “open” state, looking for another number to add.
If the first alternate does not match, the parser reaches the second
alternate, which does nothing. That closes the rule and heads back up
to the val rule. The empty alternate is needed, because when no
alternate matches, that’s a parse error.
The parse is now back up the stack, closing the val rule:
...
val:
...
// Define the "closing" phase of the rule.
close: [
{} // Ending "alternate" - does nothing.
]
...Again there is nothing to do, so the parse completes. The grammar
has accepted the input, and the parser returns the value of the top
level node:
tn.parse('1+2+3') // => 6You can define grammars at three levels: ABNF, a declarative rule table, or the programmatic API. Use the API for parameterised grammars, as @tabnas/directive and @tabnas/expr do.
Choosing a grammar representation#
ABNF provides a notation for grammar structure. A rule table exposes the alternates and actions directly. The programmatic API constructs those same rules from code. All three use the same parsing engine.
Extending an existing plugin means maintaining its interaction with the rules you add. Validate the combined grammar and test valid inputs, invalid inputs, and any syntax whose meaning changes. See test a grammar for that workflow.
Related reading:
- Quickstart: the addition grammar above, in full.
- How it works: rules, alternates, and the stack.
- ABNF grammars: the dialect, and left recursion.
- Other parsers: compare grammar formats, algorithms, and tooling.
- FAQ: what it does, what it won’t do, and why.
- Playground: edit a grammar in the browser.