PEGO

Generating TypeScript Parsers with pego gen -lang ts

pego gen -lang ts turns a grammar into a TypeScript module: a parser you commit next to your code and import like any other module. The module has no dependencies, not even on Node.js: it runs in Node.js, Deno, Bun and browsers, and it returns the same trees and the same errors as the engine and the generated Go parsers.

This guide shows how to generate a parser, how to call it, how it maps the engine's Go types to TypeScript, what is different in JavaScript (strings, ints, the stack), and how to keep it up to date. It assumes you know the Go generator's guide, code generation, only for the comparisons.

Why generate TypeScript

Use it when a grammar has to run where Go does not: in a browser (a syntax checker, an editor, a playground), in a Node.js tool, or in a TypeScript service that shares a format with a Go one. Since both parsers come from the same .pego file and return identical results, the two sides cannot drift apart: a tree serialized on one side is the tree the other side would have built.

Engine (Go) Generated Go Generated TypeScript
Runs in Go programs Go programs Node.js, Deno, Bun, browsers
Dependencies github.com/ornew/pego Standard library None
Trees and errors Reference Identical Identical (also as JSON, byte for byte)
Recognition (no tree) RecognizeOnly Recognize recognize, with -recognize
Typed values (-types) — Yes No
Streaming, incremental, backends Yes No No

Quick start

The grammar, pairs.pego, is the one of the Go guide:

type Pair struct { Key Match, Value Match }

def main = ws ps:pair+ $$ -> $ps
def pair: Pair = k:key "=" v:value ";"? ws -> new Pair{Key: $k, Value: $v}
def key = @(?a-zあ-ん)+
def value = @(?0-9)+
def ws = (? \t\n)*

Generate the module (pego is installed with go install github.com/ornew/pego/cmd/pego@latest, or run as go run github.com/ornew/pego/cmd/pego):

pego gen -lang ts -g pairs.pego -o pairs.ts

The result is pairs.ts, 2,761 lines and 79,670 bytes for this grammar, starting with // Code generated by pego. DO NOT EDIT.. Use it:

// main.ts
import { parse, parseRule, marshal, Bytes, Node, SyntaxError } from "./pairs.ts";

const { node, error } = parse("abc=12; あい=3");
if (error !== null) {
  throw error;
}
console.log(String(node));
for (const pair of node!.children) {
  const key = pair!.field("Key") as Node;
  console.log(pair!.type, JSON.stringify(key.text), "at", key.start + "-" + key.end);
}

// Positions in UTF-8 bytes.
console.log(parse("あい=3", Bytes).node!.children[0]!.end);

// Any rule can be the start rule.
const k = parseRule("key", "abc");
console.log(String(k.node), k.error);
console.log(parseRule("nope", "abc").error?.message);

// Syntax errors carry the position and what was expected.
const bad = parse("abc=x");
if (bad.error instanceof SyntaxError) {
  console.log(bad.error.line, bad.error.col, bad.error.expected, bad.error.message);
}

// JSON: the same as the engine's, byte for byte with marshal.
console.log(marshal(k.node));
console.log(JSON.stringify(k.node));
$ node main.ts   # Node.js 22.18 or later runs TypeScript directly
[(Pair Key="abc"@key Value="12"@value) (Pair Key="あい"@key Value="3"@value)]
Pair "abc" at 0-3
Pair "あい" at 8-10
8
"abc"@key null
rule nope is not defined
1 5 [ '(?0-9)' ] 1:5: syntax error: expected (?0-9)
{"type":"Match","rule":"key","start":0,"end":3,"text":"abc"}
{"type":"Match","rule":"key","start":0,"end":3,"text":"abc"}

Apart from the formatting of console.log, this is what the Go program of the Go guide prints for the same grammar and inputs.

The command

pego gen -lang ts -g <grammar> [-s <rule>] [-o <file>] [-recognize]
Flag Default Meaning
-lang go ts generates TypeScript
-g (required) The grammar: .pego, .json, or a .pegoc that contains the AST
-s saved in a .pegoc, otherwise main The start rule of parse
-o standard output The output file
-recognize off Also generate recognize, which checks input without building a tree

-pkg and -types are for Go and are rejected with -lang ts. As for Go, the output is deterministic (generating twice gives identical files), and grammar errors are reported with positions before any file is written.

The same generator is available from Go, for build tools:

code, err := pego.GenerateTypeScript(g, "main") // grammar, start rule
code, err = pego.GenerateTypeScript(g, "main", pego.WithRecognize())

pego.WithTypes() makes it return an error: typed values are generated for Go only.

The generated API

// Parses the whole input with the start rule (or the rule name).
export function parse(input: string | Uint8Array, unit?: Unit): ParseResult;
export function parseRule(name: string, input: string | Uint8Array, unit?: Unit): ParseResult;
// With -recognize: checks the input without building a tree; null if it matches.
export function recognize(input: string | Uint8Array, unit?: Unit): Error | null;

export interface ParseResult { node: Node | null; error: Error | null }

export type Unit = 0 | 1;
export const CodePoints: Unit; // the default
export const Bytes: Unit;

export class Node {
  type: string; rule: string; start: number; end: number; text: string;
  children: readonly (Node | null)[];
  fields: readonly NodeField[];   // in the order they were set
  field(name: string): Value;     // null if absent
  isTerminal(): boolean;
  toString(): string;             // the S-expression form of Node.String, quoted with Go's printable characters
  toJSON(): NodeJSON;             // for JSON.stringify
}
export interface NodeField { readonly name: string; value: Value }
export type Value = Node | number | bigint | string | boolean | null;
export function marshal(n: Node | null): string; // JSON exactly as encoding/json writes the engine's Node

export class SyntaxError extends Error { pos; line; col; expected: string[]; messages: string[]; reason(): string }
export class SyntaxErrors extends Error { errors: SyntaxError[] }

Compared with the Go API:

  • A result object instead of (node, err). parse does not throw for bad input: it returns { node, error }, with the same combinations as Go: a node and no error; no node and an error; or, after #recover, both a node and SyntaxErrors. It throws only for bugs (an exception that is not a parse error).
  • SyntaxError shadows the global of the same name in modules that import it, as in parser generators such as Peggy. Import it under another name if you need both: import { SyntaxError as PegoSyntaxError } from "./pairs.ts". Its message is Go's Error() ("1:5: syntax error: expected (?0-9)"), and reason() is Go's Message().
  • Field values are Node, numbers (or bigints, see values), strings, booleans or null; narrow them with instanceof Node or typeof.
  • JSON. marshal(node) gives exactly the bytes Go's json.Marshal gives, including its escapes of <, > and &, which is what the tests compare. JSON.stringify(node) gives the same keys in the same order (fields sorted), with JavaScript's escaping, and the same values except in two cases: ints beyond the safe integers become rounded numbers, and in Bytes mode an invalid input byte in text is written as the surrogate that stands for it (see units), such as "a\udcff" where marshal and the engine write "a�".
  • Deep trees. Left recursion, left-associative Pratt operators and foldl build trees as deep as the input is long, without deep calls. marshal and toString write trees of any depth (encoding/json gives up beyond 10,000 levels; marshal does not). JSON.stringify recurses and throws a RangeError on Node.js's default stack beyond a few thousand levels (between 3,000 and 5,000 for def e = e "+" n / n): use marshal for such trees.
  • One module per grammar. Every generated module exports the same names, so import two parsers under namespaces: import * as pairs from "./pairs.ts".
  • A parser keeps no state between calls; it is safe to call from several workers.

Running it

The module uses only erasable TypeScript syntax (no enums or namespaces), so runtimes that strip types run it as it is: Node.js 22.18 and later (or 22.6 with --experimental-strip-types), Deno and Bun. For browsers and older Node.js, compile it with your build (tsc, esbuild, Vite, ...). It needs ES2020 (for BigInt) and no DOM or Node.js types.

The tests type-check every generated parser with tsc --strict and also --noUnusedLocals, --noUnusedParameters, --noUncheckedIndexedAccess, --exactOptionalPropertyTypes and --erasableSyntaxOnly, so it compiles under strict project settings. Exclude it from linters and formatters: it starts with /* eslint-disable */.

Input, positions and units

JavaScript strings are UTF-16, but positions count Unicode code points (the default) or UTF-8 bytes (Bytes), as in the engine. The parser decodes the input first, so positions, columns, len and startPos/endPos in actions are the engine's:

  • A string is parsed as its UTF-8 encoding would be. A lone surrogate (which UTF-8 cannot encode) becomes U+FFFD, as TextEncoder does.
  • A Uint8Array is parsed as UTF-8 bytes, which is what the engine parses. An invalid sequence decodes as U+FFFD, one per byte, exactly as Go decodes it; in Bytes mode the positions count the original bytes.
  • In Bytes mode, text keeps the bytes, as Go strings do: in node text and in text(...), each invalid byte b is the lone surrogate U+DC00+b (U+DC80–U+DCFF, like Python's surrogateescape), so len, comparisons and toString see the bytes the engine sees. marshal writes such a byte as the character U+FFFD, as Go's JSON does for PEGO's module (depending on the Go release and settings, Go may write the escape \ufffd instead, the same JSON value). Valid input never contains these surrogates.

Node start and end are positions in the chosen unit, not indexes into the JavaScript string. To slice the input, use node.text (terminals) or keep the positions in code points and convert: Array.from(input).slice(n.start, n.end).join("").

Strings in actions behave as Go strings do: + concatenates, < and the other comparisons order by code point (Go's byte order of UTF-8, not JavaScript's UTF-16 order), and len counts code points or UTF-8 bytes. Expected tokens in error messages are sorted in the same order.

Values of actions

Ints in actions are 64-bit, as in Go: 9223372036854775807 + 1 wraps around to -9223372036854775808, / truncates toward zero and % takes the sign of the dividend. An int is a JavaScript number while it is a safe integer (within ±253−1) and a bigint beyond, so equal ints are always ===, and code that reads a field gets a number in all ordinary cases:

const n = node.field("Count");
if (typeof n === "number" || typeof n === "bigint") { /* an int */ }

marshal writes big ints exactly; JSON.stringify cannot represent them and writes them as (rounded) numbers.

Errors

The errors are those of the engine, with the same messages:

Error When error
Syntax error The input does not match SyntaxError: pos, line, col, expected, messages
Recovered errors #recover skipped parts of the input SyntaxErrors (and node is set)
Runtime error in an action e.g. action in main: division by zero Error
Too deep nesting more than 100,000 nested rule calls Error: nesting too deep: more than 100000 rule calls
JavaScript stack exhausted before the limit, see below Error: nesting too deep: the JavaScript stack overflowed at N rule calls

With the grammar rec.pego of the Go guide, generated with -recognize into recparser.ts:

import { parse, recognize, SyntaxErrors } from "./recparser.ts";

const { node, error } = parse("let a;let 1;let b;");
console.log(String(node));
if (error instanceof SyntaxErrors) {
  for (const e of error.errors) {
    console.log(e.line, e.col, e.reason());
  }
}
console.log(recognize("let a;let 1;let b;")?.message);
console.log(recognize("let a;"));
(Seq [(Seq "let" " " "a"@name ";") Error"let 1;"{message=`1:11: syntax error: expected (?a-z)`} (Seq "let" " " "b"@name ";")])@main
1 11 syntax error: expected (?a-z)
1:11: syntax error: expected (?a-z)
null

Deep nesting and the JavaScript stack

The generated parser calls rules recursively, like the generated Go parser, and has the same limit of 100,000 nested rule calls. But a JavaScript stack is far smaller than a goroutine's: Node.js's default stack (about 1 MB) holds some hundreds to a few thousand nested rule calls, depending on the grammar and on how warm the code is. Measured on Node.js 24: def n = "(" n ")" / "x" parses about 1,100 levels of parentheses, and the JSON grammar about 390 levels of nested arrays (each level is several rule calls). Beyond that, the parse returns an error instead of crashing:

nesting too deep: the JavaScript stack overflowed at 2292 rule calls

That is the way in which a TypeScript parser differs from the engine in practice: input the engine accepts can exhaust a small stack (the design record lists the others, which only input that is not valid UTF-8 can show). To accept deeper input, give the parser a bigger stack. Each nested rule call takes about 1 KB:

  • In a worker thread, set the stack size. With 128 MB, the parentheses grammar reaches the limit of 100,000:

    import { readFileSync } from "node:fs";
    import { Worker, isMainThread, parentPort, workerData } from "node:worker_threads";
    import { parse, marshal } from "./parser.ts";
    
    if (isMainThread) {
      const input = readFileSync(process.argv[2]!, "utf8");
      const w = new Worker(new URL(import.meta.url), { workerData: input, resourceLimits: { stackSizeMb: 128 } });
      w.on("message", (json: string) => console.log(json));
    } else {
      const r = parse(workerData as string);
      parentPort!.postMessage(r.error === null ? marshal(r.node) : r.error.message);
    }
    

    Nodes are class instances, which do not cross threads intact: send JSON (as here) or do the work in the worker.

  • On the main thread, node --stack-size=<KB> raises V8's limit, but not the operating system's limit on the thread's stack (typically 8 MB); beyond that the process crashes. Stay below it (--stack-size=7000 reaches about 15,000 rule calls).

The parity tests parse in a worker with a 1 GB stack, so that they also check the engine's error exactly at the limit.

What is supported

Everything pego gen supports for Go, except typed values:

Feature Generated TypeScript
All parsing expressions, captures, atomic and discarded expressions Yes
Pratt expressions, left recursion (direct and indirect) Yes
Predicates and variables, lookahead, cut Yes
Actions, types, built-in functions (foldl, map, concat, ...) Yes
#error and #recover, SyntaxErrors Yes
Position units (code points, bytes) unit argument
Another start rule parseRule
Recognition without a tree recognize, with -recognize
Typed values (-types) No: trees are Nodes
#stream and ParseStream, Document No (#stream is an ordinary repetition, as in Go)
WithMaxDepth No: fixed at 100,000 nested rule calls, and bounded by the stack (see above)

Keeping generated code up to date

As with Go, generated code is source: commit it, and regenerate it when the grammar changes or you upgrade PEGO (the embedded runtime comes from the version of pego that generated it). In a Go module, a go:generate line next to the grammar keeps both parsers in step:

//go:generate go tool pego gen -g pairs.pego -pkg pairsparser -o pairsparser/parser.go
//go:generate go tool pego gen -lang ts -g pairs.pego -o web/src/pairs.ts

In a JavaScript project, a script does the same:

{ "scripts": { "generate": "go run github.com/ornew/pego/cmd/pego@latest gen -lang ts -g pairs.pego -o src/pairs.ts" } }

(replace latest with the version you use, so that everybody generates the same file). To check in CI that the file is current, regenerate and compare: npm run generate && git diff --exit-code src/pairs.ts.

To test that the TypeScript parser agrees with the engine for your inputs, compare its tree with the JSON that pego parse prints (indented, so compare the parsed values):

pego parse -g pairs.pego -i 'abc=12; あい=3' > want.json
node --input-type=module -e '
import { readFileSync } from "node:fs";
import { isDeepStrictEqual } from "node:util";
import { parse, marshal } from "./pairs.ts";
const want = JSON.parse(readFileSync("want.json", "utf8"));
const got = JSON.parse(marshal(parse("abc=12; あい=3").node));
console.log(isDeepStrictEqual(want, got) ? "same" : "different");
'

PEGO's own tests do this for every grammar in examples/ and the engine's test grammars, in both units, with the inputs as strings and as bytes (design record 013).

Trade-offs

  • Size. The runtime is about 2,400 lines in every file, with a 3 KB table of printable characters; the rest grows with the grammar. pairs.pego generates 80,480 bytes, the JSON grammar (parsers/json) 114,192 bytes and the minilang grammar 153,227 bytes. Minified and compressed by a bundler, it is a fraction of that.
  • Speed. On Node.js 24 (Apple M-series), the JSON grammar parses a 260 KB document in about 18 ms in code points and 25 ms in bytes, once the code is warm. That is fine for editors and tools, but no match for JSON.parse: use a generated parser for formats that have no native parser.
  • Stack. Deep nesting needs a bigger stack than JavaScript gives by default (see above).
  • No typed values. Trees are generic Nodes, as from Parse in Go. Write a small conversion to your own types, or read fields with field(name) and instanceof Node.