PEGO

Generating Standalone Go Parsers with pego gen

pego gen turns a grammar into a Go source file: a parser you commit to your repository, build with your program, and call like any other package. The generated code depends only on the standard library. It does not import PEGO, and it does not load or compile the grammar at run time.

This guide shows how to generate a parser, how to call it, what it supports compared with the engine (the pego.Parser of the runtime guide), how to keep it up to date, and when it is the right choice. pego gen -lang ts generates a TypeScript module with the same behavior instead; it has a guide of its own, TypeScript parsers.

Why generate

Engine (pego.CompileSource) Generated parser
Speed Closure backend by default The fastest backend: 28–45% faster than the closure backend, and typed values faster still (benchmarks, analysis)
Dependencies github.com/ornew/pego Standard library only
Grammar at run time Loaded and compiled (or loaded from .pegoc) Gone: it is compiled into Go code
Grammar changes Edit the grammar, restart Regenerate and rebuild
Recognition (no tree) RecognizeOnly Recognize, with -recognize
Streaming, incremental, backends Yes No (see what is supported)

Generate a parser when the grammar is fixed, the parser is on a hot path or in a small binary, or you want no dependency on PEGO at run time. Use the engine when the grammar is data, or when you need streaming or incremental parsing. A good deal of the time you can use both: the engine in tools and tests, and the generated code in production (see Recipes).

The generated parser returns the same tree and the same errors as the engine, so switching between them does not change what your code sees. The test suite checks this for every grammar in examples/ (design record 008).

Quick start

A complete, runnable setup. The grammar, pairs.pego, is the one used in the runtime guide:

type Pair struct { Key Match, Value Match }

def main = ws ps:pair+ $$ -> $ps
def pair: Pair = k:key "=" v:value ";"? ws -> new Pair{Key: $k, Value: $v}
def key = @(?a-zあ-ん)+
def value = @(?0-9)+
def ws = (? \t\n)*

The layout is a module with the grammar next to the code that generates from it, and the generated parser in its own package (see why its own package):

example.com/pairs
├── go.mod
├── pairs.pego
├── generate.go
├── main.go
└── pairsparser/        # created by go generate
    └── parser.go

Record pego as a tool dependency of the module, so that go generate runs the pinned version and nobody needs a separate install:

go get -tool github.com/ornew/pego/cmd/pego

This adds a tool github.com/ornew/pego/cmd/pego line to go.mod. Then declare the generation step with a go:generate directive, and create the output directory (pego gen -o does not create directories):

// generate.go
package main

//go:generate go tool pego gen -g pairs.pego -pkg pairsparser -o pairsparser/parser.go
mkdir pairsparser
go generate ./...

The result is pairsparser/parser.go, 1,919 lines for this grammar, starting with:

// Code generated by pego. DO NOT EDIT.

// Package pairsparser is a parser generated from a PEGO grammar.
package pairsparser

Now use it:

// main.go
package main

import (
	"encoding/json"
	"errors"
	"fmt"
	"log"

	"example.com/pairs/pairsparser"
)

func main() {
	node, err := pairsparser.Parse("abc=12; あい=3")
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(node)
	for _, pair := range node.Children {
		key := pair.Field("Key").(*pairsparser.Node)
		fmt.Printf("%s %q at %d-%d\n", pair.Type(), key.Text, key.Start, key.End)
	}

	// Positions in UTF-8 bytes.
	node, _ = pairsparser.Parse("あい=3", pairsparser.Bytes)
	fmt.Println(node.Children[0].End)

	// Any rule can be the start rule.
	node, err = pairsparser.ParseRule("key", "abc")
	fmt.Println(node, err)
	_, err = pairsparser.ParseRule("nope", "abc")
	fmt.Println(err)

	// Errors have the same fields as the engine's, but are types of the generated package.
	_, err = pairsparser.Parse("abc=x")
	var se *pairsparser.SyntaxError
	fmt.Println(errors.As(err, &se), se.Line, se.Col, se.Expected, err)

	out, _ := json.Marshal(node)
	fmt.Println(string(out))
}
[(Pair Key="abc"@key Value="12"@value) (Pair Key="あい"@key Value="3"@value)]
Pair "abc" at 0-3
Pair "あい" at 8-10
8
"abc"@key <nil>
rule nope is not defined
true 1 5 [(?0-9)] 1:5: syntax error: expected (?0-9)
{"type":"Match","rule":"key","start":0,"end":3,"text":"abc"}

The output is exactly what the engine prints for the same grammar and input.

The repository uses the same setup for the ready-made parsers (JSON and others): each is a module of its own with the grammar, the generated parser.go, code written around it and tests against the language's reference implementation; parsers/generate.go holds the go:generate lines and parsers/parsers_test.go the checks against the engine. The benchmark parsers in bench/gen are generated by go generate ./bench.

The pego gen command

pego gen -g <grammar> -pkg <package> [-s <rule>] [-o <file>] [-types] [-recognize] [-nodoc]
Flag Default Meaning
-lang go The language: go, or ts for TypeScript (see TypeScript parsers)
-g (required) The grammar: .pego, .json, or a .pegoc that contains the AST
-pkg (required for Go) The package name of the generated code. It must be a valid Go identifier.
-s saved in a .pegoc, otherwise main The start rule of the generated Parse function
-o standard output The output file
-types off Also generate Go types for the grammar's types and ParseAST (see typed values)
-recognize off Also generate Recognize, which checks input without building a tree (see recognition)
-nodoc off Leave out the package comment (// Package p is a parser generated from a PEGO grammar.), for a package that has its own documentation in another file

Details:

  • Without -o the code is written to standard output, which is handy for a look (pego gen -g g.pego -pkg p | less).

  • The output is formatted with gofmt, passes go vet, and is deterministic: generating twice gives identical files. That is what makes the up-to-date check possible.

  • Grammar errors are reported with positions, before any file is written:

    $ pego gen -g broken.pego -pkg pairs
    pego: broken.pego:2:1: expected an expression, found end of file
    $ pego gen -g typeerr.pego -pkg pairs -o /dev/null
    pego: typeerr.pego:2:34: undefined capture $y
    $ pego gen -g pairs.pego -pkg pairs -s nope -o /dev/null
    pego: pairs.pego: start rule nope is not defined
    $ pego gen -g pairs.pego -pkg foo-bar -o /dev/null
    pego: pairs.pego: invalid package name "foo-bar"
    
  • A .pegoc saved with -no-ast cannot be used, because code generation needs the grammar: pego: pairs-noast.pegoc: the compiled grammar omits the AST.

  • As with the other commands, without -s the start rule saved in a .pegoc is used (main for other grammars).

From Go

The same generator is available as a function, for build tools that produce parsers programmatically:

g, err := pego.ParseGrammar(src)
if err != nil {
	log.Fatal(err)
}
code, err := pego.GenerateGo(g, "pairsparser", "main") // grammar, package name, start rule
if err != nil {
	log.Fatal(err)
}
err = os.WriteFile("pairsparser/parser.go", code, 0o644)

Options follow the start rule: pego.GenerateGo(g, "pairsparser", "main", pego.WithTypes()) is pego gen -types, and pego.WithRecognize() is -recognize, and pego.WithoutPackageDoc() is -nodoc.

pego gen is a thin wrapper around pego.GenerateGo (and pego.GenerateTypeScript with -lang ts). Because a grammar built with the grammar package (see the runtime guide) is just an AST, it can be generated from, too.

The generated API

Every generated package has the same small API, defined in the file itself:

// Parse parses the whole input with the start rule. unit selects the position unit (CodePoints by default).
func Parse(input string, unit ...Unit) (*Node, error)

// ParseRule parses the whole input with the rule name.
func ParseRule(name, input string, unit ...Unit) (*Node, error)

type Unit int
const (
	CodePoints Unit = iota // the default
	Bytes
)

type Node struct { Start, End int32; Text string; Children []*Node; Fields Fields } // and Type() and Rule()
func (n *Node) Field(name string) any
func (n *Node) IsTerminal() bool
func (n *Node) String() string
type Fields []NodeField         // with Get(name) (any, bool); encodes to JSON with sorted keys
type NodeField struct { Name string; Value any }

type SyntaxError struct { Pos, Line, Col int; Expected, Messages []string }
func (e *SyntaxError) Error() string
func (e *SyntaxError) Message() string
type SyntaxErrors []*SyntaxError

Compared with the engine's API:

  • Parse and ParseRule replace Parser.Parse. The position unit is an optional trailing argument (Parse(input, pairsparser.Bytes)) instead of WithUnit. ParseRule(name, input) plays the role of Parser.WithStart(name).Parse(input); an unknown name is an error: rule nope is not defined. Like the engine, both require the rule to consume the whole input.

  • The types belong to the generated package. pairsparser.Node has the same fields, methods and JSON encoding as pego.Node, but it is a different Go type: you cannot pass one where the other is expected, and error checks use errors.As with the generated *pairsparser.SyntaxError (and pairsparser.SyntaxErrors for errors recovered with #recover). If other code needs pego.Node, convert through JSON, which is identical for both, or write code against the fields you need.

  • Fields. As with the engine, a field value is a *Node, int, string, bool or nil: assert it to *pairsparser.Node (as the example does with Field("Key")), not to *pego.Node.

  • Errors carry the same positions and messages. Recovered errors work too. With this grammar, rec.pego:

    def main = stmt* $$
    def stmt = s:(("let" " " name ";")) #recover(skip=(?^;)* ";") -> $s
    def name = @(?a-z)+
    

    generated into a package recparser, the node is returned together with the recovered errors:

    node, err := recparser.Parse("let a;let 1;let b;")
    fmt.Println(node)
    var errs recparser.SyntaxErrors
    if errors.As(err, &errs) {
    	for _, e := range errs {
    		fmt.Println(e.Line, e.Col, e.Message())
    	}
    }
    
    (Seq [(Seq "let" " " "a"@name ";") Error"let 1;"{message=`1:11: syntax error: expected (?a-z)`} (Seq "let" " " "b"@name ";")])@main
    1 11 syntax error: expected (?a-z)
    
  • A generated parser is safe for concurrent use: each call has its own state, and the tables are read-only after start-up. (Checked here with sixteen goroutines under the race detector.)

  • Runtime errors in actions (such as division by zero) and exceeding the nesting limit return plain errors, as in the engine.

What is inside the file

Part Contents
Runtime The parser state, memoization, left recursion, Pratt loop, attribute handling and the built-in functions of actions. About 1,800 lines, the same in every generated file, embedded from internal/engine/genrt/runtime.go.
Rule table For each rule, the results of the engine's static analysis (memoized? left-recursion leader? capture names?).
Parser expressions One method per expression, with literals, character classes and rule indices embedded as constants, and one per rule that is never memoized and has no captures, which calls the rule's body directly.
Actions and predicates Go expressions; lambdas become Go function literals.

That is why even a tiny grammar produces a file of a few dozen kilobytes. Measured sizes: pairs.pego (7 lines) generates 56,970 bytes; the JSON grammar (parsers/json, 1,514 bytes of source) generates 92,537 bytes; the minilang grammar (4,416 bytes) generates 134,245 bytes. The growth is in the per-grammar part.

Typed values (-types)

With -types (pego.WithTypes()), the generated package also defines a Go type for each type of the grammar and a function that returns the result as values of those types:

// ParseAST parses the input like Parse and returns the result as typed values.
func ParseAST(input string, unit ...Unit) (T, error) // T: the Go type of the start rule's type

For the JSON grammar (parsers/json), whose start rule has the type Value, the generated types are:

type Span struct{ Start, End int }

type String struct {
	Span
	Text string
}
// ... Number, Bool and Null likewise (terminal types)

type Member struct {
	Span
	Key   *String
	Value Value
}

type Object struct {
	Span
	Members []*Member
}
// ... Array likewise (struct types)

// Value is the union type Value = Object | Array | String | Number | Bool | Null.
type Value interface{ isValue() }

func ParseAST(input string, unit ...Unit) (Value, error)

and a type switch replaces the field lookups and type assertions of *Node:

v, err := jsonparser.ParseAST(`{"name": "pego", "tags": ["peg", "go"], "stars": 42}`)
if err != nil {
	log.Fatal(err)
}
for _, m := range v.(*jsonparser.Object).Members {
	switch x := m.Value.(type) {
	case *jsonparser.Array:
		fmt.Println(m.Key.Text, "array of", len(x.Elements))
	case *jsonparser.Number:
		fmt.Println(m.Key.Text, "number", x.Text, "at", x.Start)
	case *jsonparser.String:
		fmt.Println(m.Key.Text, "string", x.Text)
	}
}
name string pego
tags array of 2
stars number 42 at 49

How grammar types map to Go:

Grammar type Go type
int, string, bool the same
struct type T *T: the same fields, plus an embedded Span with the value's range
terminal type T *T: Span and Text
Match *Match: Span and Text
Error (left by #recover) *Error: Span, Text and Message
union U = A | B of node types an interface U, implemented by *A, *B and *Error (and by *Node if a member is a CST type)
[]T a slice; an empty list is an empty slice, not nil
*T *int, *string or *bool for basic types; otherwise the Go type of T, nil when absent
Seq, List, Operator and records *Node, as from Parse
node, terminal, any, unions with basic members any, holding the converted value

Details:

  • Names. Go types have the names of the grammar types. A name that the generated runtime already uses (Node, Match, Error, Parse, Unit, ...) gets a trailing underscore: a grammar type Node becomes Node_. Span is renamed the same way if a type or field is called Span.
  • Errors. ParseAST returns the errors of Parse. After #recover, an Error in a list of a union type stays in the list as *Error; where a struct or terminal type is expected it becomes nil (it is among the returned SyntaxErrors anyway).
  • Speed. When every value ParseAST can return has a Go type of its own (no Seq, List, Operator, records, node, terminal or any, also in struct fields) and no action reads a field of a struct value, the generated parser builds the values directly, without the nodes Parse makes: on the benchmarks ParseAST then takes a half to two thirds of the time of Parse and a quarter to a third of the memory (JSON 6.7 → 3.1 ms and 10.4 → 2.4 MB on a 262 KB input, less than Recognize; XML, the left-recursive calculator and the outline grammar alike; the Pratt calculator about 11% faster). Otherwise ParseAST runs Parse and converts the tree, about 10% slower than Parse (minilang). The design record explains both. The typed runtime adds about as much code to the generated file as the parser itself.
  • ParseRule still returns *Node; typed values exist for the start rule only.

Recognition (-recognize)

With -recognize (pego.WithRecognize()), the generated package also has

// Recognize checks that the whole input matches the start rule without building a tree, and returns the
// syntax errors Parse would return (SyntaxErrors for errors recovered with #recover).
func Recognize(input string, unit ...Unit) error

It is the engine's RecognizeOnly in generated form: the grammar is compiled once more with no rule building a value (except those whose values predicates read), into a second rule table in the same file. Actions are not evaluated, so Recognize does not report runtime errors in actions. On the benchmark workloads it takes about half the time of Parse (JSON 8.8 → 4.9 ms on a 262 KB input, CSV 4.0 → 1.8 ms), and it is the fastest way PEGO has to validate input. The price is a larger file: the second table adds about as much code as the first.

What is supported

Everything in the language that works with Parse is generated; the entry points that need engine-level machinery are not.

Feature Engine Generated
All parsing expressions, captures, atomic and discarded expressions Yes Yes
Pratt expressions, left recursion (direct and indirect) Yes Yes
Predicates and variables, lookahead, cut Yes Yes
Actions, types, built-in functions (foldl, map, concat, ...) Yes Yes
#error and #recover, SyntaxErrors Yes Yes
Position units (code points, bytes) WithUnit trailing Unit argument
Another start rule WithStart ParseRule
#stream and ParseStream Yes No: #stream is generated as an ordinary repetition, and there is no ParseStream
Document (incremental parsing) Yes No
RecognizeOnly Yes Recognize, with -recognize
WithMaxDepth Yes No: fixed at the default of 100,000 nested rule calls
Choice of backend Yes (it is a backend of its own)
Parser.Grammar(), saving, loading Yes No: the grammar is gone at run time

Notes:

  • Depth. Generated code uses the Go stack for rule calls and fails with nesting too deep: more than 100000 rule calls at the fixed limit. With the nested-parentheses grammar def nested = "(" nested ")" / "x", 50,000 levels parse and 200,000 levels fail with that error. If you must accept deeper nesting, the engine's BytecodeIterative backend with WithMaxDepth is the tool (runtime guide).
  • Memory. Generated parsers memoize in the same way as the engine, so memory still grows with the input. For input larger than memory you need the engine's ParseStream.
  • Using a stream grammar with the generator is fine: the #stream attribute is simply ignored, so the same grammar file serves both the engine (streaming) and the generated parser (whole-input).

Keeping generated code up to date

Generated code is source: commit it, review its diffs, and regenerate it whenever the grammar changes or you upgrade PEGO (the embedded runtime comes from the version of pego that generated it). Three habits keep it honest.

1. Generate with go generate. Put the directive next to the grammar, as in Quick start and in parsers/generate.go, and run go generate ./.... Pin the generator with go get -tool as shown above. (go install github.com/ornew/pego/cmd/pego@latest and //go:generate pego gen ... works, too, but every developer then has to install a matching version.) With several grammars, add one directive per grammar, as bench/generate.go does.

2. Test that it is up to date. A test regenerates in memory and compares with the committed file; it fails with a reminder when somebody changed the grammar and forgot to regenerate:

//go:embed pairs.pego
var grammarSrc string

// pairsparser/parser.go is up to date with pairs.pego.
func TestGeneratedIsUpToDate(t *testing.T) {
	g, err := pego.ParseGrammar(grammarSrc)
	if err != nil {
		t.Fatal(err)
	}
	want, err := pego.GenerateGo(g, "pairsparser", "main")
	if err != nil {
		t.Fatal(err)
	}
	got, err := os.ReadFile("pairsparser/parser.go")
	if err != nil {
		t.Fatal(err)
	}
	if string(got) != string(want) {
		t.Error("pairsparser/parser.go is out of date; run go generate ./...")
	}
}

If you change pairs.pego (here: def value = @(?0-9a-f)+) and skip regeneration, the test fails. Among the output:

--- FAIL: TestGeneratedIsUpToDate (0.00s)
    generated_test.go:49: pairsparser/parser.go is out of date; run go generate ./...
FAIL

This is the same test as TestGeneratedParsersAreUpToDate in parsers/parsers_test.go. In CI, go generate ./... && git diff --exit-code achieves the same without a test.

3. Test that it behaves like the engine. The engine is the reference implementation, and a one-time parity test protects you from the two diverging (for example after an upgrade):

func TestGeneratedMatchesEngine(t *testing.T) {
	p, err := pego.CompileSource(grammarSrc, "main")
	if err != nil {
		t.Fatal(err)
	}
	for _, in := range []string{"abc=12", "abc=12; あい=3", "", "abc=", "abc=1 x"} {
		want, werr := p.Parse(in)
		got, gerr := pairsparser.Parse(in)
		wj, _ := json.Marshal(want)
		gj, _ := json.Marshal(got)
		if string(wj) != string(gj) || fmt.Sprint(werr) != fmt.Sprint(gerr) {
			t.Errorf("%q:\n engine    %s %v\n generated %s %v", in, wj, werr, gj, gerr)
		}
	}
}

Comparing JSON and error text covers the whole tree (with positions) and all error details. Include inputs that fail, and inputs that exercise each feature of your grammar.

Trade-offs

  • Speed. The benchmarks in benchmarks.md measure the generated parsers 28–45% faster than the closure backend on every workload (for example JSON 6.7 ms against 11.8 ms for the closure backend and 3.6 ms for encoding/json, on a 262 KB input), and ParseAST faster still (JSON 3.1 ms). The standard library is still faster where it applies, because it builds no positioned typed tree. The numbers are from one machine and one commit; run go run ./bench/report to measure yours.
  • Binary and repository size. Each generated parser adds tens of kilobytes of source (see the sizes above) that you commit and that compiles into your binary. It does not pull in PEGO's compiler, analyzer or VMs; the imports are bytes, encoding/json, fmt, sort, strconv, strings and unicode/utf8.
  • Start-up. There is none: no grammar to parse or compile. If start-up time is the only concern and a source file is acceptable, loading a .pegoc is the cheaper alternative.
  • Flexibility. The grammar is frozen into the binary. Changing it means regenerating, rebuilding and redeploying; if users supply grammars at run time, use the engine.
  • Features. No streaming, incremental parsing, depth option or backend choice (see what is supported).
  • Review noise. Large diffs on regeneration (particularly after upgrading PEGO) are expected. Mark the files as generated for your review tool: the first line, // Code generated by pego. DO NOT EDIT., follows the Go convention that tools recognize.
  • Typed AST. Parse returns generic *Node trees like the engine, which is what makes the two interchangeable. For Go types matching the grammar's types, generate with -types and call ParseAST; for types of your own, write a small conversion from *Node or from the typed values, as parsers/json does (ToAny).

Recipes

One grammar, one package

Each generated file defines Node, SyntaxError, Parse and the rest at package level, so two generated files in the same package collide:

both/b.go:17:6: Node redeclared in this block
	both/a.go:17:6: other declaration of Node

Give every grammar its own package (and directory), as parsers/* and bench/gen/* do. If you also want several entry points of one grammar, use ParseRule rather than generating it twice.

Engine in development, generated in production

Keep one .pego file. Use pego parse and the engine (CompileSource) in tools, experiments and tests, where start-up and features matter more than speed; generate the production parser from the same file; and let the parity test prove they agree.

Choosing the start rule

Pass the usual entry rule as -s (for example pego gen -g lang.pego -pkg lang -s program) so that Parse is the common case, and use ParseRule("expr", input) for the others, for example for an expression evaluator embedded in a larger tool.

Wrapping the result in your own types

If the grammar declares the types you want, -types generates them. Otherwise: the generated Node is the same loose shape for every grammar. A thin layer keeps that out of the rest of your program:

type Pair struct{ Key, Value string }

func Pairs(input string) ([]Pair, error) {
	node, err := pairsparser.Parse(input)
	if err != nil {
		return nil, err
	}
	var out []Pair
	for _, n := range node.Children {
		out = append(out, Pair{
			Key:   n.Field("Key").(*pairsparser.Node).Text,
			Value: n.Field("Value").(*pairsparser.Node).Text,
		})
	}
	return out, nil
}

Design the grammar's types and actions (trees-and-actions.md) so that this layer is thin, and unit-test it with the engine and the generated parser alike.