Running Parsers: the Go API, Backends, Compiled Grammars and the CLI
This guide is about the runtime side of PEGO: how to turn a grammar into a parser in Go, how to call it, what comes back,
and how to choose between the ways of executing it (closure, bytecode, iterative bytecode, generated Go code). It also
covers saving grammars as .pegoc files (in brief) and the pego command-line tool.
It assumes you can already write a grammar. If not, start with the language specification. Related guides: compiled grammars, generating a standalone Go parser, streaming input and reparsing edited text.
- The pipeline
- Compiling a grammar
- Choosing a start rule
- Parsing, errors and concurrency
- The
Nodetype - Position units
- Backends
- Recognition mode
- Compiled grammars (
.pegoc) - Memoization
- The
pegocommand - Choosing: a cheat sheet
All examples on this page use one small grammar, pairs.pego, that reads key=value pairs separated by optional
semicolons:
type Pair struct { Key Match, Value Match }
def main = ws ps:pair+ $$ -> $ps
def pair: Pair = k:key "=" v:value ";"? ws -> new Pair{Key: $k, Value: $v}
def key = @(?a-zあ-ん)+
def value = @(?0-9)+
def ws = (? \t\n)*
Install the library and the tool with:
go get github.com/ornew/pego # library
go install github.com/ornew/pego/cmd/pego@latest # command-line tool
From a checkout of the repository, go run ./cmd/pego ... is equivalent to pego ....
The pipeline
PEGO source (.pego) ──ParseGrammar──▶ grammar AST ──Compile──▶ *pego.Parser ──Parse──▶ *pego.Node
grammar JSON (.json) ──UnmarshalJSON─▶ ▲ │ ▲
grammar package (Go values) ───────────────┘ MarshalBinary LoadParser
▼ │
compiled grammar (.pegoc)
- The grammar AST (package
grammar) is the in-memory form of a grammar. It can come from PEGO source, from JSON, or be built by hand. - Compiling analyzes the AST, type-checks it (reporting grammar errors), and prepares it for execution. A
*pego.Parseris the result: a compiled grammar plus a start rule. - A compiled grammar (
.pegoc) is a savedParser. Loading one skips parsing, analysis and type checking (compiled-grammars.md). - A standalone Go parser generated by
pego genbypasses all of this at run time. See code-generation.md.
Compiling a grammar
From source
pego.CompileSource(src, start) does everything in one call:
package main
import (
_ "embed"
"fmt"
"log"
"github.com/ornew/pego"
)
//go:embed pairs.pego
var src string
func main() {
p, err := pego.CompileSource(src, "main")
if err != nil {
log.Fatal(err)
}
node, err := p.Parse("abc=12; あい=3")
if err != nil {
log.Fatal(err)
}
fmt.Println(node)
}
[(Pair Key="abc"@key Value="12"@value) (Pair Key="あい"@key Value="3"@value)]
Embedding the grammar with go:embed keeps it in a real .pego file that pego fmt and pego parse can also use.
Compile once and keep the *pego.Parser. Compilation is the expensive step, and a Parser is cheap to call.
Errors in the grammar (syntax errors, undefined rules, type errors) are reported by CompileSource, with positions
when the grammar came from source:
_, err = pego.CompileSource("def main = (", "main")
fmt.Println(err) // 1:13: expected an expression, found end of file
_, err = pego.CompileSource("def main = other", "main")
fmt.Println(err) // 1:12: undefined rule other
From an AST
CompileSource is ParseGrammar followed by Compile. Use the two steps when you want to look at or change the AST
(for example, to list the rules with g.Rules() or to reformat with grammar.Format(g)) before compiling:
g, err := pego.ParseGrammar(src)
if err != nil {
log.Fatal(err)
}
fmt.Println(len(g.Rules()), "rules")
p, err := pego.Compile(g, "main")
From Go values (programmatic grammars)
The grammar package defines the AST, so a grammar can also be built without any PEGO source: for a tool that
generates grammars, or to embed a tiny grammar in a test. This builds the equivalent of
def main = word $$ and def word = @(?a-z)+:
g := &grammar.Grammar{Statements: []grammar.Statement{
&grammar.RuleDef{
Name: "main",
Expr: &grammar.Seq{Items: []grammar.Expr{
&grammar.Ref{Name: "word"},
&grammar.EndInput{},
}},
},
&grammar.RuleDef{
Name: "word",
Expr: &grammar.Atomic{Expr: &grammar.Repeat{
Expr: &grammar.CharClass{Ranges: []grammar.CharRange{{Lo: 'a', Hi: 'z'}}},
Min: 1,
Max: -1, // no upper bound
}},
},
}}
fmt.Print(grammar.Format(g))
p, err := pego.Compile(g, "main")
if err != nil {
log.Fatal(err)
}
n, err := p.Parse("hello")
fmt.Println(n, err)
def main = word $$
def word = @(?a-z)+
(Seq "hello"@word)@main <nil>
grammar.Format prints an AST as PEGO source, which is a quick way to check what you built. A hand-built grammar goes
through the same analysis and type checking as one parsed from source, so mistakes are reported by Compile; the
errors have no line numbers because there is no source:
bad := &grammar.Grammar{Statements: []grammar.Statement{
&grammar.RuleDef{Name: "main", Expr: &grammar.Ref{Name: "missing"}},
}}
_, err = pego.Compile(bad, "main")
fmt.Println(err) // undefined rule missing
An AST can be exchanged as JSON with grammar.MarshalJSON and grammar.UnmarshalJSON, which round-trip the grammar
(grammar.Format of the decoded grammar equals the original). The command line does the same with
pego convert.
Choosing a start rule
Every rule can be a start rule. The second argument of Compile and CompileSource names it, and it must be defined:
_, err = pego.CompileSource(src, "nope")
fmt.Println(err) // start rule nope is not defined
Parser.WithStart(name) returns a parser for another start rule. It shares the compiled grammar, so it costs almost
nothing, and it is how you expose several entry points of one grammar (for instance expr for an expression
evaluator and main for a whole file):
keyParser, err := p.WithStart("key")
if err != nil {
log.Fatal(err)
}
n, err := keyParser.Parse("abc")
fmt.Println(n, err) // "abc"@key <nil>
A parse always has to consume the whole input. The start rule does not need an $$ of its own: if the rule matches only a
prefix, the parse fails with end of input among the expected items at the first unconsumed position:
_, err = keyParser.Parse("abc=1")
fmt.Println(err) // 1:4: syntax error: expected (?a-zあ-ん), end of input
On the command line, -s selects the start rule for parse, compile and gen (default: the one saved in a
.pegoc, otherwise main).
Parsing, errors and concurrency
func (p *Parser) Parse(input string, opts ...ParseOption) (*Node, error)
Parse returns the value of the start rule, or an error:
| Result | Meaning |
|---|---|
node, nil |
The input matched. |
nil, *pego.SyntaxError |
The input does not match. The error holds the farthest position reached and what was expected there. |
node, pego.SyntaxErrors |
The input matched after recovering from errors with #recover. The tree contains Error nodes. See errors-and-recovery.md. |
nil, error |
Anything else that stops the parse: a run-time error in an action (such as action in main: division by zero), or exceeding the nesting limit. These are plain errors, not syntax errors, and they are not subject to backtracking. |
_, err = p.Parse("abc=x")
var se *pego.SyntaxError
if errors.As(err, &se) {
fmt.Println(se.Pos, se.Line, se.Col, se.Expected, se.Messages)
fmt.Println(err)
}
4 1 5 [(?0-9)] []
1:5: syntax error: expected (?0-9)
Pos is the offset in the position unit, Line and Col are 1-based, Expected lists what would
have been accepted, and Messages holds the messages set with #error. se.Message() is the text without the
position.
A *pego.Parser is safe for concurrent use: share one among goroutines and call Parse on all of them (each call
has its own state). A Document, in contrast, is not. Concurrent first use of different backends and of recognition
mode is also safe.
The Node type
A parse result is a tree of *pego.Node:
type Node struct {
Type string // "Match", "Seq", "List", "Operator", "Error", or the name of a type from the grammar
Rule string // the rule that produced the node, if it had no action
Start int // start offset in the input
End int // end offset, exclusive
Text string // the text of a terminal
Children []*Node // children of Seq, List and Operator nodes (omitted elements are nil)
Fields Fields // struct fields and captures
}
Type |
Produced by | Where the data is |
|---|---|---|
Match |
A terminal: a literal, a character class, an atomic expression @e |
Text |
Seq |
A sequence in a rule without an action | Children (the elements that have a value) and Fields (captures) |
List |
A repetition | Children |
Operator |
A Pratt operator application without an action | Children: the operands and the operator (left, operator, right for an infix operator) |
Error |
A range skipped by #recover |
Text (the skipped input) and the message field |
| a type name | new T{...} in an action, or a terminal type |
Fields for struct types, Text for terminal types |
How to build the shapes you want is covered in trees-and-actions.md. Here is how to read them.
Fields. A field value is a *Node, an int, a string, a bool or nil. node.Field(name) returns the value
(nil if absent or if node itself is nil). node.Fields.Get(name) also tells you whether the field exists, and
ranging over node.Fields yields NodeField{Name, Value} in the order the fields were set:
for _, pair := range node.Children {
key := pair.Field("Key").(*pego.Node)
fmt.Printf("%s %q at %d-%d\n", pair.Type(), key.Text, key.Start, key.End)
}
first := node.Children[0]
v, ok := first.Fields.Get("Value")
fmt.Println(v, ok) // "12"@value true
_, ok = first.Fields.Get("Nope")
fmt.Println(first.Field("Nope") == nil, ok) // true false
Pair "abc" at 0-3
Pair "あい" at 8-10
Because a field is an any, assert its type: .(*pego.Node) for nodes and captures, .(int) for integers.
node.IsTerminal() reports whether a node is a terminal (Match or a terminal type); for those, Text holds the text.
Positions. Start and End (int32, which Go accepts as slice indexes; convert with int(n.Start) for
arithmetic with ints) delimit the input a node covers, as a half-open range. Note that a rule that
includes trailing whitespace covers it: the first Pair above ends at 8, after "; ". To get the matched text of any
node, slice the input (with byte positions this is a plain Go slice) or use text(...) in an action.
Printing. node.String() (also used by fmt.Println) renders an S-expression-like form meant for tests and
debugging: "a" is a terminal, Number"12" a terminal of type Number, (Seq ...) a sequence, [...] a list,
(Pair Key=... Value=...) a struct, and @rule marks a node that came from a rule without an action.
JSON. Nodes encode with encoding/json. Fields are an object with keys in lexicographic order, so the output is
deterministic, and rule, text, children and fields are omitted when empty:
out, _ := json.Marshal(first)
fmt.Println(string(out))
{"type":"Pair","start":0,"end":8,"fields":{"Key":{"type":"Match","rule":"key","start":0,"end":3,"text":"abc"},"Value":{"type":"Match","rule":"value","start":4,"end":6,"text":"12"}}}
This is also what pego parse prints by default. Nodes cannot be decoded back from JSON; the JSON is for inspection and
for exchanging results with other programs.
Position units
Positions can count Unicode code points (the default) or UTF-8 bytes. The unit affects node Start and End,
the Pos and Col of syntax errors, startPos and endPos in actions, the len builtin, and the ranges of
Document.Edit. It never changes what matches: matching always proceeds by code point.
n, _ := p.Parse("あい=3")
fmt.Println(n.Children[0].Field("Key").(*pego.Node).End) // 2
n, _ = p.Parse("あい=3", pego.WithUnit(pego.Bytes))
fmt.Println(n.Children[0].Field("Key").(*pego.Node).End) // 6
Syntax errors follow the unit, too (the same あい=x input, with the error at the x):
$ pego parse -g pairs.pego -i 'あい=x'
pego: 1:4: syntax error: expected (?0-9)
$ pego parse -g pairs.pego -unit bytes -i 'あい=x'
pego: 1:8: syntax error: expected (?0-9)
Choosing:
- Use bytes when you index into the Go string (
input[n.Start:n.End]), pass positions to other byte-oriented tools, or want to avoid converting offsets. Bytes is the natural unit for Go. - Use code points (the default) when positions are shown to people as character counts, or are exchanged with a system that counts characters (editors and language servers often count UTF-16 units, which neither unit matches).
- Invalid UTF-8 is read as one U+FFFD character per invalid byte, so positions stay well defined.
The benchmarks show almost no speed difference between the units.
The CLI flag is -unit codepoints|bytes.
Backends
A backend is the way a compiled grammar is executed. All of them produce identical results: the same trees (including positions), the same syntax errors and the same recovered errors. The test suite checks this, and the closure backend serves as the reference. They differ in speed, memory, how they treat deep nesting, and what they need.
| Backend | Option | How it runs | Needs the grammar AST |
|---|---|---|---|
| Closure | pego.Closure |
Compiles parsing expressions into Go closures | Yes |
| Bytecode (recursive) | pego.Bytecode |
A VM running portable bytecode; rule calls use Go recursion | No |
| Bytecode (iterative) | pego.BytecodeIterative |
The same bytecode on a VM with its own call stack | No |
| Generated Go | (not an option) | A Go parser generated ahead of time by pego gen |
At generation time |
The first three are selected per parse:
n, err := p.Parse("abc=12", pego.WithBackend(pego.BytecodeIterative))
pego.DefaultBackend, used when you give no option, is the closure backend when the grammar AST is available and
recursive bytecode otherwise (that is, for a parser loaded from a .pegoc file saved without the AST). On the command
line, -backend closure|bytecode|bytecode-iterative does the same.
Which one
The measurements are in benchmarks.md and their analysis in
performance.md; the figures below are summarized from them (one machine, so
treat them as orders of magnitude rather than promises). Run go test ./bench -bench . -benchmem to measure your own
grammar and input.
| You want | Choose | Why |
|---|---|---|
| The best default for a Go program that loads a grammar at start-up | Closure (the default) | It is faster than both VMs on every benchmarked workload. |
| The fastest parsing, with a fixed grammar | Generated Go | The benchmarks show it 28–45% faster than the closure backend, and its typed values (-types) faster still. |
| Input that can nest very deeply (untrusted JSON-like data, generated code) | Bytecode (iterative) | Rule calls live on the VM's own stack, not the Go stack. |
| A grammar distributed as a data file, not source | Bytecode, from a .pegoc |
See the compiled grammars guide. |
| Checking validity only | Any backend with RecognizeOnly |
No tree is built. |
The recursive bytecode VM is 1.1–1.3× slower than the closure backend in the benchmarks and the iterative VM 1.3–1.9× slower; you pay that for portability (the same bytecode is specified for other runtimes in bytecode.md) and, for the iterative VM, for the independence from the Go stack. Do not pick bytecode inside a Go program for speed.
Deep nesting and WithMaxDepth
Recursive grammars (parentheses, nested arrays, nested blocks) use one rule call per level of nesting. The closure and
recursive bytecode backends use the Go stack for those calls. To keep runaway input from exhausting memory, a parse
fails with an error once rule calls nest deeper than a limit: 100,000 by default, 10,000,000 for
BytecodeIterative. WithMaxDepth(n) changes it. This is the grammar def main = nested $$ with
def nested = "(" nested ")" / "x" (one rule call per level of parentheses, plus the calls of main and of the
innermost x):
Input (parentheses around x) |
Closure | Bytecode | Iterative |
|---|---|---|---|
| 50,000 deep | ok | ok | ok |
| 200,000 deep | nesting too deep: more than 100000 rule calls |
same error | ok |
| 1,000,000 deep | same error | same error | ok |
n, err := p.Parse(input, pego.WithBackend(pego.BytecodeIterative), pego.WithMaxDepth(100))
// nesting too deep: more than 100 rule calls
A Pratt expression reads the operand of a prefix operator and the right operand of an infix operator by recursion,
so each operator of a chain of prefix or right-associative operators (---x, a ^ b ^ c ^ ...) counts as one more
nested call; chains of left-associative operators are read in a loop and do not nest.
Two things to know:
- Lowering the limit is safe on every backend, and is a good idea for untrusted input: it bounds the work and the memory
a hostile input can cause. A nesting error is an ordinary
error, so treat it like any parse failure. - Raising the limit on the closure and recursive bytecode backends is limited by the Go stack. Go's goroutine
stack is capped at 1 GB by default; running past it is not a recoverable error but a fatal
stack overflowthat kills the process. In a trial with the grammar above, the closure backend parsed 300,000 levels withWithMaxDepth(1_000_000), but 1,000,000 levels withWithMaxDepth(2_000_000)on the recursive bytecode backend crashed. If you need depth beyond the default, useBytecodeIterativerather than raising the limit.
Generated parsers have the same default of 100,000 and no option to change it (see code-generation.md).
Recognition mode
RecognizeOnly() checks whether the input matches without building a tree. Parse then returns a nil node and the
same syntax errors as a full parse. It is faster and allocates less: the benchmarks measure recognition at 1.2–1.6×
faster than a full parse on the closure backend.
n, err := p.Parse("abc=12", pego.RecognizeOnly())
fmt.Println(n, err) // nil <nil>
n, err = p.Parse("abc=", pego.RecognizeOnly())
fmt.Println(n, err) // nil 1:5: syntax error: expected (?0-9)
What it does not do:
- Actions are not evaluated, so run-time errors in actions are not reported. With the grammar
def main: Ratio = a:@(?0-9)+ "/" b:@(?0-9)+ $$ -> new Ratio{Num: len($a) / (len($b) - len($b))}, the input12/34fails a full parse withaction in main: division by zerobut is accepted by recognition (pego parse -checkprintsok). Use recognition for syntax checks, and a full parse where action errors matter. - Captures that predicates read are still built, since a predicate can depend on them.
- It needs the grammar AST, so a parser loaded from a
.pegocsaved without the AST returns an error (recognition needs the grammar, which the compiled grammar omits). - It cannot be combined with
ParseStreamorNewDocument: both return errors (recognition is not supported for stream parsing,recognition is not supported for documents).
Typical uses are validating uploads or configuration before doing real work, linting loops, and pre-filtering large
inputs. On the command line: pego parse -g grammar.pego -check < input, which prints ok or the syntax error (and
exits with status 1).
Compiled grammars (.pegoc)
A compiled grammar is a saved *pego.Parser: bytecode, optionally the grammar AST, and the start rule, in a binary file
with a checksum. Loading one skips parsing, analysis, type checking and compilation, so a program starts faster and a
grammar can ship as data (a file or a go:embed byte slice).
data, err := p.MarshalBinary() // bytecode + grammar AST + start rule
slim, err := p.Marshal(pego.WithoutAST()) // bytecode + start rule only
q, err := pego.LoadParser(slim)
pego compile -g pairs.pego -o pairs.pegoc # with the AST
pego compile -g pairs.pego -no-ast -o pairs-noast.pegoc # bytecode only
pego parse -g pairs-noast.pegoc -f sexpr -i 'abc=12' # a .pegoc is accepted wherever a grammar is
A file without the AST runs only on the bytecode backends, and it cannot be used for recognition mode, linting,
conversion back to source or code generation. The compiled grammars guide covers what is kept
and lost, sizes and load times, embedding with go:embed, versions and compatibility, loading untrusted data, and when
to use a .pegoc rather than pego gen or compiling at start-up. Load only trusted files.
Memoization
PEGO parses with packrat memoization: the result of a rule at an input position is remembered, so the rule is evaluated at most twice at that position, however much backtracking the grammar does. This is what keeps parsing time roughly linear in the input. It is on by default, automatic, and has no public switch. The engine decides per rule:
- Rules that call no other rule, and rules referenced only once in the grammar, are not memoized in a normal parse, because their entries could never be reused.
- Other rules are memoized at a position from their second call there: most rules are called only once per position, and an entry that is never reused only costs memory. A rule whose calls do repeat is memoized from the first call for the rest of the parse.
- Rules in a left-recursive cycle are not memoized one by one; only the cycle's leader, the rule that grows the seed, is.
- Rules that read variables are memoized per combination of those variables' values, so memoization never changes the result; it only changes the cost.
What this means for you:
- Memory grows with input size. The memo table, like the tree, is proportional to the input. For very large or
unbounded input use a
#streamrepetition andParseStream, which discards memo entries and input already handed to you (see streaming.md). - Incremental parsing reuses the memo.
Documentkeeps it between edits, and memoizes the rules that a normal parse skips, so that more of a previous parse can be reused. - Generated parsers memoize in the same way. Memoization decisions are made by the same analysis and are saved in
.pegocfiles. - The design is in 003-packrat-parsing.md, and the history of how the engine's speed was tuned (including the memo table) in performance.md.
The pego command
pego with no arguments prints the usage and exits with status 1. Every command takes -h to list its flags.
Errors are printed to standard error as pego: ... and give exit status 1. fmt and convert accept flags before or
after their file arguments.
A <grammar> argument is PEGO source (.pego), a grammar in JSON (.json) or a compiled grammar (.pegoc). The type
is recognized by the contents of a .pegoc and by the .json suffix; anything else is parsed as PEGO source.
parse
pego parse -g <grammar> [-s <rule>] [-i <input>] [-f json|sexpr] [-stream] [-check]
[-unit codepoints|bytes] [-backend closure|bytecode|bytecode-iterative]
| Flag | Default | Meaning |
|---|---|---|
-g |
(required) | The grammar |
-s |
saved in a .pegoc, otherwise main |
Start rule |
-i |
standard input | The input text. Without -i, input is read from standard input. |
-f |
json |
Output format: json (indented, like encoding/json's MarshalIndent) or sexpr (Node.String()) |
-check |
off | Recognition only: print ok, or the syntax error |
-unit |
codepoints |
Position unit |
-backend |
chosen for the grammar | closure, bytecode or bytecode-iterative |
-stream |
off | Print each #stream element as soon as it matches, one per line (see streaming.md) |
$ pego parse -g pairs.pego -f sexpr -i 'abc=12; あい=3'
[(Pair Key="abc"@key Value="12"@value) (Pair Key="あい"@key Value="3"@value)]
$ printf 'abc=12; def=3\n' | pego parse -g pairs.pego -f sexpr
[(Pair Key="abc"@key Value="12"@value) (Pair Key="def"@key Value="3"@value)]
$ pego parse -g pairs.pego -s key -f sexpr -i 'abc'
"abc"@key
$ pego parse -g pairs.pego -check -i 'abc=12'
ok
$ pego parse -g pairs.pego -check -i 'abc=x'; echo "exit $?"
pego: 1:5: syntax error: expected (?0-9)
exit 1
Without -f sexpr the tree is printed as JSON, for instance (shortened) for abc=12:
{
"type": "List",
"start": 0,
"end": 6,
"children": [
{
"type": "Pair",
"start": 0,
"end": 6,
"fields": {
"Key": { "type": "Match", "rule": "key", "start": 0, "end": 3, "text": "abc" },
"Value": { "type": "Match", "rule": "value", "start": 4, "end": 6, "text": "12" }
}
}
]
}
(The tool prints the nested objects over several lines.) When the parse recovered from errors with #recover, the tree is
printed first and then the errors, with exit status 1, so you can still inspect the tree.
fmt
pego fmt [-w] [-l] [<file.pego> ...]
Formats PEGO source and keeps comments. Without -w it prints the result; -w rewrites the files in place;
-l lists the files whose formatting differs (use it as a CI check). Without file arguments it formats standard input
(-w is then an error). It accepts only .pego files; for JSON and compiled grammars it points you to convert.
$ cat messy.pego
// A messy grammar.
def main = ws "x"+ ws $$ // trailing comment
def ws = (? \t\n)*
$ pego fmt -l messy.pego
messy.pego
$ pego fmt -w messy.pego && cat messy.pego
// A messy grammar.
def main = ws "x"+ ws $$ // trailing comment
def ws = (? \t\n)*
$ pego fmt -l messy.pego # nothing listed: already formatted
convert
pego convert [-to pego|json] [-o <file>] <grammar>
Converts between PEGO source and the JSON form of the grammar AST. -to defaults to the other format (.pego becomes
JSON, .json becomes PEGO source), and is required for a .pegoc input. Without -o, the result goes to standard
output. JSON and compiled grammars carry no comments or layout, so converting back to source gives canonical formatting
without comments; keep the .pego as the source of truth.
pego convert -to json pairs.pego -o pairs.json
pego parse -g pairs.json -f sexpr -i 'abc=12' # JSON grammars work everywhere a grammar is expected
pego convert -to pego pairs.pegoc # recover the source of a compiled grammar that has the AST
compile
pego compile -g <grammar> [-s <rule>] [-no-ast] -o <file.pegoc>
Saves a compiled grammar; see compiled grammars. -s sets the default start rule stored in
the file, and -no-ast omits the AST. Both -g and -o are required.
gen
pego gen -g <grammar> -pkg <package> [-s <rule>] [-o <file>] [-types] [-recognize]
pego gen -lang ts -g <grammar> [-s <rule>] [-o <file>] [-recognize]
Generates a standalone Go parser (see code-generation.md) or, with -lang ts, a TypeScript
module (see typescript.md). -g is required, and -pkg for Go; without -o the code goes to
standard output.
explain, trace and profile
pego explain -g <grammar> [-s <rule>] [-i <input>] [-n <calls>]
pego trace -g <grammar> [-s <rule>] [-i <input>] [-max-depth <n>] [-rule <name>] [-failures] [-f text|json]
pego profile -g <grammar> [-s <rule>] [-i <input>] [-sort <column>] [-n <rows>] [-f text|json]
Debugging tools; all three also take -unit and -backend, and read the input from standard input without -i.
explain shows which rule calls recorded what a syntax error says was expected, trace prints every rule call of a
parse as an indented tree, and profile prints the cost of each rule with hints. In Go, the same information comes
from pego.WithTrace and pego.WithProfile. See debugging.md.
sample
pego sample -g <grammar> [-s <rule>] [-n 10] [-seed <n>] [-max-depth <d>] [-max-repeat <r>] [-max-len <bytes>] [-budget <steps>] [-coverage] [-invalid] [-f lines|json]
Prints distinct inputs that the grammar accepts, one Go-quoted string per line or as a JSON document; -invalid
prints near-miss inputs that it rejects instead. See sampling-and-fuzzing.md.
Choosing: a cheat sheet
| Situation | Use |
|---|---|
| Prototyping a grammar, or a tool that reads grammars at run time | CompileSource and the default backend; pego parse to experiment |
| A service that parses with a fixed grammar and wants the fastest parser | pego gen |
| A fixed grammar where start-up matters, or the grammar must be a data file | pego compile plus go:embed and LoadParser (see compiled grammars) |
| Validating input, no tree needed | RecognizeOnly() (or pego parse -check) |
| Untrusted or machine-generated input that may nest deeply | BytecodeIterative (or a lowered WithMaxDepth) |
| Input larger than memory | #stream and ParseStream |
| Editor-style reparsing after small edits | NewDocument |
| Positions that index a Go string | WithUnit(pego.Bytes) |
| Several entry points for one grammar | Parser.WithStart |
| Many goroutines | Share one *pego.Parser |
| Finding why an input fails or a parse is slow | pego explain, pego trace, pego profile (see debugging.md) |