PEGO

CUE

github.com/ornew/pego/parsers/cue parses CUE, as defined by The CUE Language Specification and implemented by the reference parser, cuelang.org/go/cue/parser (v0.17.1, the latest release), with a parser generated by PEGO from cue.pego. It reads syntax only: it does not evaluate, unify, load packages or format. It depends only on the standard library.

go get github.com/ornew/pego/parsers/cue

Use

// A file, parsed and checked as cuelang.org/go/cue/parser.ParseFile does.
f, err := cue.ParseFile(`package config

import "strings"

#Service: {
	name!:    string
	replicas: *2 | int & >=1
	host:     "\(name).example.com"
}

web: #Service & {name: strings.ToLower("WEB")} @go(Web)
`)
for _, d := range f.Decls {
	if fd, ok := d.(*cue.Field); ok {
		name := fd.Label.(*cue.Ident)
		fmt.Println(name.Text, name.IsDefinition(), fd.Start, fd.End) // #Service true 34 124, web false 126 181
	}
}

// Only the syntax tree of the grammar, without the checks that come after it (see Conformance).
f, err = cue.ParseAST(`a: b + 1 * c.d`)
be := f.Decls[0].(*cue.Field).Value.(*cue.BinaryExpr) // be.Op.Text is "+"; be.Y is the product

// Values of literals.
s, err := str.Unquote() // a *cue.String: escapes, hashes, multi-line indentation
n, err := num.Value()   // a *cue.Int: 0xff, 1_000, 1.5Ki; Float has Rat and Float64

// Walk the tree, and list the comments.
cue.Inspect(f, func(n any) bool { /* ... */ return true })
cs := cue.Comments(src)

// Errors.
_, err = cue.ParseFile("a: 1\nb: [1, 2\nc: 3\n")
var se *cue.SyntaxError // a syntax error: Line, Col, Pos, Expected
_, err = cue.ParseFile("a: 1\nX=a: 2\nX=b: 3\n")
var ce *cue.SemanticError // found after the syntax: Msg, Line, Col, Span
ParseFile(input, unit...) (*File, error) The syntax tree, accepted exactly when cue/parser.ParseFile accepts the input: ParseAST, then Check, and an error for invalid UTF-8
Valid(input) bool Whether ParseFile accepts the input
ParseAST(input, unit...) (*File, error) The syntax tree as the grammar reads it: the syntax of every experiment, without the checks of Check
(*File).Check() error The errors cue/parser reports besides those of the syntax (see Conformance), as a *SemanticError
Parse(input, unit...) (*Node, error), Recognize(input, unit...) error The tree of *Node, as the engine returns it; only checking (the same syntax as ParseAST)
File, Package, ImportDecl, Field, Comprehension, StructLit, ListLit, BinaryExpr, ... The syntax tree, modeled on cuelang.org/go/cue/ast, each node with its Span
Inspect(node, f), SpanOf(node) Depth-first traversal as ast.Inspect; the span of any node
Comments(src, unit...) []Comment The comments of the source, in order, with their spans
(*String).Unquote(), (*Int).Value(), (*Int).Int64(), (*Float).Rat(), (*Float).Float64(), (*Bool).Value() The values of literals, decoded as cuelang.org/go/cue/literal does (see literal_test.go)
(*Ident).IsDefinition(), IsHidden(), (*Attribute).Split() Identifier kinds; the name and the body of an attribute
*SyntaxError, *SemanticError Errors with positions

Positions are in code points by default (Span is Start, End, with End exclusive); cue.ParseFile(src, cue.Bytes) counts bytes, as the reference parser does.

The types follow cuelang.org/go/cue/ast: ast.BasicLit is split by kind into Int, Float and String (with Bool, Null and Bottom as in ast); the package clause, the imports and the file attributes are in File.Decls; the fields, embeddings, let clauses, comprehensions and ... are declarations; a field has a Label, an Alias (a postfix alias, a~X), a Constraint (? or !), a Value and Attrs. A string with interpolations is an Interpolation whose Elts alternate between String fragments and expressions.

Conformance

The reference for the language is cuelang.org/go/cue/parser at v0.17.1, which defines in code what the specification leaves open, and the specification is its documentation. The grammar follows the scanner and the parser, not the prose of the specification where they differ (the comma that the scanner inserts, which words are keywords where, 1.5K being an integer, the hashes of strings, try, ...). The specification's own examples (66 marked as CUE) are in the corpus below.

The reference's results are vendored, and the tests compare with them, with no tool but Go. internal/refgen (a module of its own, since it needs cuelang.org/go) ran cue/parser over a corpus and wrote testdata/:

Test What is compared Cases Result
TestCorpus The 121 .cue files of cue-lang/cue, the 3,726 CUE files inside its .txtar test archives, the 148 inputs of the tests of cue/parser, and the 66 examples of the specification: acceptance by ParseFile, and for accepted sources the syntax tree (every node with its kind, byte offsets, operators and literal text) and the comments 4,061 (3,963 accepted, 98 rejected) 0 differ
TestMutants Each source with 8 random single edits (a byte or a span of bytes inserted, deleted or replaced, a line deleted, duplicated, swapped or joined, a token deleted, duplicated, swapped, inserted or replaced from a vocabulary of CUE tokens, truncation): acceptance and the hash of the tree 32,488 0 differ
TestGenerated Random strings (hashes, quotes, multi-line, escapes, interpolations, flaws), numbers and token sequences: acceptance and the hash of the tree 15,000 (1,493 accepted) 0 differ
TestLiterals Unquote, Int.Value and Float.Rat against cue/literal: whether decoding fails, and the value 4,126 literals (3,626 strings, 359 integers, 141 floats) 0 differ
FuzzParse ParseAST, Parse and Recognize accept the same inputs, no panic, spans nest go test -fuzz FuzzParse; 1.5 million inputs in 90 s no failure

The same tests run on sets that are too large to vendor, made by refgen with other seeds (the mutants and the generated inputs are recorded with the hash of the input, so the test detects that its generator and refgen's are out of step):

git clone --branch v0.17.1 https://github.com/cue-lang/cue /tmp/cue
cd internal/refgen
go run . -cue /tmp/cue -mutants /tmp/mutants.txt.gz -per 150 -seed 2000000     # 609,150 mutants
go run . -cue /tmp/cue -generated /tmp/generated.txt.gz -count 250000 -seed 5000000   # 750,000 inputs
cd ../..
CUE_MUTANTS=/tmp/mutants.txt.gz go test -run TestMutants -v .     # CUE_REPORTS=all lists every difference
CUE_GENERATED=/tmp/generated.txt.gz go test -run TestGenerated -v .

Measured at the commit of this README: 609,150 mutants (150 per source, seed 2,000,000) and 750,000 generated inputs (250,000 per family, seed 5,000,000) were checked, with no difference in acceptance or in the tree.

What ParseFile checks after the syntax, as cue/parser does (Check): a label in square brackets that does not have exactly one element; syntax that an @experiment attribute enables (try clauses with else and otherwise, aliasv2 postfix aliases, explicitopen's postfix ...) in a file that has not enabled it, an experiment that does not exist, and a prefix alias in a file that has enabled aliasv2; import paths with characters that a path cannot have; the indentation of the lines of a string with interpolations on several lines; the scope rules of astutil.Resolve (an alias or let declared twice in a scope, a field named as an alias of the scope, a postfix alias with the blank identifier or that clashes with a label alias, a reference in a pattern constraint to a field of its struct); and input that is not UTF-8. Of the 98 sources of the corpus that the reference rejects, 16 are rejected by these checks and not by the syntax: ParseAST, Parse and Recognize accept them. The grammar reads the syntax of all experiments because a variable of the grammar that says which experiments are on would make every memoized rule a function of it.

Deviations

  • Errors. ParseFile returns the first error. cue/parser returns a list and recovers, and its messages and positions differ from the messages of PEGO (expected "(", ..., with the line and column of the farthest failure). Only whether an input is accepted is the same. The messages of Check are the reference's.
  • Comments are listed by Comments (compared with the reference's list on the corpus), not attached to nodes as doc and line comments, and there are no relative positions (token.RelPos: the spacing between tokens), no identifier resolution (ast.Ident.Node, Scope) and no ast.BadExpr for a syntax error.
  • Nesting. cue/parser rejects expressions nested 10,000 levels deep ("expression exceeds maximum nesting depth"); this parser stops at the generated parser's limit of 100,000 rule calls and reports an error, at a depth that depends on the construct (TestDepth measures it: 8,331 levels of structs, 14,282 of lists, 19,995 of parentheses, about 50,000 of calls, 99,981 prefix operators in a chain). Input nested between the two limits is accepted by one and rejected by the other.
  • Language versions. Only the latest syntax is read: there is no @lang version selection of older releases.
  • Invalid UTF-8. ParseFile rejects it as cue/parser does; ParseAST reads it inside strings and comments as U+FFFD, and rejects it elsewhere.
  • The Unicode tables (letters and digits in identifiers) are those of Go 1.27.1's unicode package (Unicode 17.0.0), written into the grammar by internal/unitab; TestUnicodeTables checks the edges of every range against the Go in use (CUE_UNICODE_ALL=1 checks every code point) and fails after an upgrade that changes them.

Performance

On 256 KB of CUE (go test -bench .; minimum of 6 runs on a shared Apple M3 Max, Go 1.27.1), against cue/parser.ParseFile with the default mode (no comments), which ParseFile equals in what it accepts:

config corpus
cue/parser.ParseFile 7.6 ms (34 MB/s), 6.0 MB, 107,946 allocs 6.2 ms (42 MB/s), 5.0 MB, 73,570 allocs
ParseFile 26.9 ms (9.8 MB/s), 14.8 MB, 141,878 allocs 20.1 ms (13 MB/s), 10.3 MB, 115,313 allocs
ParseAST 24.6 ms 17.9 ms
Recognize 26.9 ms 25.4 ms
Parse (the tree of *Node) 38.8 ms 32.1 ms

config is a synthetic configuration of services (definitions, constraints, disjunctions, comprehensions, interpolations, comments; 262,824 bytes); corpus is the sources of the test corpus that have no package clause, imports or experiments, each as the value of a field (263,156 bytes: real CUE, but dense in the unusual). ParseFile is 3.2 to 3.5 times slower than the reference's hand-written recursive-descent parser, allocating twice the bytes. (The reference's time includes the resolution of identifiers, of which Check does the part that finds errors.)

go test -run '^$' -bench . -benchmem .                                    # this module
CUE_BENCH_OUT=/tmp/corpus.cue CUE_BENCH_CONFIG_OUT=/tmp/config.cue go test -run XXX .   # write the two inputs
cd internal/refgen && go run . -bench /tmp/config.cue                      # cue/parser on them

How the grammar got there (each step measured on config, ParseAST):

  • Rules start with a lookahead on the first character, where a rule is tried at every declaration and rarely matches (a failure that is recorded in the set of expectations costs a search of the set).
  • The file is read first inside a lookahead (main_fast), where failures record nothing, and again outside it only if it is invalid, to give the error: 37 ms to 26 ms.
  • An expression that is a single operand is read without the Pratt loop (atom_expr), and the experiments are checked in Check, not in the grammar.
  • The only variable of the grammar is the number of hashes of the string being read (sv): rules that read a variable are memoized per its value, which costs an allocation for each memoized call: about 7% (when the generated runtime shares one empty environment, as a trial, 26.6 ms become 24.6 ms and 145,000 allocations 20,000). Avoiding the variable needs a rule for every number of hashes, which limits it, so the variable stays.
  • A predicate that reads a captured value makes the recognizer build the values of everything that the capture calls. The __ check of labels and the indentation check of strings with interpolations did that for most of the grammar, and are a lookahead and part of Check: Recognize went from 40 ms to 27 ms.

Recognize is not faster than ParseAST here: ParseAST runs on the typed runtime with direct rules, Recognize on the general one.

Regenerating

go generate ./parsers                         # parser.go from cue.pego (from the root of the repository)
go test ./parsers -update                     # the golden files (testdata/*.golden) from the engine
go run ./cmd/pego lint -g parsers/cue/cue.pego
go run ./parsers/cue/internal/unitab          # the letter and digit classes at the end of cue.pego
cd parsers/cue/internal/refgen && go run . -cue /tmp/cue -o ../../testdata   # the vendored results (about 3 MB)

The vendored data is described in testdata/README.md: the sources are those of cue-lang/cue at v0.17.1 (Apache License 2.0).