CUE
github.com/ornew/pego/parsers/cue parses CUE, as defined by
The CUE Language Specification and implemented by the
reference parser, cuelang.org/go/cue/parser (v0.17.1, the latest release), with a parser generated by PEGO from
cue.pego. It reads syntax only: it does not evaluate, unify, load packages or format. It depends only on
the standard library.
go get github.com/ornew/pego/parsers/cue
Use
// A file, parsed and checked as cuelang.org/go/cue/parser.ParseFile does.
f, err := cue.ParseFile(`package config
import "strings"
#Service: {
name!: string
replicas: *2 | int & >=1
host: "\(name).example.com"
}
web: #Service & {name: strings.ToLower("WEB")} @go(Web)
`)
for _, d := range f.Decls {
if fd, ok := d.(*cue.Field); ok {
name := fd.Label.(*cue.Ident)
fmt.Println(name.Text, name.IsDefinition(), fd.Start, fd.End) // #Service true 34 124, web false 126 181
}
}
// Only the syntax tree of the grammar, without the checks that come after it (see Conformance).
f, err = cue.ParseAST(`a: b + 1 * c.d`)
be := f.Decls[0].(*cue.Field).Value.(*cue.BinaryExpr) // be.Op.Text is "+"; be.Y is the product
// Values of literals.
s, err := str.Unquote() // a *cue.String: escapes, hashes, multi-line indentation
n, err := num.Value() // a *cue.Int: 0xff, 1_000, 1.5Ki; Float has Rat and Float64
// Walk the tree, and list the comments.
cue.Inspect(f, func(n any) bool { /* ... */ return true })
cs := cue.Comments(src)
// Errors.
_, err = cue.ParseFile("a: 1\nb: [1, 2\nc: 3\n")
var se *cue.SyntaxError // a syntax error: Line, Col, Pos, Expected
_, err = cue.ParseFile("a: 1\nX=a: 2\nX=b: 3\n")
var ce *cue.SemanticError // found after the syntax: Msg, Line, Col, Span
ParseFile(input, unit...) (*File, error) |
The syntax tree, accepted exactly when cue/parser.ParseFile accepts the input: ParseAST, then Check, and an error for invalid UTF-8 |
Valid(input) bool |
Whether ParseFile accepts the input |
ParseAST(input, unit...) (*File, error) |
The syntax tree as the grammar reads it: the syntax of every experiment, without the checks of Check |
(*File).Check() error |
The errors cue/parser reports besides those of the syntax (see Conformance), as a *SemanticError |
Parse(input, unit...) (*Node, error), Recognize(input, unit...) error |
The tree of *Node, as the engine returns it; only checking (the same syntax as ParseAST) |
File, Package, ImportDecl, Field, Comprehension, StructLit, ListLit, BinaryExpr, ... |
The syntax tree, modeled on cuelang.org/go/cue/ast, each node with its Span |
Inspect(node, f), SpanOf(node) |
Depth-first traversal as ast.Inspect; the span of any node |
Comments(src, unit...) []Comment |
The comments of the source, in order, with their spans |
(*String).Unquote(), (*Int).Value(), (*Int).Int64(), (*Float).Rat(), (*Float).Float64(), (*Bool).Value() |
The values of literals, decoded as cuelang.org/go/cue/literal does (see literal_test.go) |
(*Ident).IsDefinition(), IsHidden(), (*Attribute).Split() |
Identifier kinds; the name and the body of an attribute |
*SyntaxError, *SemanticError |
Errors with positions |
Positions are in code points by default (Span is Start, End, with End exclusive); cue.ParseFile(src, cue.Bytes)
counts bytes, as the reference parser does.
The types follow cuelang.org/go/cue/ast: ast.BasicLit is split by kind into Int, Float and String (with Bool,
Null and Bottom as in ast); the package clause, the imports and the file attributes are in File.Decls; the
fields, embeddings, let clauses, comprehensions and ... are declarations; a field has a Label, an Alias (a
postfix alias, a~X), a Constraint (? or !), a Value and Attrs. A string with interpolations is an
Interpolation whose Elts alternate between String fragments and expressions.
Conformance
The reference for the language is cuelang.org/go/cue/parser at v0.17.1, which defines in code what the specification
leaves open, and the specification is its documentation. The grammar follows the scanner and the parser, not the
prose of the specification where they differ (the comma that the scanner inserts, which words are keywords where,
1.5K being an integer, the hashes of strings, try, ...). The specification's own examples (66 marked as CUE) are
in the corpus below.
The reference's results are vendored, and the tests compare with them, with no tool but Go. internal/refgen
(a module of its own, since it needs cuelang.org/go) ran cue/parser over a corpus and wrote
testdata/:
| Test | What is compared | Cases | Result |
|---|---|---|---|
TestCorpus |
The 121 .cue files of cue-lang/cue, the 3,726 CUE files inside its .txtar test archives, the 148 inputs of the tests of cue/parser, and the 66 examples of the specification: acceptance by ParseFile, and for accepted sources the syntax tree (every node with its kind, byte offsets, operators and literal text) and the comments |
4,061 (3,963 accepted, 98 rejected) | 0 differ |
TestMutants |
Each source with 8 random single edits (a byte or a span of bytes inserted, deleted or replaced, a line deleted, duplicated, swapped or joined, a token deleted, duplicated, swapped, inserted or replaced from a vocabulary of CUE tokens, truncation): acceptance and the hash of the tree | 32,488 | 0 differ |
TestGenerated |
Random strings (hashes, quotes, multi-line, escapes, interpolations, flaws), numbers and token sequences: acceptance and the hash of the tree | 15,000 (1,493 accepted) | 0 differ |
TestLiterals |
Unquote, Int.Value and Float.Rat against cue/literal: whether decoding fails, and the value |
4,126 literals (3,626 strings, 359 integers, 141 floats) | 0 differ |
FuzzParse |
ParseAST, Parse and Recognize accept the same inputs, no panic, spans nest |
go test -fuzz FuzzParse; 1.5 million inputs in 90 s |
no failure |
The same tests run on sets that are too large to vendor, made by refgen with other seeds (the mutants and the
generated inputs are recorded with the hash of the input, so the test detects that its generator and refgen's are
out of step):
git clone --branch v0.17.1 https://github.com/cue-lang/cue /tmp/cue
cd internal/refgen
go run . -cue /tmp/cue -mutants /tmp/mutants.txt.gz -per 150 -seed 2000000 # 609,150 mutants
go run . -cue /tmp/cue -generated /tmp/generated.txt.gz -count 250000 -seed 5000000 # 750,000 inputs
cd ../..
CUE_MUTANTS=/tmp/mutants.txt.gz go test -run TestMutants -v . # CUE_REPORTS=all lists every difference
CUE_GENERATED=/tmp/generated.txt.gz go test -run TestGenerated -v .
Measured at the commit of this README: 609,150 mutants (150 per source, seed 2,000,000) and 750,000 generated inputs (250,000 per family, seed 5,000,000) were checked, with no difference in acceptance or in the tree.
What ParseFile checks after the syntax, as cue/parser does (Check): a label in square brackets that does not have
exactly one element; syntax that an @experiment attribute enables (try clauses with else and otherwise,
aliasv2 postfix aliases, explicitopen's postfix ...) in a file that has not enabled it, an experiment that does
not exist, and a prefix alias in a file that has enabled aliasv2; import paths with characters that a path cannot
have; the indentation of the lines of a string with interpolations on several lines; the scope rules of
astutil.Resolve (an alias or let declared twice in a scope, a field named as an alias of the scope, a postfix
alias with the blank identifier or that clashes with a label alias, a reference in a pattern constraint to a field of its
struct); and input that is not UTF-8. Of the 98 sources of the corpus that the reference rejects, 16 are rejected
by these checks and not by the syntax: ParseAST, Parse and Recognize accept them. The grammar reads the syntax of
all experiments because a variable of the grammar that says which experiments are on would make every memoized rule a
function of it.
Deviations
- Errors.
ParseFilereturns the first error.cue/parserreturns a list and recovers, and its messages and positions differ from the messages of PEGO (expected "(", ..., with the line and column of the farthest failure). Only whether an input is accepted is the same. The messages ofCheckare the reference's. - Comments are listed by
Comments(compared with the reference's list on the corpus), not attached to nodes as doc and line comments, and there are no relative positions (token.RelPos: the spacing between tokens), no identifier resolution (ast.Ident.Node,Scope) and noast.BadExprfor a syntax error. - Nesting.
cue/parserrejects expressions nested 10,000 levels deep ("expression exceeds maximum nesting depth"); this parser stops at the generated parser's limit of 100,000 rule calls and reports an error, at a depth that depends on the construct (TestDepthmeasures it: 8,331 levels of structs, 14,282 of lists, 19,995 of parentheses, about 50,000 of calls, 99,981 prefix operators in a chain). Input nested between the two limits is accepted by one and rejected by the other. - Language versions. Only the latest syntax is read: there is no
@langversion selection of older releases. - Invalid UTF-8.
ParseFilerejects it ascue/parserdoes;ParseASTreads it inside strings and comments as U+FFFD, and rejects it elsewhere. - The Unicode tables (letters and digits in identifiers) are those of Go 1.27.1's
unicodepackage (Unicode 17.0.0), written into the grammar by internal/unitab;TestUnicodeTableschecks the edges of every range against the Go in use (CUE_UNICODE_ALL=1checks every code point) and fails after an upgrade that changes them.
Performance
On 256 KB of CUE (go test -bench .; minimum of 6 runs on a shared Apple M3 Max, Go 1.27.1), against
cue/parser.ParseFile with the default mode (no comments), which ParseFile equals in what it accepts:
| config | corpus | |
|---|---|---|
cue/parser.ParseFile |
7.6 ms (34 MB/s), 6.0 MB, 107,946 allocs | 6.2 ms (42 MB/s), 5.0 MB, 73,570 allocs |
ParseFile |
26.9 ms (9.8 MB/s), 14.8 MB, 141,878 allocs | 20.1 ms (13 MB/s), 10.3 MB, 115,313 allocs |
ParseAST |
24.6 ms | 17.9 ms |
Recognize |
26.9 ms | 25.4 ms |
Parse (the tree of *Node) |
38.8 ms | 32.1 ms |
config is a synthetic configuration of services (definitions, constraints, disjunctions, comprehensions,
interpolations, comments; 262,824 bytes); corpus is the sources of the test corpus that have no package clause,
imports or experiments, each as the value of a field (263,156 bytes: real CUE, but dense in the unusual). ParseFile
is 3.2 to 3.5 times slower than the reference's hand-written recursive-descent parser, allocating twice the bytes. (The
reference's time includes the resolution of identifiers, of which Check does the part that finds errors.)
go test -run '^$' -bench . -benchmem . # this module
CUE_BENCH_OUT=/tmp/corpus.cue CUE_BENCH_CONFIG_OUT=/tmp/config.cue go test -run XXX . # write the two inputs
cd internal/refgen && go run . -bench /tmp/config.cue # cue/parser on them
How the grammar got there (each step measured on config, ParseAST):
- Rules start with a lookahead on the first character, where a rule is tried at every declaration and rarely matches (a failure that is recorded in the set of expectations costs a search of the set).
- The file is read first inside a lookahead (
main_fast), where failures record nothing, and again outside it only if it is invalid, to give the error: 37 ms to 26 ms. - An expression that is a single operand is read without the Pratt loop (
atom_expr), and the experiments are checked inCheck, not in the grammar. - The only variable of the grammar is the number of hashes of the string being read (
sv): rules that read a variable are memoized per its value, which costs an allocation for each memoized call: about 7% (when the generated runtime shares one empty environment, as a trial, 26.6 ms become 24.6 ms and 145,000 allocations 20,000). Avoiding the variable needs a rule for every number of hashes, which limits it, so the variable stays. - A predicate that reads a captured value makes the recognizer build the values of everything that the capture calls.
The
__check of labels and the indentation check of strings with interpolations did that for most of the grammar, and are a lookahead and part ofCheck:Recognizewent from 40 ms to 27 ms.
Recognize is not faster than ParseAST here: ParseAST runs on the typed runtime with direct rules, Recognize on the
general one.
Regenerating
go generate ./parsers # parser.go from cue.pego (from the root of the repository)
go test ./parsers -update # the golden files (testdata/*.golden) from the engine
go run ./cmd/pego lint -g parsers/cue/cue.pego
go run ./parsers/cue/internal/unitab # the letter and digit classes at the end of cue.pego
cd parsers/cue/internal/refgen && go run . -cue /tmp/cue -o ../../testdata # the vendored results (about 3 MB)
The vendored data is described in testdata/README.md: the sources are those of cue-lang/cue at v0.17.1 (Apache License 2.0).