PEGO

CEL

github.com/ornew/pego/parsers/cel parses CEL, the Common Expression Language, with a parser generated by PEGO from cel.pego. It depends only on the standard library.

go get github.com/ornew/pego/parsers/cel

Use

// An expression, checked as cel-go checks it: literals in range, depth and size within its limits.
e, err := cel.ParseExpr(`request.auth.claims.exists(c, c == "admin") && size(items) <= 10 ? "ok" : "no"`)

// Typed values with positions; ParseExpr is this followed by Check.
e, err = cel.ParseAST(`a.b + f(1,
  2u) // sum`)
call := e.(*cel.Binary).Right.(*cel.Call)
fmt.Println(call.Func.Name(), call.Start, call.End) // f 6 16

// Walk the tree.
cel.Inspect(e, func(e cel.Expr) bool {
	if id, ok := e.(*cel.Ident); ok {
		fmt.Println("variable", id.Name())
	}
	return true
})

// Decode literals.
s := (&cel.StringLit{Text: `'''a\nb'''`}).Value() // "a\nb"

// Print the tree back as source.
cel.Format(e) // `a.b + f(1, 2u)`

// Syntax errors.
_, err = cel.ParseExpr("a + (b *")
var se *cel.SyntaxError
if errors.As(err, &se) {
	fmt.Println(se.Line, se.Col, se.Message()) // 1 9 syntax error: expected "!", "(", "-", ...
}

// Only check.
ok := cel.Valid(`1 +`) // false
ParseAST(input, unit...) (Expr, error) The expression as the typed values below, each with its Span
ParseExpr(input, opts...) (Expr, error) ParseAST, then Check; the size of the input is checked first. It returns a *SyntaxError or a *CheckError
Valid(input, opts...) bool, Recognize(input, unit...) error Only check the input: Valid as ParseExpr does, Recognize the syntax and nothing else
Parse(input, unit...) (*Node, error) The tree of *Node, as the engine returns it
Check(e, opts...) error What the grammar does not check: literals out of range, the limits of depth, and with WithMacros the macros
CheckLimits(e, Limits) error, DefaultLimits The depth of recursion and the size that cel-go allows (250 and 100,000 code points)
CheckMacros(e) error, IsMacroCall(*Call) bool The calls of the standard macros that cel-go rejects, and the calls that are macros
WithLimits(Limits), WithMacros() Options of ParseExpr, Valid and Check
Format(e) string The expression as source, with the parentheses its precedence needs
Inspect(e, f), Children(e, f), SpanOf(e) Walk the tree; the range of any node
Constant(e) (any, bool) The Go value of a literal, or a list or map of literals
(*IntLit).Value(), (*UintLit).Value(), (*DoubleLit).Value(), (*StringLit).Value(), (*BytesLit).Value(), (*BoolLit).Value() The values of the literals, decoded as cel-go does (Text is the source)
(*Ident).Name(), (*Ident).Rooted(), (*Name).Value() Names, without the leading dot and the whitespace in them; field names without backquotes

Positions are in code points by default; cel.ParseAST(src, cel.Bytes) counts bytes (ParseExpr and the errors of Check count code points).

The tree

Type Fields
Ident Text A variable, a function or the name of a message: a, .a (the leading dot resolves in the root scope), pkg.Msg. Name() is the name without the dot and any whitespace.
IntLit, UintLit, DoubleLit Text -5 (the sign belongs to the literal, and may be followed by whitespace), 0x1F, 7u, 1.5e3, .5
StringLit, BytesLit Text The source with its quotes and its r or b prefix: 'a', """a""", r"\d", b"\xff"
BoolLit, NullLit Text
Paren X (x). cel-go drops parentheses; this tree keeps them.
Select X, Field *Name, Optional x.f, x.?f; a field escaped with backquotes (x.`a-b`) keeps them in Field.Text
Index X, Index, Optional x[i], x[?i]
Call Target, Func *Ident, Args f(a) with Target nil, x.f(a) with the receiver as Target; the macros are calls too
Unary Op, X !x, -x. A run of operators is nested: !!x is a Unary of a Unary, which cel-go reduces to x.
Binary Left, Op, Right Every binary operator, left associative: `a
Conditional Cond, Then, Else c ? a : b
ListLit Elems [a, b]; an element written ?a is an *Optional
MapLit Entries Each Entry has Key, Value and Optional ({?k: v})
MessageLit Type *Ident, Fields a.b.Msg{f: v}; each Field has Name *Name, Value and Optional
Optional X ?a as an element of a list

Expr is the union of the expression types. A value of ParseAST is one of them, or a node of Parse; every struct has an embedded Span{Start, End} (the end is exclusive).

Checks that are not syntax

cel-go reads literals and enforces limits after its ANTLR parser has accepted the input, and rejects what they find with a parse error. The grammar of this package is the syntax; Check (which ParseExpr and Valid call) reports the rest as a *CheckError:

  • an integer that does not fit int64 (9223372036854775808; the literal -9223372036854775808 does) or uint64, and a double that does not fit float64 (1e309; 1e-400 is 0);
  • an expression more than 250 levels deep. cel-go measures the depth in two ways, which CheckLimits models on the tree. While parsing, it rejects an expression for which more than 250 contexts of one rule of its grammar are open at once: each nested parenthesis, list element, map entry, call argument, index and branch of a conditional opens one of the rules for expressions, and the right operand of an operator opens one of the rule for that operator. While building the tree, it rejects more than 250 nested conditionals, selections, receiver calls, index operations and binary operators other than && and ||; so a chain of 251 selections or additions is rejected, and a chain of && or a run of ! is not. ParseExpr agrees with cel-go on all 584 sizes of 30 shapes of expression in testdata/ref/limits.tsv and on 1,000 random combinations of them;
  • an expression of more than 100,000 code points (ParseExpr only);
  • with WithMacros (or CheckMacros), a call of a standard macro with arguments it does not take: has(a), x.all(1, p). The grammar leaves the macros as calls, since the specification lets an application choose its macros; has(has(a.b)) is valid as in cel-go, the inner has being a selection after expansion. The macros of extension libraries (optMap, cel.bind, ...) are not checked.

ParseAST, Parse and Recognize do none of this, and the depth they accept is that of the generated parser's limit of 100,000 rule calls: about 20,000 nested parentheses, 16,600 nested calls, 14,000 nested lists and 12,500 nested maps; a run of !, a chain of selections, of binary operators or of conditionals is not limited by it. Deeper input fails with an error instead of exhausting the stack.

Conformance

The tests check the parser against the specification's conformance suite and against cel-go v0.32.0, the reference implementation (cel.dev/cel-go), whose parser is an ANTLR grammar:

  • The conformance tests of cel-spec (testdata/cel-spec, commit 40a3c900 of 2026-09-16, Apache 2.0): the 2,527 tests of the simple suite, 2,263 different expressions, all parse with ParseExpr, Recognize and Parse (0 failures, 0 skipped). The suite has no test that expects a parse error. go test -run TestConformanceSuite.
  • Differential tests against cel-go (TestDifferential, go test -run TestDifferential -v): for each expression of testdata/ref/*.tsv, ParseExpr must return the AST that cel-go returns, written in a canonical form that also has the source offset of the token of each node, and the string, bytes, int, uint and double values that it decodes, or must reject what cel-go rejects. The corpora are the conformance tests (2,263 expressions); the inputs of cel-go's own parser tests (358, 80 rejected); 854 hand-picked edge cases of the lexical syntax and the grammar (442 accepted); 12,000 random expressions and mutations of them (3,913 accepted); 4,000 random literals of every form of string, bytes and number (2,395 accepted); 1,000 expressions nested around the limit of depth (676 accepted, verdicts only); and 12,319 more for cel-go with the standard macros (7,496 accepted, verdicts only; compared with ParseExpr(src, WithMacros())).
  • Larger runs. internal/refgen can generate any number of expressions. A run of 389,475 expressions (go run . -seed 51 -n 300000 -literals 80000 -deep 6000 -macros 50000, then CEL_REF_DIR=... go test -run 'Differential|Limits|Agreement|Spans|Format') found no difference, nor did the 56,309 for the macros and a run of 156,308 more macro calls. Runs of 196,000 to 426,000 expressions on earlier versions of the grammar found one difference besides those that the smaller corpora had found while the grammar was written: has(has(a.b)), which is fixed.
  • ParseAST, Parse and Recognize agree on every input, also in bytes; the spans of every value lie in those of its parent in order (TestAgreement, TestSpans, over all of the above); Format prints each of the 181,336 accepted expressions of the large run as source that parses to the same expression (TestFormat); and go test -fuzz FuzzParse checks the same on arbitrary input (1.4 million executions without a failure).

How the reference results are made

internal/refgen is a module of its own, since it depends on cel-go. It parses expressions with cel-go's parser (parser.NewParser with the optional syntax and escaped identifiers on, variadic logical operators so that a || b || c is one call, and no macros) and writes one line for each expression: the expression as a Go string literal, a tab, and the AST or ERR@line:column of the first error. testdata/ref/*.tsv are its output:

git clone https://github.com/google/cel-spec        # commit 40a3c900, for the conformance expressions
cd internal/refgen
go run . -spec /path/to/cel-spec -out ../../testdata/ref -deep 1000
go run . -spec /path/to/cel-spec -out /tmp/big -seed 7 -n 150000 -literals 60000 -deep 6000 -macros 50000
cd ../.. && CEL_REF_DIR=/tmp/big go test -run Differential .   # the tests of the module, on the large set

The canonical form is described in the comment of canon_test.go and internal/refgen/canon.go. Offsets are the ones cel-go records for each node (the . of a selection, the ( of a call, the operator of a binary operation, the ? of a conditional, the : of an entry and a field, the first operator of a run of !; the start of anything else).

Known deviations

  • Syntax errors are reported where the PEG parser failed (the farthest position), with the tokens expected there; the text is not cel-go's, and the position is often not the one ANTLR reports (the same in 56% of the 209,178 rejected random expressions of the large run, 83% of the macro ones, 81% of the hand-picked ones and 15% of the literals, for which ANTLR reports the lexer's error at the start of the token). A list of expected tokens is not exhaustive: the lookaheads that select the alternatives of an operand do not appear in it, except for the characters that start an operand, and after a name ( is not listed.
  • The tree is not cel-go's AST: parentheses and runs of ! and - are kept, && and || are binary and left nested, a negative number is one literal as in cel-go, an optional selection is a Select, and the macros are not expanded. The tests reduce the tree to cel-go's form to compare them.
  • cel-go's choices are followed where the specification's grammar differs or is silent. A sign before a signed number is part of the literal, and -1u is the negation of 1u; --9223372036854775808 is an error; [,], {,} and M{,} are valid; \u and \U are errors in bytes literals; if{} is a message and if is an error; a.if and a.if() are valid; a message type may be dotted, with whitespace around the dots; a triple-quoted string ends at the first closing triple quote.
  • Invalid UTF-8 in the input is read as U+FFFD, one for each invalid byte, in names, literals and anywhere else, as cel-go reads it (the source is converted to runes first). Value of a literal decodes the same way.
  • Expressions are limited as cel-go does by default (DefaultLimits); its other limits (the number of syntax errors it reports, the size of error recovery, and the number of nodes created by macros) are not modeled.

Grammar

cel.pego follows the productions of the specification and the ANTLR grammar of cel-go. Where it differs for speed, the comments say what it measurably cost. The table has the time of Recognize and ParseAST on the 28 expressions of testdata/bench/policies.txt (2.9 KB) and on one expression of 34 KB made of them, for the grammar as it is and for variants that differ in one thing (each generated with pego gen -types -recognize -nodoc into a copy of the module, and run in turn, the minimum of 8 rounds of 0.3 s on one CPU of an Apple M3 Max):

Variant Recognize ParseAST long Recognize long ParseAST
The grammar as committed 196 µs 686 µs 2.32 ms 3.61 ms
First version (an alternative for each literal, a message before a call before an identifier, rules for characters) +39% +16% +38% +24%
No lookaheads to choose the alternative of an operand +26% +8% +29% +18%
A left-recursive rule for the members +51% +11% +70% +18%
A rule for the characters of a name +10% +4% +9% +5%
ws calls a rule (so it is memoized) +14% +9% +14% +7%
Five postfix lines instead of three +4% +4% +4% +13%
A call before an identifier +4% +1% +2% +4%
No lookahead for a message's name +3% +2% +3% +5%

A Pratt expression for the binary operators took as long as a rule for each precedence level, so the grammar has the shorter. The Go reference implementation is not in the table: the standard library has no CEL parser.

Performance

go test -bench . (Apple M3 Max, Go 1.27.1, 16 CPUs), on the 28 typical expressions (2.9 KB; access policies, admission checks and validation rules of 40 to 200 characters, in testdata/bench/policies.txt), parsed in turn, and on one expression of 256 KB made of them. cel-go is measured in internal/refgen (go test -bench .), outside this module, with its parser (cel.dev/cel-go/parser v0.32.0) on the same expressions:

28 expressions Throughput Allocations One of 256 KB Throughput Allocations
ParseAST 518 µs 5.6 MB/s 3.2 MB in 283 19.3 ms 13.6 MB/s 5.2 MB in 476
ParseExpr 570 µs 5.1 MB/s 3.2 MB in 482
Recognize 202 µs 14.3 MB/s none 18.0 ms 14.6 MB/s 346 KB in 25
Parse 374 µs 7.7 MB/s 1.4 MB in 135 25.9 ms 10.1 MB/s 25.8 MB in 2,180
cel-go's parser, with the standard macros 1,314 µs 2.2 MB/s 975 KB in 13,970 114 ms 2.3 MB/s 82 MB in 1,021,558

ParseAST takes 2.5 times less time than cel-go's parser on typical expressions and 6 times less on a long one, with 15 times less memory on the long one. The 3.2 MB allocated for 28 expressions is the slabs of the typed values: the generated parser takes the values of each type (256 of each, 1,024 of each list) from chunks that it allocates when the type is first used, about 115 KB for each small expression; with chunks of 8 and 64, ParseAST of the 28 expressions took 22% less time on one CPU (547 µs instead of 703 µs; the time of a long expression does not change). Recognize builds nothing, but it takes as long as ParseAST on a long expression: the generator compiles the rules of the typed parser into direct code, and not those of Recognize (see the code generation guide).

Development

go generate ./parsers                  # regenerate parser.go after changing cel.pego (from the repository root)
go test ./parsers                      # parser.go up to date; golden files on every backend of the engine
cd parsers/cel && go test ./...        # the conformance tests, the differential tests and the rest
go test -fuzz FuzzParse                # fuzz the parser
go test -bench .                       # benchmarks
CEL_REF_DIR=/path/to/ref go test .     # the differential tests on larger reference results (see above)
cd internal/refgen && go test -bench . # cel-go on the same expressions