PEGO

Go

github.com/ornew/pego/parsers/golang parses Go source files as go/parser parses them, with a parser generated by PEGO from golang.pego: it accepts exactly the files go/parser accepts and builds the same tree, positions included. It depends only on the standard library.

go get github.com/ornew/pego/parsers/golang

Use

// A go/ast tree, as go/parser.ParseFile(fset, "x.go", src, parser.SkipObjectResolution) returns it,
// for go/printer, go/format, go/types and the like.
fset := token.NewFileSet()
f, err := golang.ParseFile(fset, "x.go", src, 0) // golang.ParseComments collects File.Comments too

// Typed values modeled on go/ast, each with its Span in the input (Bytes counts bytes).
file, err := golang.ParseAST(string(src), golang.Bytes)
for _, d := range file.Decls {
	if fn, ok := d.(*golang.FuncDecl); ok {
		fmt.Println(fn.Name.Text, fn.Start, fn.End)
	}
}

// Only check.
ok := golang.Valid(string(src))
ParseFile(fset, filename, src, mode) (*ast.File, error) The go/ast tree, with the line table, //line directives and comment groups of go/parser; an error is a scanner.ErrorList with one error
ParseAST(input, unit...) (*File, error) The tree as the Go types of the grammar (File, GenDecl, FuncDecl, CallExpr, ...), each with its Span
ToGoAST(fset, filename, src, f, mode) *ast.File The conversion ParseFile uses; f must be parsed with Bytes
Valid(src) bool, Recognize(input, unit...) error Only check the input
Parse(input, unit...) (*Node, error) The tree of *Node, as the engine returns it
(*StringLit).Value(), (*CharLit).Value() The string or rune a literal denotes
(*IntLit).Constant(), (*FloatLit).Constant(), (*ImagLit).Constant() The go/constant value of a number
*SyntaxError The position (Line, Col, Pos) and the expected tokens of a syntax error

Positions are in code points by default; pass Bytes for byte offsets (which ToGoAST needs, and which makes the parser reject invalid UTF-8 everywhere; see Deviations).

Conformance

The grammar mirrors the decisions of go/parser and go/scanner, not only the language specification: types are expressions, composite literals in statement headers, the type parameter or array decision in type declarations, the receive-or-channel-type reassociation, go and defer calls, type switch guards, case clauses, //line directives and the literal checks of the scanner. Where go/parser accepts more than the specification (a type where an expression is expected, x.(type) outside a type switch), so does this parser.

Every test compares the parser with the go/parser of the Go toolchain that runs it: acceptance, and when both accept, the whole tree (every node, every position, the line table and the comments). Measured with go1.27.1 (go version), cd parsers/golang && go test -v -count=1 .:

Test Input Result
TestGOROOT every .go file under GOROOT/src (8,078 files, 45 rejected by go/parser) 0 differ
TestGOROOTComments every 4th of those files (2,020), with ParseComments: File.Comments 0 differ
TestMutations 7,032 copies of those files with one random change each (insertions of tokens, keywords, comments, malformed literals and characters that go/scanner rejects; deletions, truncations, duplicated lines); 3,077 accepted by go/parser 0 differ
TestEmbeddedSources 1,285 programs embedded as string literals in the tests of GOROOT/src (1,055 accepted), many invalid on purpose 0 differ
TestOtherSources 131 files of the testdata of go/parser, go/printer, gofmt and go/format that do not end in .go (104 accepted) 0 differ
TestCases, TestCasesInBodies over 500 hand-written programs, one or more for each decision the grammar mirrors, again inside function bodies 0 differ
TestInvalidUTF8, TestNesting, TestLiterals, TestGolden* invalid UTF-8, the depth limit, literal decoding, the tree shapes in testdata/ pass
FuzzParseFile arbitrary input, with and without ParseComments see below

The tests that need GOROOT/src skip when it is absent (runtime.GOROOT() locates it); the others run with Go alone.

More than the default run:

PEGO_GO_MUTATIONS=20000 go test -run TestMutations -v .   # 233,640 mutated files: 102,394 accepted by go/parser, 0 differ
go test -run xxx -fuzz FuzzParseFile -fuzztime 120s .     # 3,560,117 executions, no failure
go test -short .                                          # a sample of each corpus, for quick checks

PEGO_GO_DUMP=<dir> makes a failing mutation test write the input of each failure to the directory.

The comparison ignores what the parser does not set (see Deviations): Doc and Comment fields, object resolution, and the text and number of errors.

Deviations

  • Errors. A syntax error is reported once, as a *SyntaxError listing the expected tokens (ParseAST, Parse, Recognize) or as a scanner.ErrorList with one error (ParseFile), and no tree is returned. go/parser reports up to ten errors with messages of its own and returns the part of the tree that it could parse. Whether a file is accepted is the same (as measured above); the position and text of the error are not compared.
  • Mode. ParseFile behaves like go/parser with SkipObjectResolution: Scope, Unresolved and Obj are not set. The only mode is ParseComments, which fills File.Comments with the groups go/parser makes but does not set the Doc and Comment fields of declarations, fields and specs. There is no ImportsOnly, PackageClauseOnly or AllErrors.
  • Nesting. The generated parser stops at 100,000 nested rule calls: 16,600 levels of parentheses (25,000 of blocks, 50,000 of array types), where go/parser stops at 100,000 levels. Deeper input fails with an error ("nesting too deep") instead of exhausting the stack. TestNesting pins this.
  • Invalid UTF-8. With the default position unit, an invalid byte is read as U+FFFD, so the grammar cannot tell it from a spelled-out U+FFFD in a comment, which go/scanner accepts. ParseFile and Valid check the encoding first; with Recognize and ParseAST pass Bytes, with which the grammar rejects invalid UTF-8 everywhere (TestInvalidUTF8).
  • Memory. The parser memoizes while it parses: ParseAST allocates about 14 times the size of the source (go/parser about 8 times; see Performance).
  • The toolchain. The grammar follows the Go 1.27 go/parser. With another toolchain the comparison tests measure the difference, if any.

Performance

go test -bench . -benchmem parses three large files of the standard library (go/parser/parser.go, net/http/server.go, runtime/proc.go; 457 KB together) per iteration. Apple M3 Max, go1.27.1, medians of five runs:

time throughput allocated allocations
go/parser.ParseFile (SkipObjectResolution) 4.5 ms 103 MB/s 3.9 MB 98,776
golang.ParseAST 23.4 ms 19.6 MB/s 6.6 MB 4,950
golang.ParseFile (ParseAST and ToGoAST) 28.8 ms 16.0 MB/s 12.3 MB 71,800
golang.Parse (*Node) 38.8 ms 11.8 MB/s 51.0 MB 9,786
golang.Recognize 37.8 ms 12.1 MB/s 50.7 MB 9,758

ParseAST is about five times slower than the hand-written parser of the standard library, and ParseFile about six, for a grammar that mirrors its decisions: the generated parser is a packrat parser with memoization, where go/parser is a predictive parser that never backtracks. ParseAST builds the typed values directly (every value of the grammar has a Go type of its own), and is faster and lighter here than Parse and Recognize, which use the generic path (not analyzed further).

The grammar's optimizations are commented in golang.pego: not memoizing the whitespace rules, guarding operators and operands by their first character so that failed attempts do not merge long lists of expected tokens (the comments there give about a tenth, a third and a fifth of the time of a parse), and a shortcut for expressions without a binary operator (15%, measured by removing it). Guarding the postfix operators and trying expression statements first did not help and were dropped. These changes made the first version of the grammar about 2.5 times faster (ParseAST on the same files: 61 ms before, 23 ms after). See the benchmarks of PEGO for the other backends.