Go
github.com/ornew/pego/parsers/golang parses Go source files as go/parser parses them, with a parser generated by
PEGO from golang.pego: it accepts exactly the files go/parser accepts and builds the same tree,
positions included. It depends only on the standard library.
go get github.com/ornew/pego/parsers/golang
Use
// A go/ast tree, as go/parser.ParseFile(fset, "x.go", src, parser.SkipObjectResolution) returns it,
// for go/printer, go/format, go/types and the like.
fset := token.NewFileSet()
f, err := golang.ParseFile(fset, "x.go", src, 0) // golang.ParseComments collects File.Comments too
// Typed values modeled on go/ast, each with its Span in the input (Bytes counts bytes).
file, err := golang.ParseAST(string(src), golang.Bytes)
for _, d := range file.Decls {
if fn, ok := d.(*golang.FuncDecl); ok {
fmt.Println(fn.Name.Text, fn.Start, fn.End)
}
}
// Only check.
ok := golang.Valid(string(src))
ParseFile(fset, filename, src, mode) (*ast.File, error) |
The go/ast tree, with the line table, //line directives and comment groups of go/parser; an error is a scanner.ErrorList with one error |
ParseAST(input, unit...) (*File, error) |
The tree as the Go types of the grammar (File, GenDecl, FuncDecl, CallExpr, ...), each with its Span |
ToGoAST(fset, filename, src, f, mode) *ast.File |
The conversion ParseFile uses; f must be parsed with Bytes |
Valid(src) bool, Recognize(input, unit...) error |
Only check the input |
Parse(input, unit...) (*Node, error) |
The tree of *Node, as the engine returns it |
(*StringLit).Value(), (*CharLit).Value() |
The string or rune a literal denotes |
(*IntLit).Constant(), (*FloatLit).Constant(), (*ImagLit).Constant() |
The go/constant value of a number |
*SyntaxError |
The position (Line, Col, Pos) and the expected tokens of a syntax error |
Positions are in code points by default; pass Bytes for byte offsets (which ToGoAST needs, and which makes the
parser reject invalid UTF-8 everywhere; see Deviations).
Conformance
The grammar mirrors the decisions of go/parser and go/scanner, not only the language specification: types are
expressions, composite literals in statement headers, the type parameter or array decision in type declarations, the
receive-or-channel-type reassociation, go and defer calls, type switch guards, case clauses, //line directives and
the literal checks of the scanner. Where go/parser accepts more than the specification (a type where an expression is
expected, x.(type) outside a type switch), so does this parser.
Every test compares the parser with the go/parser of the Go toolchain that runs it: acceptance, and when both accept,
the whole tree (every node, every position, the line table and the comments). Measured with go1.27.1 (go version),
cd parsers/golang && go test -v -count=1 .:
| Test | Input | Result |
|---|---|---|
TestGOROOT |
every .go file under GOROOT/src (8,078 files, 45 rejected by go/parser) |
0 differ |
TestGOROOTComments |
every 4th of those files (2,020), with ParseComments: File.Comments |
0 differ |
TestMutations |
7,032 copies of those files with one random change each (insertions of tokens, keywords, comments, malformed literals and characters that go/scanner rejects; deletions, truncations, duplicated lines); 3,077 accepted by go/parser |
0 differ |
TestEmbeddedSources |
1,285 programs embedded as string literals in the tests of GOROOT/src (1,055 accepted), many invalid on purpose |
0 differ |
TestOtherSources |
131 files of the testdata of go/parser, go/printer, gofmt and go/format that do not end in .go (104 accepted) |
0 differ |
TestCases, TestCasesInBodies |
over 500 hand-written programs, one or more for each decision the grammar mirrors, again inside function bodies | 0 differ |
TestInvalidUTF8, TestNesting, TestLiterals, TestGolden* |
invalid UTF-8, the depth limit, literal decoding, the tree shapes in testdata/ |
pass |
FuzzParseFile |
arbitrary input, with and without ParseComments |
see below |
The tests that need GOROOT/src skip when it is absent (runtime.GOROOT() locates it); the others run with Go alone.
More than the default run:
PEGO_GO_MUTATIONS=20000 go test -run TestMutations -v . # 233,640 mutated files: 102,394 accepted by go/parser, 0 differ
go test -run xxx -fuzz FuzzParseFile -fuzztime 120s . # 3,560,117 executions, no failure
go test -short . # a sample of each corpus, for quick checks
PEGO_GO_DUMP=<dir> makes a failing mutation test write the input of each failure to the directory.
The comparison ignores what the parser does not set (see Deviations): Doc and Comment fields, object resolution,
and the text and number of errors.
Deviations
- Errors. A syntax error is reported once, as a
*SyntaxErrorlisting the expected tokens (ParseAST,Parse,Recognize) or as ascanner.ErrorListwith one error (ParseFile), and no tree is returned.go/parserreports up to ten errors with messages of its own and returns the part of the tree that it could parse. Whether a file is accepted is the same (as measured above); the position and text of the error are not compared. - Mode.
ParseFilebehaves likego/parserwithSkipObjectResolution:Scope,UnresolvedandObjare not set. The only mode isParseComments, which fillsFile.Commentswith the groupsgo/parsermakes but does not set theDocandCommentfields of declarations, fields and specs. There is noImportsOnly,PackageClauseOnlyorAllErrors. - Nesting. The generated parser stops at 100,000 nested rule calls: 16,600 levels of parentheses (25,000 of blocks,
50,000 of array types), where
go/parserstops at 100,000 levels. Deeper input fails with an error ("nesting too deep") instead of exhausting the stack.TestNestingpins this. - Invalid UTF-8. With the default position unit, an invalid byte is read as U+FFFD, so the grammar cannot tell it
from a spelled-out U+FFFD in a comment, which
go/scanneraccepts.ParseFileandValidcheck the encoding first; withRecognizeandParseASTpassBytes, with which the grammar rejects invalid UTF-8 everywhere (TestInvalidUTF8). - Memory. The parser memoizes while it parses:
ParseASTallocates about 14 times the size of the source (go/parserabout 8 times; see Performance). - The toolchain. The grammar follows the Go 1.27
go/parser. With another toolchain the comparison tests measure the difference, if any.
Performance
go test -bench . -benchmem parses three large files of the standard library (go/parser/parser.go,
net/http/server.go, runtime/proc.go; 457 KB together) per iteration. Apple M3 Max, go1.27.1, medians of five runs:
| time | throughput | allocated | allocations | |
|---|---|---|---|---|
go/parser.ParseFile (SkipObjectResolution) |
4.5 ms | 103 MB/s | 3.9 MB | 98,776 |
golang.ParseAST |
23.4 ms | 19.6 MB/s | 6.6 MB | 4,950 |
golang.ParseFile (ParseAST and ToGoAST) |
28.8 ms | 16.0 MB/s | 12.3 MB | 71,800 |
golang.Parse (*Node) |
38.8 ms | 11.8 MB/s | 51.0 MB | 9,786 |
golang.Recognize |
37.8 ms | 12.1 MB/s | 50.7 MB | 9,758 |
ParseAST is about five times slower than the hand-written parser of the standard library, and ParseFile about six,
for a grammar that mirrors its decisions: the generated parser is a packrat parser with memoization, where go/parser
is a predictive parser that never backtracks. ParseAST builds the typed values directly (every value of the grammar
has a Go type of its own), and is faster and lighter here than Parse and Recognize, which use the generic path (not
analyzed further).
The grammar's optimizations are commented in golang.pego: not memoizing the whitespace rules, guarding
operators and operands by their first character so that failed attempts do not merge long lists of expected tokens (the
comments there give about a tenth, a third and a fifth of the time of a parse), and a shortcut for expressions without a
binary operator (15%, measured by removing it). Guarding the postfix operators and trying expression statements first did
not help and were dropped. These changes made the first version of the grammar about 2.5 times faster (ParseAST on the
same files: 61 ms before, 23 ms after). See the benchmarks of PEGO for the other backends.