JSON
github.com/ornew/pego/parsers/json parses JSON as defined by RFC 8259,
with a parser generated by PEGO from json.pego. It depends only on the standard library.
go get github.com/ornew/pego/parsers/json
Use
// Go values, as encoding/json decodes into an any: map[string]any, []any, string, float64, bool, nil.
v, err := json.Decode(`{"name": "pego", "stars": 42}`)
// Typed values with positions.
ast, err := json.ParseAST(`{"a": [1, true], "b": "x\ty"}`)
for _, m := range ast.(*json.Object).Members {
fmt.Printf("%q at %d-%d: %T\n", m.Key.Value(), m.Span.Start, m.Span.End, m.Value)
}
// "a" at 1-15: *json.Array
// "b" at 17-28: *json.String
// Only check.
ok := json.Valid(`[1, 2, 3,]`) // false
// Syntax errors.
_, err = json.ParseAST(`{"a": [1, 2,]}`)
var se *json.SyntaxError
if errors.As(err, &se) {
fmt.Println(se.Line, se.Col, se.Message())
// 1 13 syntax error: expected "-", "0", "[", "\"", "false", "null", "true", "{", (? \t\r\n), (?1-9)
}
ParseAST(input, unit...) (Value, error) |
The value: *Object, *Array, *String, *Number, *Bool or *Null, each with its Span |
Decode(input) (any, error) |
Go values, as encoding/json decodes into an any |
ToAny(Value) (any, error) |
The conversion Decode uses |
Valid(input) bool, Recognize(input, unit...) error |
Only check the input |
Parse(input, unit...) (*Node, error) |
The tree of *Node, as the engine returns it |
(*String).Value() string |
The string with its escapes decoded (Text is the source between the quotes) |
(*Number).Float64(), (*Number).Int64() |
The number (Text is the source) |
(*Bool).Value() bool |
The boolean |
Positions are in code points by default; json.ParseAST(src, json.Bytes) counts bytes.
Conformance
The parser accepts exactly the JSON texts of RFC 8259: any value at the top level, surrounded by optional whitespace. The tests check:
- every case of JSONTestSuite (in
testdata/JSONTestSuite): the 95y_files are accepted and decode to whatencoding/jsondecodes, the 188n_files are rejected, andRecognizeandParseASTagree on the 35i_files; Decodeagainstencoding/jsonon 2,000 random documents, and on arbitrary input inFuzzDecode(go test -fuzz FuzzDecode): both must accept the same inputs and return the same values.
Where RFC 8259 leaves the behavior to the implementation, the package behaves like encoding/json:
- invalid UTF-8 in strings is accepted and decodes to U+FFFD, and so does an unpaired surrogate escape (
"\ud800"); - of duplicate keys,
Decodekeeps the last (ParseASTkeeps them all, in order); - a number outside the range of
float64(1e400) is a valid*Numberbut an error inDecodeandFloat64; - a byte order mark is rejected.
Nesting is limited by the generated parser's depth limit of 100,000 rule calls, about 33,000 levels of arrays (where
encoding/json stops at 10,000); deeper input fails with an error instead of exhausting the stack.
Performance
On an array of objects of 256 KB (go test -bench .), ParseAST is about a sixth faster than encoding/json decoding
into an any, and allocates a quarter more bytes in 230 times fewer allocations; Decode (which converts the typed
values) takes about 15% longer than encoding/json. Valid builds nothing, but it is not faster than
ParseAST and far slower than encoding/json's Valid, which is a hand-written scanner. See the
benchmarks of PEGO for the other backends.