XML
github.com/ornew/pego/parsers/xml parses XML as defined by
Extensible Markup Language (XML) 1.0 (Fifth Edition), as a non-validating processor,
with a parser generated by PEGO from xml.pego. It also applies
Namespaces in XML 1.0 (Third Edition) on request. It depends only on the standard
library.
go get github.com/ornew/pego/parsers/xml
Use
// A decoded tree: entities expanded, text and attribute values normalized, attribute defaults applied, and every
// well-formedness constraint checked.
tree, err := xml.Decode(`<!DOCTYPE catalog [<!ENTITY pego "PEGO">]>
<catalog><book id="b1"><title>&pego; & XML</title></book></catalog>`)
for _, b := range tree.Root.Elements() {
id, _ := b.Attr("id")
fmt.Println(id, b.Text()) // b1 PEGO & XML
}
// From bytes in UTF-8, UTF-16, ISO-8859-1 or US-ASCII, detected from the byte order mark and the declaration.
tree, err = xml.DecodeBytes(data)
// Namespaces: Space and Local of every element and attribute, and the namespace constraints.
err = tree.ResolveNamespaces()
// The syntax tree, with positions.
doc, err := xml.ParseAST(`<a x="1"><b>text</b><!-- note --></a>`)
for _, c := range doc.Root.Content {
if e, ok := c.(*xml.Element); ok {
fmt.Println(e.Name.Text, e.Start, e.End) // b 9 20
}
}
// Errors.
_, err = xml.Decode(`<a x="1" x="2"/>`)
var we *xml.WFError // a well-formedness constraint checked after the parse
var se *xml.SyntaxError // a syntax error, or the element type match
if errors.As(err, &we) {
fmt.Println(we.Line, we.Col, we.Msg) // 1 10 attribute x appears twice
}
Decode(input) (*Tree, error) |
The document decoded and checked (see Conformance) |
DecodeBytes(data) (*Tree, error) |
Transcode, then Decode |
WellFormed(input) error |
The first error Decode finds |
Transcode(data) (string, error) |
The text of a document given as bytes, in UTF-8 without a byte order mark |
(*Tree).ResolveNamespaces() error |
Namespace names, and the constraints of Namespaces in XML 1.0 |
Tree |
Version, Encoding, Standalone, DTD, Prolog, Root, Epilog and the syntax tree Doc |
Elem, Attr, Text |
Decoded elements (name, attributes, children), attributes (normalized value, Defaulted) and text (adjacent text merged) |
DTD, Entity, AttributeDecl |
The declarations of the internal subset that were processed |
ParseAST(input, unit...) (*Document, error) |
The syntax tree: Document, Element, Attribute, CharData, EntityRef, CharRef, CDSect, PI, Comment, Doctype and the declarations, each with its Span |
Parse(input, unit...) (*Node, error), Recognize(input, unit...) error |
The tree of *Node, as the engine returns it; only checking |
(*CharData).Value(), (*AttValue).Value(), (*CharRef).Rune(), ... |
Values of the syntax tree: line ends normalized, references decoded |
In the decoded tree, children are *Elem, *Text, *PI, *Comment, or *EntityRef for a reference to an entity
that was not expanded (an external entity, which a non-validating processor does not read, or an undeclared entity in a
document whose DTD was not read entirely). Elements that come from the replacement text of an entity name it in
Elem.Entity, and their Src spans are relative to that text.
Positions are in code points by default; xml.ParseAST(src, xml.Bytes) counts bytes. Decode reports code points.
Conformance
The grammar checks the syntax of a document entity with an internal DTD subset (every production of the
specification except those of the external subset and of external entities: conditional sections and text
declarations), the Char production everywhere, the name characters of the fifth edition, and the element type
match. Decode checks every other well-formedness constraint that does not need an external entity:
- Unique Att Spec, Legal Character (character references), No < in Attribute Values, No External Entity References,
Parsed Entity (no reference to an unparsed entity), No Recursion, Entity Declared (where it applies: no DTD, an
internal subset without parameter-entity references, or
standalone="yes"; in an attribute default, the entity must be declared before it); - the replacement text of every internal general entity that is referenced matches
content, and that of every internal parameter entity referenced between declarations matches the declarations of the internal subset (PE Between Declarations; it is parsed with the same grammar); - the input is valid UTF-8 (
DecodeBytesandTranscodealso check the encoding declaration against the encoding).
It processes the internal subset as XML requires: the first declaration of an entity or an attribute binds; after a reference to a parameter entity it does not read (an external one, or an undeclared one), it does not process entity and attribute-list declarations unless the document is standalone. It normalizes line ends, expands character and entity references, normalizes attribute values (and collapses spaces in attributes declared with a type other than CDATA), and adds the default values of declared attributes.
The W3C conformance suite
TestConformanceSuite runs the
W3C XML Conformance Test Suite (xmlts20130923, 2,585 tests) when XMLCONF names
its xmlconf directory:
curl -LO https://www.w3.org/XML/Test/xmlts20130923.tar.gz
tar xzf xmlts20130923.tar.gz
XMLCONF=$PWD/xmlconf go test -run ConformanceSuite -v . # XMLCONF_VERBOSE=1 lists the tests that do not pass
The suite is not vendored: its parts carry the licenses of their authors (the James Clark tests may be redistributed only unmodified, as the original archive; the Sun tests are "All Rights Reserved"), so a curated subset cannot be copied here. The core tests (xml_test.go) cover each constraint with documents of their own.
Each test is run with DecodeBytes, and the tests of the namespace recommendations also with ResolveNamespaces.
Valid and invalid documents must be accepted (a non-validating processor accepts invalid documents), and where the
suite gives the canonical output, the decoded tree must produce it; not-well-formed documents must be rejected. Every
test is listed with its outcome; measured at the commit of this README:
| Type | Tests | Pass | Not run |
|---|---|---|---|
| valid | 812 | 728 (228 of them with the canonical output compared) | 84: XML 1.1 or Namespaces 1.1 |
| invalid | 242 | 229 (34 with the canonical output) | 13: XML 1.1 |
| not-wf | 1,498 | 958 | 172: XML 1.1 or Namespaces 1.1; 309: editions 1–4 only (names that the fifth edition allows); 59: accepted, the error is in an external entity |
| error | 33 | 33: errors a processor may report or not (27), XML 1.1 (5), editions 1–4 (1) |
No test fails. The 59 not-well-formed documents that are accepted depend on external entities that a non-validating
processor does not read (their descriptions say so: conditional sections and text declarations in external subsets
and entities, recursion through an external entity); the test accepts them only because their ENTITIES attribute
is not none. Mutating the checks of Decode one at a time (unique attributes, legal characters, entity declared,
attribute defaults, normalization of tokenized types, undeclared namespace prefixes) makes the suite fail, so it
does test them; it does not test that declarations after an unread parameter entity are skipped, which the core
tests do.
encoding/xml
diff_test.go compares Decode (with ResolveNamespaces) with encoding/xml's Decoder.Token on
3,000 random documents without a DTD (elements, namespaces, attributes, references, CDATA sections, comments,
processing instructions), and FuzzDecode checks that encoding/xml accepts every such document this package
accepts, with the same tokens (go test -fuzz FuzzDecode; 5 minutes found nothing more than the differences below).
encoding/xml is not a conforming XML processor, so the suite is the authority. It differs in these ways (the
tests allow for each):
- it does not process the DTD: it reports it as a directive, and rejects references to the entities it declares;
- it does not normalize attribute values (white space stays as it is) nor line ends in comments and processing instructions;
- it rejects names with characters that the fifth edition allows and earlier editions did not (
<ީ/>), and versions other than1.0(XML 1.0 processes1.xas 1.0); - it accepts what is not well-formed: duplicate attributes, several root elements, text after the root element, an
XML declaration after the start, undeclared namespace prefixes,
xmlns:p="", and�(as U+FFFD).
Deviations and limits
- Not read: the external subset and external entities. A reference to an external parsed entity stays as an
*EntityRefin the tree; conditional sections and text declarations, which occur only there, are not parsed. The replacement text of an internal parameter entity is parsed like the internal subset: no conditional sections and no parameter-entity references inside declarations. - XML 1.1 is not supported: a document that declares version
1.1(or any1.x) is processed as XML 1.0, as the fifth edition requires of a 1.0 processor: what only 1.1 allows is rejected (control characters, also as character references), and NEL (U+0085) and LS (U+2028) are characters, not line ends. XML 1.1 is rarely used, and supporting it would mean a second set of character classes in the grammar. - Encodings:
Transcodesupports UTF-8, UTF-16, ISO-8859-1 and US-ASCII; other encodings are a fatal error, as XML allows. The suite has no test that needs another encoding and is not of typeerror. ParseAST,ParseandRecognizealone check the syntax and the element type match, not the constraints thatDecodechecks, and read invalid UTF-8 as U+FFFD (which is aChar).- Entity expansion is limited by
ExpansionLimit(16 Mi units of replacement text expanded per document), which stops exponential expansion ("billion laughs"). - Nesting is limited by the generated parser's depth limit of 100,000 rule calls: 24,999 levels of elements; deeper input fails with an error, not a stack overflow.
- The declarations of the predefined entities (
lt,amp, ...) are not checked against the forms XML prescribes, and the predefined meaning is used. - Errors inside the replacement text of an entity are reported at the outermost reference to it in the document.
Performance
On a catalog of 256 KB (go test -bench .; elements, attributes, references, comments and CDATA sections), on a
shared Apple M3 Max (minimum of 6 runs): ParseAST takes 3.8 ms and Recognize 3.9 ms, about the time of
encoding/xml's Decoder.Token loop over the document (4.0 ms), which builds no tree and records no positions;
ParseAST allocates 1.2 times the bytes in 320 times fewer allocations (251 against 80,914). Decode, which also
checks the constraints and builds the decoded tree, takes 4.3 ms (7.3 MB, 1,308 allocations). Parse (the tree of
*Node) takes 6.8 ms.