PEGO

CSV

github.com/ornew/pego/parsers/csv parses CSV as defined by RFC 4180, with a parser generated by PEGO from csv.pego. It depends only on the standard library.

go get github.com/ornew/pego/parsers/csv

Use

// Records as strings, as encoding/csv's ReadAll returns them.
recs, err := csv.Records("name,note\r\nalice,\"likes \"\"quotes\"\", commas\"\r\n")
// [["name" "note"] ["alice" "likes \"quotes\", commas"]]

// A file with a header: every record must have as many fields as the header.
header, rows, err := csv.Table("id,name\n1,alice\n2,bob\n")
name := slices.Index(header, "name")

// Typed values with positions.
f, err := csv.ParseAST("a,\"b\"\"c\"\n\nd\n")
for _, r := range f.Records {
	for _, fl := range r.Fields {
		fmt.Println(fl.Start, fl.End, fl.Text, fl.Quoted(), fl.Value()) // 2 8 "b""c" true b"c
	}
}

// Only check.
ok := csv.Valid("a,\"b\" \n") // false: a space after the closing quote

// Syntax errors.
_, err = csv.Records("a,b\nc,d\"e\n")
var se *csv.SyntaxError
if errors.As(err, &se) {
	fmt.Println(se.Line, se.Col) // 2 4
}
Records(input) ([][]string, error) The records, each a list of field values; empty lines left out
Table(input) (header []string, rows [][]string, err error) The first record as the header and the others, each with as many fields as the header (*FieldCountError otherwise; ErrNoHeader for an input without records)
ParseAST(input, unit...) (*File, error) The file: a *File of *Records of *Fields, each with its Span
Valid(input) bool, Recognize(input, unit...) error Only check the input
Parse(input, unit...) (*Node, error) The tree of *Node, as the engine returns it
(*Field).Value() string The field's value: quotes removed and "" decoded (Text is the source)
(*Field).Quoted() bool Whether the field is quoted
(*Record).Strings() []string, (*File).Strings() [][]string The values of a record, of a file (as Records returns them)
(*Record).Blank() bool Whether the record is an empty line

Positions are in code points by default; csv.ParseAST(src, csv.Bytes) counts bytes. A record's span includes its line break; its fields end at the End of the last one. With code points, the Text of a field holds U+FFFD for each byte that is not part of valid UTF-8; with Bytes it keeps the bytes, and so do Records and Table.

The grammar works with the engine too: its #stream attribute hands records over one at a time in a stream parse (pego parse -g csv.pego -stream, or ParseStream in Go), for files larger than memory.

Where RFC 4180 and practice differ

The RFC defines the syntax with an ABNF grammar that real files often stray from. The package accepts every file of the RFC's grammar, and in the points where the RFC and practice differ it follows encoding/csv with its default settings (except FieldsPerRecord), so that it can replace it, with two exceptions where encoding/csv loses information.

Point RFC 4180 This package encoding/csv Why
Line break CRLF CRLF or LF CRLF or LF LF-only files are the norm on Unix and in most tools; the RFC warns that "some implementations may use other values" and asks readers to be "liberal in what you accept"
CR outside quotes not allowed in a field data, except at the end of the input, where it ends the record the same Compatibility with encoding/csv
Final line break optional optional optional
Empty input (by the grammar) one record of one empty field no records no records An empty file holds no data
Empty line (by the grammar) a record of one empty field ParseAST: a record of one empty field (Blank); Records, Table: skipped skipped Blank lines (often at the end of a file) are almost never meant as records; tools that need them have them in ParseAST. A file of one column loses its empty values in Records, as with encoding/csv
Records of different lengths "should" have the same number of fields Records: allowed; Table: an error (*FieldCountError) an error unless FieldsPerRecord is -1 (ErrFieldCount) The RFC makes it a recommendation, and ragged files are common; where columns matter, Table checks them against the header
Trailing comma (a,b,) an empty last field (the prose forbids the comma, the grammar reads it as a field) an empty last field an empty last field
Spaces around a quoted field ( "a", "a" ) not allowed syntax error syntax error (TrimLeadingSpace allows the leading one) Conformance; such a field is ambiguous
Spaces in an unquoted field data data data
Double quote in an unquoted field (a"b) not allowed syntax error syntax error (LazyQuotes allows it) Conformance; a bare quote usually means a broken file
Text after a closing quote ("a"b) not allowed syntax error syntax error
Characters printable ASCII (TEXTDATA) any, including non-ASCII, tabs, control characters and bytes that are not UTF-8 any Files are UTF-8 (or other encodings) in practice
CRLF inside a quoted field data kept: "a\r\nb" is a\r\nb turned into LF: a\nb The value is the text between the quotes; the conversion cannot be undone. Replace "\r\n" yourself if you want LF
Byte order mark at the start not mentioned skipped kept in the first field (and a quoted first field after it is a syntax error) Excel writes one in "CSV UTF-8"; it is an encoding signature, not data, and kept it breaks lookups of the first column name
Header optional, indicated by the MIME type's header parameter a record like the others; Table reads the first as the header no header support The syntax cannot tell; the caller knows
Delimiter, comments comma; none comma; none configurable (Comma, Comment) The grammar is generated with a fixed delimiter; use encoding/csv for other delimiters

Conformance

There is no official conformance suite for CSV. The tests check:

  • 73 cases (edgeCases in csv_test.go): the examples of section 2 of RFC 4180, the cases of csv-spectrum (rewritten from their descriptions, not vendored), and the points in the table above, 13 of them rejected. Each has its expected records and is checked against encoding/csv (FieldsPerRecord -1), which gives the same result on all but 7, the differences in the table: a byte order mark (4 cases) and CRLF in quoted fields (3 cases);
  • Records against encoding/csv on 2,000 files of random records written by encoding/csv's Writer with LF and with CRLF, and against the records written;
  • the same comparison on arbitrary input in FuzzRecords (go test -fuzz FuzzRecords): both must accept the same inputs and return the same records once the byte order mark is removed and CRLFs in values are turned into LF (6 million inputs passed when it was written);
  • the positions of ParseAST, the lines and columns of syntax errors, Table, and that Valid, Recognize, ParseAST and Records agree on every input.

All of them run with Go alone.

Performance

On a file of 5,000 records (210 KB) with quoted fields that contain commas, doubled quotes and line breaks (go test -bench ., Apple M3 Max, minimum of 8 runs):

Time Memory Allocations
encoding/csv ReadAll 0.75 ms 1.2 MB 10,033
ParseAST 1.59 ms 1.8 MB 172
ParseAST (Bytes) 1.79 ms 1.8 MB 173
Records 1.84 ms 3.6 MB 1,048
Valid 1.29 ms 0 0
Parse 2.46 ms 5.5 MB 308

encoding/csv is a hand-written reader that builds nothing but strings; ParseAST takes about twice its time while recording the position of every field and record, and Records 2.5 times. Valid builds nothing but runs the general code of the generated parser (one method per expression), not the direct rules ParseAST runs. The comments in csv.pego give the measured effect of the grammar's choices.