CSV
github.com/ornew/pego/parsers/csv parses CSV as defined by RFC 4180,
with a parser generated by PEGO from csv.pego. It depends only on the standard library.
go get github.com/ornew/pego/parsers/csv
Use
// Records as strings, as encoding/csv's ReadAll returns them.
recs, err := csv.Records("name,note\r\nalice,\"likes \"\"quotes\"\", commas\"\r\n")
// [["name" "note"] ["alice" "likes \"quotes\", commas"]]
// A file with a header: every record must have as many fields as the header.
header, rows, err := csv.Table("id,name\n1,alice\n2,bob\n")
name := slices.Index(header, "name")
// Typed values with positions.
f, err := csv.ParseAST("a,\"b\"\"c\"\n\nd\n")
for _, r := range f.Records {
for _, fl := range r.Fields {
fmt.Println(fl.Start, fl.End, fl.Text, fl.Quoted(), fl.Value()) // 2 8 "b""c" true b"c
}
}
// Only check.
ok := csv.Valid("a,\"b\" \n") // false: a space after the closing quote
// Syntax errors.
_, err = csv.Records("a,b\nc,d\"e\n")
var se *csv.SyntaxError
if errors.As(err, &se) {
fmt.Println(se.Line, se.Col) // 2 4
}
Records(input) ([][]string, error) |
The records, each a list of field values; empty lines left out |
Table(input) (header []string, rows [][]string, err error) |
The first record as the header and the others, each with as many fields as the header (*FieldCountError otherwise; ErrNoHeader for an input without records) |
ParseAST(input, unit...) (*File, error) |
The file: a *File of *Records of *Fields, each with its Span |
Valid(input) bool, Recognize(input, unit...) error |
Only check the input |
Parse(input, unit...) (*Node, error) |
The tree of *Node, as the engine returns it |
(*Field).Value() string |
The field's value: quotes removed and "" decoded (Text is the source) |
(*Field).Quoted() bool |
Whether the field is quoted |
(*Record).Strings() []string, (*File).Strings() [][]string |
The values of a record, of a file (as Records returns them) |
(*Record).Blank() bool |
Whether the record is an empty line |
Positions are in code points by default; csv.ParseAST(src, csv.Bytes) counts bytes. A record's span includes its line
break; its fields end at the End of the last one. With code points, the Text of a field holds U+FFFD for each byte
that is not part of valid UTF-8; with Bytes it keeps the bytes, and so do Records and Table.
The grammar works with the engine too: its #stream attribute hands records over one at a time in a stream parse
(pego parse -g csv.pego -stream, or ParseStream in Go), for files larger than memory.
Where RFC 4180 and practice differ
The RFC defines the syntax with an ABNF grammar that real files often stray from. The package accepts every file of the
RFC's grammar, and in the points where the RFC and practice differ it follows encoding/csv with its default settings
(except FieldsPerRecord), so that it can replace it, with two exceptions where encoding/csv loses information.
| Point | RFC 4180 | This package | encoding/csv |
Why |
|---|---|---|---|---|
| Line break | CRLF | CRLF or LF | CRLF or LF | LF-only files are the norm on Unix and in most tools; the RFC warns that "some implementations may use other values" and asks readers to be "liberal in what you accept" |
| CR outside quotes | not allowed in a field | data, except at the end of the input, where it ends the record | the same | Compatibility with encoding/csv |
| Final line break | optional | optional | optional | |
| Empty input | (by the grammar) one record of one empty field | no records | no records | An empty file holds no data |
| Empty line | (by the grammar) a record of one empty field | ParseAST: a record of one empty field (Blank); Records, Table: skipped |
skipped | Blank lines (often at the end of a file) are almost never meant as records; tools that need them have them in ParseAST. A file of one column loses its empty values in Records, as with encoding/csv |
| Records of different lengths | "should" have the same number of fields | Records: allowed; Table: an error (*FieldCountError) |
an error unless FieldsPerRecord is -1 (ErrFieldCount) |
The RFC makes it a recommendation, and ragged files are common; where columns matter, Table checks them against the header |
Trailing comma (a,b,) |
an empty last field (the prose forbids the comma, the grammar reads it as a field) | an empty last field | an empty last field | |
Spaces around a quoted field ( "a", "a" ) |
not allowed | syntax error | syntax error (TrimLeadingSpace allows the leading one) |
Conformance; such a field is ambiguous |
| Spaces in an unquoted field | data | data | data | |
Double quote in an unquoted field (a"b) |
not allowed | syntax error | syntax error (LazyQuotes allows it) |
Conformance; a bare quote usually means a broken file |
Text after a closing quote ("a"b) |
not allowed | syntax error | syntax error | |
| Characters | printable ASCII (TEXTDATA) | any, including non-ASCII, tabs, control characters and bytes that are not UTF-8 | any | Files are UTF-8 (or other encodings) in practice |
| CRLF inside a quoted field | data | kept: "a\r\nb" is a\r\nb |
turned into LF: a\nb |
The value is the text between the quotes; the conversion cannot be undone. Replace "\r\n" yourself if you want LF |
| Byte order mark at the start | not mentioned | skipped | kept in the first field (and a quoted first field after it is a syntax error) | Excel writes one in "CSV UTF-8"; it is an encoding signature, not data, and kept it breaks lookups of the first column name |
| Header | optional, indicated by the MIME type's header parameter |
a record like the others; Table reads the first as the header |
no header support | The syntax cannot tell; the caller knows |
| Delimiter, comments | comma; none | comma; none | configurable (Comma, Comment) |
The grammar is generated with a fixed delimiter; use encoding/csv for other delimiters |
Conformance
There is no official conformance suite for CSV. The tests check:
- 73 cases (
edgeCasesin csv_test.go): the examples of section 2 of RFC 4180, the cases of csv-spectrum (rewritten from their descriptions, not vendored), and the points in the table above, 13 of them rejected. Each has its expected records and is checked againstencoding/csv(FieldsPerRecord-1), which gives the same result on all but 7, the differences in the table: a byte order mark (4 cases) and CRLF in quoted fields (3 cases); Recordsagainstencoding/csvon 2,000 files of random records written byencoding/csv'sWriterwith LF and with CRLF, and against the records written;- the same comparison on arbitrary input in
FuzzRecords(go test -fuzz FuzzRecords): both must accept the same inputs and return the same records once the byte order mark is removed and CRLFs in values are turned into LF (6 million inputs passed when it was written); - the positions of
ParseAST, the lines and columns of syntax errors,Table, and thatValid,Recognize,ParseASTandRecordsagree on every input.
All of them run with Go alone.
Performance
On a file of 5,000 records (210 KB) with quoted fields that contain commas, doubled quotes and line breaks
(go test -bench ., Apple M3 Max, minimum of 8 runs):
| Time | Memory | Allocations | |
|---|---|---|---|
encoding/csv ReadAll |
0.75 ms | 1.2 MB | 10,033 |
ParseAST |
1.59 ms | 1.8 MB | 172 |
ParseAST (Bytes) |
1.79 ms | 1.8 MB | 173 |
Records |
1.84 ms | 3.6 MB | 1,048 |
Valid |
1.29 ms | 0 | 0 |
Parse |
2.46 ms | 5.5 MB | 308 |
encoding/csv is a hand-written reader that builds nothing but strings; ParseAST takes about twice its time while
recording the position of every field and record, and Records 2.5 times. Valid builds nothing but runs the general
code of the generated parser (one method per expression), not the direct rules ParseAST runs. The comments in
csv.pego give the measured effect of the grammar's choices.