PEGO

Lexical Structure

This chapter defines how the text of a grammar file is divided into tokens: characters, line terminators, white space, comments, identifiers, keywords, literals, escape sequences and punctuation. The structure of the file built from the tokens is the subject of Grammar Files, and the whole syntax is summarized in Syntax Summary.

Source text

A grammar file is a sequence of Unicode characters encoded in UTF-8. A character outside a comment, a string literal and a character class that is not white space and does not start a token is an error, reported together with the characters that follow it up to the next white space or token.

Errors in a grammar are reported with a 1-based line and a 1-based column. The column counts characters (a tab is one character), not bytes.

Line terminators

A line ends at a line feed (U+000A), at a carriage return followed by a line feed, or at a carriage return alone. A line terminator matters only where this specification says so: it ends a comment, it numbers the lines of error positions, and a line feed MUST NOT appear in a string literal or in a character class.

(Positions in the input being parsed count lines differently: see Positions.)

White space

White space is any Unicode white space character (the Unicode property White_Space, which includes the line terminators). It separates tokens and is otherwise insignificant, except where tokens MUST be adjacent (see Adjacency).

Comments

A comment starts with // and extends to the end of the line, up to but not including the line terminator. A comment is white space. There are no block comments. The characters // inside a string literal or a character class do not start a comment.

def digits = (?0-9)+ // one or more digits

Tokens

The tokens are identifiers, string literals, character classes, integer literals, capture references and punctuation. The tokens of a file are formed left to right, and the longest possible token is taken at each position (see Adjacency).

Identifiers

An identifier starts with a letter or _ and continues with letters, digits and _. Letters and digits are Unicode letters and digits. Identifiers are case-sensitive and there is no length limit.

identifier = ( letter | "_" ) { letter | digit | "_" } .

Identifiers name rules, types, fields, captures, variables, levels, attributes and the parameters of lambdas in actions. The case of the first letter matters for type and field names; see Names.

Keywords

The following identifiers are keywords. They MUST NOT be used as rule names or be referred to as rules:

def  infix  level  new  operand  package  postfix  pratt  prefix  skip  struct  terminal  type

A keyword in a position where it cannot appear is an error. The keywords def, infix, level, operand, package, postfix, prefix, skip and type also end a sequence of parsing expressions, because each of them starts the next definition or item; for that reason they cannot be used as capture labels either. The reference implementation accepts a definition whose name is a keyword, but no other rule can call it.

Some identifiers have a special meaning only in certain contexts and are not keywords:

Identifier Meaning Where
_ The top expression In a parsing expression
left, right, none Associativity After infix in a Pratt expression
true, false, nil Literal values In actions and predicates
len, text, foldl, foldr, map, list, concat Built-in functions In actions and predicates
startPos, endPos, children Built-in fields In field access

String literals

A string literal is a sequence of characters between double quotes. It matches exactly that sequence of characters (see Literals), and it is used as a value in actions. A line terminator MUST NOT appear inside it. The character " and the character \ are written as escape sequences (\" and \\).

"if"   "\n"   "say \"hi\""   "\u{1F600}"

Character classes

A character class starts with (?, ends with ) and lists characters and ranges between them. Its syntax and meaning are defined in Character classes. A character class MUST contain at least one item and MUST NOT contain a line feed. The sequence (? always starts a character class.

(?a-z_)   (?^"\\)   (?\-(?\))

Integer literals

An integer literal is a sequence of decimal digits. It is used as a repetition count in a{n,m} and in actions and predicates. A number that does not fit in an int of the implementation is an error. A negative number in an action is written with the unary - operator.

"a"{2,3}   [len($s) > 4]

Capture references

$ immediately followed by an identifier ($label) or by decimal digits ($1) is a capture reference, used in actions and predicates. $ followed by anything else is the end-of-line anchor, and $$ is the end-of-input anchor.

Punctuation

The following are the punctuation tokens. Each is a single token, whatever follows it.

_|_  ^^  $$  --  ->  =>  ==  !=  <=  >=  &&  ||  []
:  =  /  (  )  {  }  [  ]  ,  .  *  +  ?  &  !  @  -  ^  $  #  |  <  >  %
Token Used for
_|_, ^^, $$, ^, $, ., -- Bottom, anchors, any character and cut
/, *, +, ?, &, !, @, -, :, (, ), {, }, , Combinators, repetition and capture in parsing expressions
[, ] Predicates
# Attributes
-> The start of an action
=> A lambda in actions
==, !=, <, <=, >, >=, &&, ||, !, +, -, *, /, % Operators in actions and predicates
=, |, [], * Definitions and types (def a = ..., type A = B | C, []T, *T)
. Field access

Escape sequences

String literals and character classes accept the following escape sequences.

Escape Character
\n line feed (U+000A)
\t horizontal tab (U+0009)
\r carriage return (U+000D)
\f form feed (U+000C)
\v vertical tab (U+000B)
\a alert (U+0007)
\b backspace (U+0008)
\0 NUL (U+0000)
\u{h...} the code point with the hexadecimal value h... (for example \u{1F600})
\uhhhh the code point with the four-digit hexadecimal value hhhh
\c for any other character c that is not an ASCII letter or digit, the character itself (for example \\, \", \-, \)). Any other escape of an ASCII letter or digit is an error.

A hexadecimal escape that is not valid is an error.

Adjacency

The following constructs require their tokens to be written without intervening white space or comments. When white space is present, the tokens are read as separate constructs.

Construct Example With white space
Capture label:a x:expr x :expr is a syntax error
Bounded repetition a{n,m} a{2,3} a {2,3} is a syntax error
Level-restricted call rule(level) expr(assignment) expr (assignment) is a call of expr followed by a group
Attribute arguments #name(...) #error(message="x") #error (...) is an attribute without arguments followed by a group
Function call f(...) in actions len($s) len ($s) is a variable reference followed by a parenthesized expression

Tokens are formed by taking the longest possible match. In particular, -- is always the cut operator (--a is a cut followed by a, not two discards), $$ is always the end-of-input anchor, -> is always the start of an action, and $ followed by a letter, _ or a digit is a capture reference ($label, $1) rather than the end-of-line anchor. [] is one token, so a predicate cannot be empty and a list type is written []T with no space.

Limits

An implementation MAY limit the nesting depth of a grammar source and the number of errors it reports for one source. The reference implementation reports at most 100 errors (the last one says that there are more) and rejects with nesting too deep a source that nests parentheses deeper than about 500 levels.