Wolfram Language
Paclet Repository
Community-contributed installable additions to the Wolfram Language
Primary Navigation
Categories
Cloud & Deployment
Core Language & Structure
Data Manipulation & Analysis
Engineering Data & Computation
External Interfaces & Connections
Financial Data & Computation
Geographic Data & Computation
Geometry
Graphs & Networks
Higher Mathematical Computation
Images
Knowledge Representation & Natural Language
Machine Learning
Notebook Documents & Presentation
Scientific and Medical Data & Computation
Social, Cultural & Linguistic Data
Strings & Text
Symbolic & Numeric Computation
System Operation & Setup
Time-Related Computation
User Interface Construction
Visualization & Graphics
Random Paclet
Alphabetical List
Using Paclets
Create a Paclet
Get Started
Download Definition Notebook
Learn More about
Wolfram Language
Parser
Tutorials
Building Language Front-Ends
Inside CodeAnalysis - How CodeStructure Parses C
Design and Compilation Strategy
Implementing the LaTeX Math Parser
MaTeX Comparison Showcase
The Parser Landscape - a Survey of What Exists Today
The Parser Zoo - language front-ends over a shared algebra
Parsing BNF Grammars (and bootstrapping a TPTP parser)
Parsing GrammarRules Locally
A Markdown Inline Parser in Parser Combinators
ParsingOpenQASM
Parsing TPTP, Auto-Generated from the Published BNF
PrattVsPEG
The Wolfram Box Typesetting Reference
Guides
Parsing in the Wolfram Language
Symbols
ASTAddSource
ASTAlgebra
ASTContainer
ASTLeafQ
ASTNodeQ
ASTStripSource
BinaryNode
BrainfuckAST
BrainfuckGrammar
BrainfuckRun
BrainfuckSemantic
CalculatorAST
CalculatorEval
CalculatorGrammar
CalculatorSemantic
CallNode
ContainerNode
EBNFParse
EBNFRules
ErrorNode
ExportLaTeX
GroupNode
InfixNode
JSONAST
JSONGrammar
JSONImport
JSONSemantic
LambdaAST
LambdaEval
LambdaGrammar
LambdaSemantic
LaTeXMathParse
LaTeXMathParser
LaTeXMathStyle
LeafNode
LispAST
LispGrammar
LispRead
LispSemantic
LispSymbol
MarkdownInlineParse
MarkdownInlineParser
MarkdownParse
MarkdownParser
ParseAction
ParseBetween
ParseChainLeft
ParseChainRight
ParseCharacter
ParseChoiceLongest
ParseChoice
ParseFail
ParseLiteral
ParseLookahead
ParseMany
Parse
ParseNotFollowedBy
ParseOperatorTable
ParseOptional
ParsePartial
ParsePosition
ParserCombinator
ParserCombinatorQ
ParserCompile
ParseRecursive
ParseRegex
ParseSepBy1
ParseSepBy
ParseSequence
ParseSome
ParseSucceed
ParseTry
PostfixNode
PrefixNode
RecCell
RecRef
SetRec
SpannedToken
TernaryNode
ToCodeParser
TPTPExport
TPTPImport
Overviews
WolframParser
Wolfram`Parser`
S
p
a
n
n
e
d
T
o
k
e
n
S
p
a
n
n
e
d
T
o
k
e
n
[
t
o
k
e
n
,
w
s
,
b
u
i
l
d
]
m
a
t
c
h
e
s
t
o
k
e
n
,
c
a
p
t
u
r
e
s
i
t
s
s
o
u
r
c
e
s
p
a
n
v
i
a
P
a
r
s
e
P
o
s
i
t
i
o
n
,
e
a
t
s
t
r
a
i
l
i
n
g
w
h
i
t
e
s
p
a
c
e
w
s
,
b
u
i
l
d
s
t
h
e
l
e
a
f
w
i
t
h
b
u
i
l
d
,
a
n
d
s
t
a
m
p
s
S
o
u
r
c
e
{
s
t
a
r
t
,
e
n
d
}
(
c
h
a
r
a
c
t
e
r
o
f
f
s
e
t
s
)
o
n
t
o
t
h
e
n
o
d
e
.
D
e
t
a
i
l
s
a
n
d
O
p
t
i
o
n
s
▪
S
p
a
n
n
e
d
T
o
k
e
n
is how the parser-zoo grammars give their leaves source spans. It brackets
token
with two
P
a
r
s
e
P
o
s
i
t
i
o
n
probes -
P
a
r
s
e
P
o
s
i
t
i
o
n
[
]
~
~
t
o
k
e
n
~
~
P
a
r
s
e
P
o
s
i
t
i
o
n
[
]
~
~
w
s
- then applies
build
to the matched text and
s
e
t
S
o
u
r
c
e
s the captured
{
s
t
a
r
t
,
e
n
d
}
onto the result.
▪
token
is the parser for the leaf's text (often a
P
a
r
s
e
R
e
g
e
x
or
P
a
r
s
e
L
i
t
e
r
a
l
);
ws
is the trailing-whitespace parser to consume
after
the span is closed (so the span covers the token only);
build
is the leaf constructor, e.g.
L
e
a
f
N
o
d
e
[
"
I
n
t
e
g
e
r
"
,
#
,
]
&
.
▪
The captured span is a raw
{
s
t
a
r
t
,
e
n
d
}
pair of
1-based character offsets
, where
end
is one past the last character
token
matched -
e
n
d
-
s
t
a
r
t
is the token's length. The trailing whitespace
ws
is eaten
outside
the two
P
a
r
s
e
P
o
s
i
t
i
o
n
s, so it never widens the span.
▪
A non-node
build
result passes through
unstamped
: when the grammar runs over a semantic algebra,
build
returns a value (a number, say) rather than a node, and there is nothing to carry
S
o
u
r
c
e
, so
S
p
a
n
n
e
d
T
o
k
e
n
returns it unchanged.
▪
The
{
s
t
a
r
t
,
e
n
d
}
offsets here are
not yet
line/column.
A
S
T
A
d
d
S
o
u
r
c
e
is the finalizer that spans the composites and converts every offset span to a
{
{
s
t
a
r
t
L
i
n
e
,
s
t
a
r
t
C
o
l
u
m
n
}
,
{
e
n
d
L
i
n
e
,
e
n
d
C
o
l
u
m
n
}
}
pair against the source.
Examples
(
3
)
Basic Examples
(
1
)
Match an integer token, skip trailing whitespace, and build a
L
e
a
f
N
o
d
e
carrying its span -
S
o
u
r
c
e
-
>
{
1
,
3
}
is the character offsets of
"
4
2
"
:
I
n
[
1
]
:
=
P
a
r
s
e
S
p
a
n
n
e
d
T
o
k
e
n
P
a
r
s
e
R
e
g
e
x
[
"
[
0
-
9
]
+
"
]
,
P
a
r
s
e
M
a
n
y
P
a
r
s
e
C
h
a
r
a
c
t
e
r
[
W
h
i
t
e
s
p
a
c
e
C
h
a
r
a
c
t
e
r
]
,
L
e
a
f
N
o
d
e
[
"
I
n
t
e
g
e
r
"
,
#
,
]
&
,
"
4
2
"
O
u
t
[
1
]
=
L
e
a
f
N
o
d
e
[
I
n
t
e
g
e
r
,
4
2
,
S
o
u
r
c
e
{
1
,
3
}
]
The span is
{
s
t
a
r
t
,
e
n
d
}
with
end
one past the last character, so
e
n
d
-
s
t
a
r
t
is the token length - here
3
-
1
=
2
, the two digits of
"
4
2
"
.
The trailing whitespace
ws
is consumed
outside
the span, so it never widens it - parsing
"
4
2
"
still yields the span
{
1
,
3
}
:
I
n
[
2
]
:
=
P
a
r
s
e
S
p
a
n
n
e
d
T
o
k
e
n
P
a
r
s
e
R
e
g
e
x
[
"
[
0
-
9
]
+
"
]
,
P
a
r
s
e
M
a
n
y
P
a
r
s
e
C
h
a
r
a
c
t
e
r
[
W
h
i
t
e
s
p
a
c
e
C
h
a
r
a
c
t
e
r
]
,
L
e
a
f
N
o
d
e
[
"
I
n
t
e
g
e
r
"
,
#
,
]
&
,
"
4
2
"
O
u
t
[
2
]
=
L
e
a
f
N
o
d
e
[
I
n
t
e
g
e
r
,
4
2
,
S
o
u
r
c
e
{
1
,
3
}
]
S
c
o
p
e
(
1
)
P
r
o
p
e
r
t
i
e
s
&
R
e
l
a
t
i
o
n
s
(
1
)
S
e
e
A
l
s
o
P
a
r
s
e
P
o
s
i
t
i
o
n
▪
A
S
T
A
d
d
S
o
u
r
c
e
▪
A
S
T
S
t
r
i
p
S
o
u
r
c
e
▪
L
e
a
f
N
o
d
e
▪
P
a
r
s
e
A
c
t
i
o
n
R
e
l
a
t
e
d
G
u
i
d
e
s
▪
P
a
r
s
e
r
Z
o
o
"
"