The Parser Landscape - a Survey of What Exists Today
Why this note exists
Parsing is a long-solved problem - in most languages. Haskell has Parsec. Rust has nom. Python has pyparsing. Lua has LPeg. Even modest scripting languages tend to ship at least one composable parser library that handles strings, tokens, and structured trees uniformly.
The Wolfram Language is an unusual exception. It has many partial answers -
StringExpression
,
RegularExpression
,
Interpreter
,
GrammarRules
,
CodeParser
,
XMLObject
- but no single library that lets a user build a custom parser of arbitrary complexity, locally, with the kind of compositional ergonomics Parsec made famous. This note surveys what's actually available, in order to be honest about why
WolframParser
is being written: not because nothing exists, but because what exists is fragmented along axes that don't compose.
The survey is structured into three parts:
1
.
Wolfram built-ins - what ships with the kernel.
2
.
Community libraries - what's on the Paclet Repository / GitHub.
3
.
Outside Wolfram - the design heritage
WolframParser
borrows from.
Each section says what the tool does well, where it stops, and how
What it does well. Concise, fast (C-backed), integrates with everything in the string namespace, expressive for regular grammars. Captures via
x:pat
are clean. The
Shortest
/
Longest
modifiers handle backtracking control.
Where it stops.
StringExpression
is fundamentally a regex - it cannot express recursion or balanced brackets. A
Shortest["("~~___~~")"]
will happily match
(a(b)
:
In[2]:=
StringCases["(a(b)c)",Shortest["("~~___~~")"]]
Out[2]=
{(a(b)}
There is also no native way to attach actions to subpatterns and assemble a structured result, other than collecting the substring captures and post-processing them yourself. That works for simple cases and creaks for grammars with more than a handful of productions.
RegularExpression
A thin PCRE wrapper. Useful when you have a regex from somewhere else and need to drop it in unchanged. The same limitations as
returns a function that tries to interpret an input string as something of the given type. Built-in types are extensive - numbers, dates, colors, locations, units, even free-form English noun phrases:
In[4]:=
Interpreter["Color"]["skyblue"]
Out[4]=
Restricted[type,constraints]
narrows the interpretation:
In[5]:=
Interpreter[Restricted["Integer",{0,100}]]["50"]
Out[5]=
50
What it does well. Outstanding for "extract a typed value from text" tasks where the type is one of the built-in entity / quantity / data interpretations. Often the right answer for input forms, configuration parsing, and natural-language commands.
Where it stops. Not user-extensible in any deep sense - the type system is closed. You can compose interpreters at the list level (
Interpreter[{type1,type2}]
) but not write your own type with custom production rules. Failure modes are opaque - you get a
Failure[...]
object with limited diagnostic information.
GrammarRules
,
GrammarApply
,
GrammarToken
The cleanest WL-native API for higher-level grammars. A grammar is a list of rules with slot syntax, and
GrammarApply
walks the input against the rules to extract structured matches:
is the most ergonomic WL grammar surface, and it is also the one with the hardest dependency: you can use it in production only if your workflow tolerates round-tripping through a CloudObject. For a paclet that wants to embed a grammar in a piece of local computation - parsing a config file, lexing a DSL, post-processing CodeParser output - this is a non-starter.
The design of
GrammarRules
, however, is the right shape.
WolframParser
borrows the slot-rule notation directly, and runs it locally.
CodeParser
A first-party paclet (
CodeParser\
`) that tokenises and parses Wolfram Language source code into a typed AST. The implementation is in C with a thin WL surface, so it is very fast and very precise about source positions:
What it does well. The gold standard for WL source - used by the front-end, the linter, the formatter, and a small ecosystem of paclets. The AST is well-typed and easy to walk.
Where it stops. It parses one language - Wolfram. There is no way to retarget it to parse JSON, TOML, a custom DSL, or even a near-relative like a
.wlt
test file. The C backend means you cannot add productions; you can only consume what the parser emits.
WolframParser
does not try to compete with
CodeParser
for WL source - on the contrary, it interoperates: you can feed a
CodeParser
AST into a
WolframParser
grammar that walks the tree and extracts higher-level structure (e.g. "find all
Module
definitions with a particular shape"). That dual-mode of "string parser" + "expression-tree parser" is what the token-oriented design enables.
ResourceFunction["CodeStructure"]
- the
CodeAnalysis
paclet
The
CodeParser
story, one language over: a first-party parser for C/C++ source, delivered through the Function Repository.
ResourceFunction["CodeStructure"]
is itself only a loader shim - its entire definition is three lines that install and load the
paclet (v0.9.6 at writing), which parses C by shelling out to Clang and post-processing its AST dump into a Wolfram expression tree. The options give the backend away:
, etc. A second argument picks an alternative representation -
"SyntaxTree"
,
"SyntaxAnnotation"
,
"SourceAnnotation"
,
"TokenAnnotation"
,
"CallGraph"
,
"FileCallGraph"
- and the companion
CodeCases
walks a
CodeElement
tree the way
Cases
walks any expression.
What it does well. A precise, source-accurate C/C++ AST with zero Clang plumbing on your part - the paclet handles the install, the compiler invocation, and the AST-dump parsing. The right tool for "find every function that calls
malloc
", call-graph extraction, and source-to-source tooling over C.
Where it stops. It parses C and only C, through an external Clang it shells out to: you cannot retarget it, add productions, or run it on a machine with no Clang on the path. Like
CodeParser
, it is an analyzer for one fixed language rather than a construction kit - and, as with
CodeParser
, the part that matters to this survey is the structured-AST output. A
CodeElement
tree is exactly what
WolframParser
's expression-tree input mode is built to walk: the same combinators that lex a string can match
CodeElement[_,"FunctionDecl",_]
nodes to pull out higher-level structure without re-parsing the C.
XMLObject
,
JSON
,
ImportString
ImportString[text,"XML"]
,
ImportString[text,"JSON"]
, and friends cover the cases where someone else has already done the parsing work for a well-known format. They are not really parsers in the construction sense - they are importers. Useful when applicable, irrelevant when you have a format that doesn't already have an importer.
Part 2 - Community libraries
AntonAntonov/FunctionalParsers
The most complete pure-WL parser combinator library on the Paclet Repository. Anton Antonov has been working on it for years; the API is a faithful Wolfram port of the Haskell
What it does well. Genuinely complete - every combinator a Parsec user would expect is there. The EBNF generator means you can write a grammar declaratively and get a working parser without hand-wiring combinators. It is pure WL and runs locally. Several large projects (especially Anton's own NLP-oriented work) use it in production.
Where it stops.
◼
Naming overhead.
ParseSequentialCompositionPickRight
is descriptive and unambiguous, but a parser line full of
. A short, combinable operator vocabulary is part of the ergonomics story and is not present here.
◼
Speed. Pure interpretation - every combinator allocates closures and walks lists. Acceptable for grammar-sized inputs (kilobytes), painful for megabyte-scale inputs. There is no compilation path.
◼
Token shape. The library always tokenises first via
ToTokens
. Operating directly on a string (without explicit tokenisation) or on a list of Wolfram expressions (e.g. a
CodeParser
AST) is not idiomatic.
◼
Diagnostics. Failure is a missing match - you get
{}
back. There is no positional error message of the form "expected
)
at line 3 col 12, saw
]
".
WolframParser
aims to share the combinator vocabulary, but add an operator surface, a token-or-string-or-expression-uniform input model, and structured parse-failure diagnostics.
Other paclet-repository entries
A handful of related entries on the Paclet Repository -
WolframLanguageExtras
,
JSONParser
,
XMLConverter
, etc. - target specific formats rather than the general-purpose niche. None compete with
FunctionalParsers
for the role of "general parser combinator library."
Part 3 - Outside Wolfram (design heritage)
This is the literature
WolframParser
reads to figure out what good looks like.
Parsec (Haskell)
Daan Leijen's
Parsec
(2001) is the reference design. Two key ideas have spread to nearly every modern parser library:
1
.
A parser is a function (or monadic value) that consumes input and returns either a result with the remainder, or a failure with diagnostic information.
2
.
Bigger parsers are built from smaller ones with a tiny vocabulary of combinators: alternation (
<|>
), sequence (
<*>
/
>>
), zero-or-more (
many
), one-or-more (
many1
), optional (
optional
), backtracking control (
try
).
Parsec's brilliance is the applicative shorthand: most useful parsers don't need full monadic power; they look like
expr = (+) <$> term <* char '+' <*> expr
PEG - Parsing Expression Grammars (Ford, 2004)
Bryan Ford's PEG formalism is the other big design tradition. A PEG looks similar to a context-free grammar but with two key differences:
The combination is expressive enough for most real-world languages, parses in linear time with packrat memoisation, and gives unambiguous parse trees by construction. Implementations: LPeg (Lua, by Roberto Ierusalimschy - widely cited as one of the best parser libraries in any language), pest (Rust), PEG.js (JavaScript).
ANTLR / LALR generators
Earley / GLR / GLL
Cross-reference table
The shape of WolframParser
The survey above defines the niche by exclusion. Concretely, the paclet aims to provide:
What the paclet is not trying to be:
◼
an ANTLR (the generator-based heavyweight school is a different ecosystem)
◼
a general CFG parser (Earley / GLR / GLL backends are out of scope for v0.1)
◼
a tokeniser for any specific format (those belong in companion paclets that use this library to define their lexers)