Wolfram Language Paclet Repository
Community-contributed installable additions to the Wolfram Language
Parser combinators for the Wolfram Language - GrammarRules compatible, locally compiled, with a LaTeX math parser
Contributed by: Nikolay Murzin, Claude (Anthropic)
The package provides Parse and ParserCompile as the entry points, ParserCombinator as the single computable head every constructor returns, and the Parse* family of constructors - ParseLiteral, ParseCharacter, ParseSequence, ParseChoice, ParseMany, ParseSome, ParseOptional, ParseBetween, ParseAction, ParseRecursive, ParseLookahead, ParseNotFollowedBy, ParseTry. GrammarRules is accepted as input to Parse and lowered locally. LaTeXMathParse parses LaTeX math-mode source to a tree of Wolfram boxes.
To install this paclet in your Wolfram Language environment,
evaluate this code:
PacletInstall["Wolfram/Parser"]
To load the code after installation, evaluate this code:
Needs["Wolfram`Parser`"]
A literal-string parser:
| In[1]:= |
| Out[1]= |
A one-or-more digit parser with an action that folds the captured digits into an integer:
| In[2]:= | ![]() |
| Out[2]= |
A GrammarRules slot template, parsed locally (no CloudDeploy round-trip):
| In[3]:= |
| Out[3]= |
LaTeXMathParse on an inline math source - the output is a tree of Wolfram boxes ready to drop into a notebook cell:
| In[4]:= |
| Out[4]= |
Wrap those boxes in RawBoxes to typeset them in a cell, and restyle with LaTeXMathStyle into the same Computer-Modern face LaTeX itself uses:
| In[5]:= |
| Out[5]= |
Every constructor returns a ParserCombinator; the wrapper carries the constructor's name, args, and an options Association. Parse interprets the tree, ParserCompile lowers it to a FunctionCompile / PEG-VM backend.
A bare-literal parser is the smallest non-trivial example - it matches its argument exactly and returns the matched string:
| In[6]:= |
| Out[6]= |
ParseCharacter matches one character against a pattern (a literal, an alternation, or a named character class):
| In[7]:= |
| Out[7]= |
ParseSequence runs combinators in order, returning the list of their results; ParseChoice returns the first one that succeeds (PEG-ordered):
| In[8]:= |
| Out[8]= |
ParseAction wraps a parser with a transformer; the transformer is splatted across the parser's result list, so ParseSome followed by StringJoin rejoins the matched chars:
| In[9]:= |
| Out[9]= |
ParseLookahead succeeds iff its argument would match - consuming nothing. ParseNotFollowedBy is its negation. Together they bound a body's content without committing to the delimiter:
| In[10]:= | ![]() |
| Out[10]= |
ParseRecursive defers binding until parse time, so a parser may name itself or its peers without pre-declaration. A balanced-parentheses parser is one line:
| In[11]:= | ![]() |
| Out[11]= |
A GrammarRules expression lowers to a ParserCombinator and runs locally - no CloudDeploy. The pattern form mirrors what the cloud accepts; the slot-template form (<name:Type>) is the convenient surface for word / digit captures:
| In[12]:= |
| Out[12]= |
The pattern form accepts the same shapes CloudDeploy'd GrammarRules does - FixedOrder, AnyOrder, OptionalElement, DelimitedSequence, RegularExpression, Pattern, and GrammarToken. The Parsing GrammarRules Locally tech note walks through every pattern.
ParserCompile lowers a ParserCombinator to a FunctionCompile function for the small / non-recursive case, or to a PEG-VM instruction table for large / recursive grammars (LaTeX, TPTP). The choice is opt-in via Method → "PEGVM":
| In[13]:= | ![]() |
| Out[13]= |
The compiled artifact can be Export'd to a WXF file and Import'd without recompiling - that is how LaTeXMathParse ships its compiled core (Assets/LaTeXMathParserCompiled.wxf).
ParserCompile takes a Method option choosing the compilation backend:
| In[14]:= |
| Out[14]= | ![]() |
Parse and ParsePartial take no user-facing options; the parser tree is the configuration.
LaTeXMathParse is the largest grammar in the paclet: a PEG over the full inline-math fragment of TeX. It handles \frac, \sqrt, sub/superscripts, \left/\right delimiters, \begin/end{matrix} environments, big operators with limits, and 40+ KaTeX macros. The build ships a WXF dump of the PEG-VM-compiled parser; first use auto-loads it if its grammar hash matches the source:
| In[15]:= |
| Out[15]= |
MarkdownInlineParse parses inline markdown - emphasis, code spans, math, links, sub/sup, escapes - to a tree of typed atoms (MdText, MdBold, MdMathInline, MdLink, …). The grammar is ~75 lines of ParseChoice over the primitives; see the Markdown inline parser tech note:
| In[16]:= |
| Out[16]= |
EBNFParse reads a BNF source (::=/:==/::-/::: arrow forms; the TPTP project's SyntaxBNF shape) and returns an Association of rule names to ParserCombinator. The BNF parser is itself written in the ParserCombinator core - a literal demonstration that the combinators are enough to parse their own meta-grammar:
| In[17]:= |
| Out[17]= | ![]() |
TPTPImport parses a TPTP problem file (.p) from the library - first-order, equational, higher-order, and typed fragments - returning a list of formulas. The parser handles the full TPTP grammar (10k+ lines of BNF) by composing the EBNFParse output with a small post-processing pass.
The combinators operate uniformly on strings, on lists of tokens, and on lists of Wolfram expressions, so the same primitives that lex a string can walk a CodeParser AST. The two domains share the implementation; only the leaf ParseCharacter / ParseLiteral definitions differ.
GrammarRules with GrammarApply requires a CloudDeploy round-trip; the cloud also resolves Interpreter-backed semantic types (City, Date, Quantity, …) via Wolfram knowledge. The local lowering trades the round-trip for offline evaluation - every documented pattern node is supported except the option flags (AllowLooseGrammar, IgnoreCase, IgnoreDiacritics). The Parsing GrammarRules Locally tech note maps the supported subset.
CodeParser parses Wolfram-language source to an AST; it is the target of a parser walk, not a combinator core itself. Wolfram`Parser`'s ParseRecursive + ParseChoice pair walks a CodeParser AST the same way it walks a token list - same primitives, different leaf type.
StringCases with StringExpression patterns is the WL idiom for one-shot string matching; it does not compose, does not capture nested structure, and does not return a typed parse tree. Wolfram`Parser` is the path to a real parse tree when StringCases outgrows the single-rule shape.
AntonAntonov/FunctionalParsers is the closest sibling - a Wolfram parser-combinator library predating this one. Differences: this paclet ships a FunctionCompile / PEG-VM compile path, GrammarRules compatibility, and the LaTeX / TPTP / Markdown front-ends out of the box.
ParseChoice is PEG-ordered: alternatives are tried left-to-right, and the first match wins (no longest-match backtracking). "north" | "northwest" matches "north" on input "northwest" and leaves "west" unconsumed. Order alternatives longer-first, or use ParseChoiceLongest when POSIX longest-match is required:
| In[18]:= |
| Out[18]= |
FunctionCompile inlines combinator graphs, so a recursive grammar (LaTeX, TPTP) hits the inliner's size cap and either compiles slowly or aborts. Use Method → "PEGVM" for those: the PEG-VM lowers to an integer instruction table that is recursive at runtime, not at compile time.
A GrammarToken whose type is not Number / Integer / Word / Automatic calls out to Interpreter at parse time. City, Country, Quantity, Date, … need the Wolfram knowledge engine reachable. Tests that depend on these are not offline-stable.
Parse requires the entire input to match the grammar. A prefix match returns a Failure with Expected → "<end of input>". Use ParsePartial to get {result, leftover} for prefix-matching grammars (REPL prompts, line-based protocols).
EBNFParse can parse the BNF that describes its own grammar. The library's EBNFParse definition (in Kernel/EBNF.wl) is built from <rule> ::= <name> ::= <body>-shaped patterns; running the parser on a BNF source that describes that very shape is the smallest non-trivial self-host.
A one-character math source resolves to its Wolfram named character:
| In[19]:= |
| Out[19]= |
Each \command resolves to its Wolfram named character - the boxes hold "α", "β", "γ", so RawBoxes typesets them as the Greek letters. The command-to-glyph table covers the full Greek alphabet plus a long tail of math symbols.
The recursive label parser in MarkdownInlineParse handles a code-styled label inside a link, producing one MdLink whose label is itself a parsed-list:
| In[20]:= |
| Out[20]= |
The "appliance controller" grammar fits in one declarative expression with no helper definitions - the same shape that round-trips through CloudDeploy runs locally with Parse:
| In[21]:= | ![]() |
| Out[21]= |
Wolfram Language Version 14.0