Wolfram Language Paclet Repository

Community-contributed installable additions to the Wolfram Language

Primary Navigation

    • Cloud & Deployment
    • Core Language & Structure
    • Data Manipulation & Analysis
    • Engineering Data & Computation
    • External Interfaces & Connections
    • Financial Data & Computation
    • Geographic Data & Computation
    • Geometry
    • Graphs & Networks
    • Higher Mathematical Computation
    • Images
    • Knowledge Representation & Natural Language
    • Machine Learning
    • Notebook Documents & Presentation
    • Scientific and Medical Data & Computation
    • Social, Cultural & Linguistic Data
    • Strings & Text
    • Symbolic & Numeric Computation
    • System Operation & Setup
    • Time-Related Computation
    • User Interface Construction
    • Visualization & Graphics
    • Random Paclet
    • Alphabetical List
  • Using Paclets
    • Get Started
    • Download Definition Notebook
  • Learn More about Wolfram Language

Parser

Tutorials

  • Building Language Front-Ends
  • Inside CodeAnalysis - How CodeStructure Parses C
  • Design and Compilation Strategy
  • Implementing the LaTeX Math Parser
  • MaTeX Comparison Showcase
  • The Parser Landscape - a Survey of What Exists Today
  • The Parser Zoo - language front-ends over a shared algebra
  • Parsing BNF Grammars (and bootstrapping a TPTP parser)
  • Parsing GrammarRules Locally
  • A Markdown Inline Parser in Parser Combinators
  • ParsingOpenQASM
  • Parsing TPTP, Auto-Generated from the Published BNF
  • PrattVsPEG
  • The Wolfram Box Typesetting Reference

Guides

  • Parsing in the Wolfram Language

Symbols

  • ASTAddSource
  • ASTAlgebra
  • ASTContainer
  • ASTLeafQ
  • ASTNodeQ
  • ASTStripSource
  • BinaryNode
  • BrainfuckAST
  • BrainfuckGrammar
  • BrainfuckRun
  • BrainfuckSemantic
  • CalculatorAST
  • CalculatorEval
  • CalculatorGrammar
  • CalculatorSemantic
  • CallNode
  • ContainerNode
  • EBNFParse
  • EBNFRules
  • ErrorNode
  • ExportLaTeX
  • GroupNode
  • InfixNode
  • JSONAST
  • JSONGrammar
  • JSONImport
  • JSONSemantic
  • LambdaAST
  • LambdaEval
  • LambdaGrammar
  • LambdaSemantic
  • LaTeXMathParse
  • LaTeXMathParser
  • LaTeXMathStyle
  • LeafNode
  • LispAST
  • LispGrammar
  • LispRead
  • LispSemantic
  • LispSymbol
  • MarkdownInlineParse
  • MarkdownInlineParser
  • MarkdownParse
  • MarkdownParser
  • ParseAction
  • ParseBetween
  • ParseChainLeft
  • ParseChainRight
  • ParseCharacter
  • ParseChoiceLongest
  • ParseChoice
  • ParseFail
  • ParseLiteral
  • ParseLookahead
  • ParseMany
  • Parse
  • ParseNotFollowedBy
  • ParseOperatorTable
  • ParseOptional
  • ParsePartial
  • ParsePosition
  • ParserCombinator
  • ParserCombinatorQ
  • ParserCompile
  • ParseRecursive
  • ParseRegex
  • ParseSepBy1
  • ParseSepBy
  • ParseSequence
  • ParseSome
  • ParseSucceed
  • ParseTry
  • PostfixNode
  • PrefixNode
  • RecCell
  • RecRef
  • SetRec
  • SpannedToken
  • TernaryNode
  • ToCodeParser
  • TPTPExport
  • TPTPImport

Overviews

  • WolframParser

A Markdown Inline Parser in Parser Combinators

What this note covers
Part 3 - Recursion and post-processing
Part 1 - The AST
Part 4 - The full document parser
Part 2 - The grammar
Why this matters

What this note covers

MarkdownInlineParse
is a parser for inline markdown - the inside-a-paragraph constructs of CommonMark: emphasis (
*italic*
,
**bold**
,
***both***
,
~~strike~~
), code spans (
`code`
,
​
literal
​
,
<code>...</code>
), inline math (
$...$
,
$$...$$
), links (
[label](url)
), images (
![alt](url)
), HTML and Pandoc sub/sup (
<sub>x</sub>
,
H~2~O
,
e^2^
), backslash escapes (
\*
,
\$
,
\\
), and underscore emphasis with CommonMark word-boundary rules (
_em_
,
__strong__
, but not
snake_case
).
It returns a flat list of inline atoms; each atom is an
Association
with a
"Type"
discriminator plus payload keys - the same convention
MarkdownParse
and M2N's block-level parser use, so downstream code can pattern-match on
"Type"
uniformly across both layers.
The whole grammar is ~75 lines of
ParseChoice
/
ParseAction
/
ParseNotFollowedBy
over the
Wolfram`Parser`
primitives. No
StringSplit
, no regex hand-tuning, no per-rule iteration. The PEG order in one
ParseChoice
resolves the precedence ambiguities (
**bold**
wins over
*italic*
opened twice;
***both***
wins over
*
-
**both**
-
*
).
This note has three parts:
1
.
The AST - what
MarkdownInlineParse
returns.
2
.
The grammar - line-by-line walkthrough of the combinator chain.
3
.
Recursion and post-processing - how nested emphasis (
**bold$x$**
) gets re-parsed and how the CommonMark word-boundary rules for underscore emphasis fall out as a small post-pass.

Part 1 - The AST

MarkdownInlineParse
[source]
returns a
List
of inline atoms. Each atom is an
Association
with a
"Type"
discriminator and per-shape payload keys:
"Type"
Other keys
Markdown form
"Text"
"Text"->str
plain text run
"Code"
"Code"->str
`code`
"LiteralCode"
"Code"->str
​
code
​
(verbatim)
"HtmlCode"
"Code"->str
<code>...</code>
"MathInline"
"Math"->str
$x$
"MathDisplay"
"Math"->str
$$x$$
"Link"
"Label"->[atoms]
,
"Url"->str
[label](url)
"Image"
"Alt"->str
,
"Url"->str
![alt](url)
"Sub"
"Children"->[atoms]
<sub>x</sub>
or
~x~
"Sup"
"Children"->[atoms]
<sup>x</sup>
or
^x^
"Bold"
"Children"->[atoms]
**x**
or
__x__
"Italic"
"Children"->[atoms]
*x*
or
_x_
"BoldItalic"
"Children"->[atoms]
***x***
"Strike"
"Children"->[atoms]
~~x~~
The
"Children"
/
"Label"
for spans are themselves
List
s of inline atoms (the body was re-parsed), so
**bold$x$**
lands as
<|"Type"->"Bold","Children"->{<|"Type"->"Text","Text"->"bold "|>,<|"Type"->"MathInline","Math"->"x"|>}|>
. Adjacent
"Text"
atoms are coalesced so consumers see one run per contiguous prose chunk, never one atom per character.
Using
Association
s with a
"Type"
key instead of a sum-type-headed expression matches the convention M2N's block-level parser uses (
<|"Type"->"Heading","Level"->n,"Text"->str|>
etc.), so downstream code reading the parse tree can pattern-match on
atom["Type"]
uniformly across the inline and block layers without learning two AST shapes.
Plain prose comes out as a single
"Text"
atom:
In[1]:=
MarkdownInlineParse
["plain text"]
Out[1]=
{TypeText,Textplain text}
Mixed prose, emphasis, math, and code:
In[2]:=
MarkdownInlineParse
["**bold $x$** and `code`"]
Out[2]=
{TypeBold,Children{TypeText,Textbold ,TypeMathInline,Mathx},TypeText,Text and ,TypeCode,Codecode}
A link with a code-styled label demonstrates recursive label parsing - the label itself is a list of inline atoms:
In[3]:=
MarkdownInlineParse
["[`Range`](paclet:ref/Range)"]
Out[3]=
{TypeLink,Urlpaclet:ref/Range,Label{TypeCode,CodeRange}}

Part 2 - The grammar

The whole parser lives in
examples/WolframParser/Kernel/Markdown.wl
. The grammar is one big
ParseChoice
and lots of small
ParseAction
arms.

The bounded-content helper

Every paired-delimiter span needs to consume "everything until the closing delimiter". The trick is
ParseNotFollowedBy
: at each position, refuse to consume a character if the next characters would form the closing delimiter.
In[4]:=
content[term_]:=
ParseAction
​​
ParseSome

ParseAction

ParseNotFollowedBy
[term]~~anyChar,#2&,​​StringJoin[{##}]&​​
Reading the body:
anyChar
is
ParseCharacter
[_]
;
ParseNotFollowedBy
[term]
is a zero-width assertion that the next characters do not start
term
. Their sequence consumes one char only when
term
doesn't start there.
ParseSome
iterates the assertion-plus-char until
ParseNotFollowedBy
fires, at which point
ParseSome
stops and the outer
ParseAction
joins the accumulated chars into a string.
content[ParseLiteral["**"]]
matches
"foo $x$"
and stops cleanly at the
**
of
"foo $x$**"
- which is exactly the bound a
**foo$x$**
body needs.

Paired spans

Every paired span is a literal-open /
content[close]
/ literal-close trio. The grammar reads a few private helper constructors (
text
,
codeAtom
,
mathIn
,
bold
, …) that each build the corresponding
"Type"
-tagged
Association
- keeping the AST construction in one place so the grammar arms stay readable.
Inline code:
In[5]:=
code=
ParseAction
​​
ParseLiteral
["`"]~~content
ParseLiteral
["`"]~~
ParseLiteral
["`"],​​codeAtom[#2]&​​
$IterationLimit::itlim:Iterationlimitof4096exceeded.
$IterationLimit::itlim:Iterationlimitof4096exceeded.
Out[5]=
ParseAction[TerminatedEvaluation[IterationLimit],codeAtom[#2]&]
Inline math:
In[6]:=
inlineMath=
ParseAction
​​
ParseLiteral
["$"]~~content
ParseLiteral
["$"]~~
ParseLiteral
["$"],​​mathIn[#2]&​​
$IterationLimit::itlim:Iterationlimitof4096exceeded.
$IterationLimit::itlim:Iterationlimitof4096exceeded.
Out[6]=
ParseAction[TerminatedEvaluation[IterationLimit],mathIn[#2]&]
Bold:
In[7]:=
boldP=
ParseAction
​​
ParseLiteral
["**"]~~content
ParseLiteral
["**"]~~
ParseLiteral
["**"],​​bold[#2]&​​
$IterationLimit::itlim:Iterationlimitof4096exceeded.
$IterationLimit::itlim:Iterationlimitof4096exceeded.
Out[7]=
ParseAction[TerminatedEvaluation[IterationLimit],bold[#2]&]
The
#2&
arm takes the second result (the content; the first is the open literal, the third is the close) and wraps it in the appropriate atom constructor.

Escapes and the plain-character catch-all

A backslash followed by ASCII punctuation collapses to that punctuation, so
\*
becomes a literal
*
instead of opening an italic span:
In[8]:=
escape=
ParseAction
​​
ParseLiteral
["\\"]~~charSat[asciiPunct],​​text[#2]&​​
$IterationLimit::itlim:Iterationlimitof4096exceeded.
Out[8]=
ParseAction[TerminatedEvaluation[IterationLimit],text[#2]&]
After every other alternative has had a chance, the catch-all
plainChar
consumes one character and wraps it in a
"Text"
atom:
In[9]:=
plainChar=
ParseAction
[anyChar,text[#1]&]
Out[9]=
ParseAction[anyChar,text[#1]&]
ParseSome
over the outer
ParseChoice
iterates this until the input is exhausted; the post-process step
mergeText
coalesces consecutive
"Text"
atoms so prose comes out as one run, not one atom per character.

Links and images

The label can contain markdown itself, but capturing it character-by-character at the top level would be wrong (we'd lose the structure). Instead capture the raw label string during parse and re-parse it after:
In[10]:=
linkP=
ParseAction
​​
ParseLiteral
"["]~~linkLabel~~
ParseLiteral
["]("]~~linkUrl~~
ParseLiteral
[")",​​link[#2,#4]&​​
$IterationLimit::itlim:Iterationlimitof4096exceeded.
Out[10]=
ParseAction[TerminatedEvaluation[IterationLimit],link[#2,#4]&]
linkLabel
is
ParseSome
[ParseNotFollowedBy[ParseLiteral["]"]]~~anyChar]
joined to a string - similar to
content[]
, but the body is "any character that isn't
]
".
linkUrl
is the same shape with
)
. The
link[#2,#4]
action stores them; recursive re-parsing of the label happens in Part 3.
imageP
is the same with a leading
!
, listed first in the
ParseChoice
so
![
opens an image and not a link with
!
prefix.

PEG ordering: longer prefixes first

The full alternation:
In[11]:=
inlineAtom=
ParseChoice
[​​escape,(*\x*)​​codeHtml,imageP,linkP,(*HTML/[/![*)​​dblCode,code,(*```*)​​displayMath,inlineMath,(*$$$*)​​strikeP,(*~~*)​​htmlSub,htmlSup,(*<sub><sup>*)​​boldItalicP,boldP,italicAst,(********)​​pandocSub,pandocSup,(*~^*)​​plainChar(*fallback*)​​]
Within each opening-character family the longest opener comes first:

Word-bounded asterisk italic

Part 3 - Recursion and post-processing

Recursive children

Adjacent-text merging

Underscore emphasis (CommonMark word boundaries)

Part 4 - The full document parser

Why this matters

© 2026 Wolfram. All rights reserved.

  • Legal & Privacy Policy
  • Contact Us
  • WolframAlpha.com
  • WolframCloud.com