# Parsing Pipeline

This page walks through the steps `transToHTML()` runs to turn Markdown into HTML, and why their order shapes the output.

## Signature

```javascript
transToHTML(body = "", path = "", target = "_blank", standard = false)
```

| Parameter | Meaning |
|---|---|
| `body` | Markdown source; all three components wrap it in `\n` before calling |
| `path` | Hashtag link prefix; an empty string disables hashtag conversion |
| `target` | Hashtag link target; only `"_blank"` is kept, any other value becomes `_self` |
| `standard` | When `true`, extended syntax is skipped; see [Standard Mode](/standard-mode) |

## Stages

| Stage | Action |
|---|---|
| 1. Escape pre-processing | `\!`, `` \` ``, `\#`, `\*`, `\_`, `\~`, `\^`, `\=`, `\<`, `\>`, `\[`, `\]`, `\(`, `\)` become internal markers such as `@excl@`; `$` becomes `@dollar@`; non-breaking spaces (U+00A0) are normalized to plain spaces |
| 2. Syntax transforms | The 16 `set*` functions in the table below run in order |
| 3. Placeholder restore | `{{tag-uuid}}` markers are replaced with their HTML repeatedly until none remain (nested fragments need several passes) |
| 4. Block whitespace cleanup | Newlines around the opening and closing tags of `h1`-`h6`, `table`, `ol`, `ul`, `pre`, `blockquote`, `details`, `hr`, `label` are removed |
| 5. Escape restore | Markers become HTML entities (`&excl;`, `&num;`, `&ast;`, ...); one or more consecutive blank lines become a single `<br>` |

## Transform Order

| Order | Function | Handles | Why here |
|---|---|---|---|
| 1 | `setPreCode` | Fenced code blocks, Mermaid | Isolated first so inner syntax is never misread |
| 2 | `setCode` | Inline code | Must be isolated before font formatting |
| 3 | `setMedia` | Images (with size and alignment), `.mp4` / `.mov` video | `![]()` resembles a link `[]()` and must match first |
| 4 | `setMediaDefault` | Images, video (no size) | Image handling for standard mode |
| 5 | `setLink` | Links, YouTube/Vimeo embeds, bare URLs, email | |
| 6 | `setLinkStandard` | Links, bare URLs, email (no embeds) | Link handling for standard mode |
| 7 | `setFont` | Bold, italic, strikethrough, marker, super/subscript | |
| 8 | `setFontStandard` | Bold, italic, strikethrough | Font handling for standard mode |
| 9 | `setHeading` | `#` headings, setext headings, headings inside table cells | Must precede lists (list items may contain `#`) |
| 10 | `setHr` | `---`, `***` | Must precede tables so `---` is not read as a table separator |
| 11 | `setTable` | Tables | |
| 12 | `setBlockquote` | Blockquotes, GitHub-style alerts | |
| 13 | `setBlockquoteStandard` | Blockquotes (no alerts) | Blockquote handling for standard mode |
| 14 | `setList` | Ordered/unordered lists, checkboxes | |
| 15 | `setTabCode` | Indented code blocks | Last, to avoid clashing with nested lists |
| 16 | `setHashtag` | `#tag` links | The final text transform |

Extended mode (`standard: false`) also runs the standard variants, but the extended function has already replaced the syntax with a placeholder, so the standard one finds nothing left to match.

## Placeholder Mechanism

`replaceWithUUID()` stores the converted fragment's `outerHTML` in `elementUUIDMap` and swaps the source text for `{{tag-<32 random chars>}}`. Later regexes never match a placeholder, which is why `**` inside a code block does not turn bold.

`elementUUIDMap` is a module-level `Map` that is never cleared; during long sessions of continuous rendering, old fragments stay in memory.

## Known Limitations (v1.11.6)

| Symptom | Cause |
|---|---|
| A `$` in the text is output as `@dollar@` | Stage 1 turns `$` into `@dollar@`, but the stage 5 restore table maps `@excl@` to `&dollar;` by mistake, so `@dollar@` is never restored |
| Raw HTML passes through unchanged | The pipeline does not sanitize; see [Security](/security) |
