Architecture

Five pure stages, three trait seams, a workspace of hosts around them, and two ratchets that keep the port honest.

The pipeline

parsevalidatetransformrenderformatASTcanonical MarkdocdiagnosticsHTMLsource textrenderable tree

The AST is the hub, and that is the shape worth noticing: validate, transform and format each read it independently. format takes an AST, not a renderable tree, so a formatter never runs the middle of the pipeline. And validate produces diagnostics beside the tree rather than gating it -- a document with errors still transforms and still renders.

Each stage is a pure function of its inputs, and each is a separate entry point. A linter stops after validate. An editor stops after transform and maps the tree onto its own components. A static site generator runs the lot. A formatter skips the middle entirely and goes from AST straight back to source.

Nothing in the crate holds state between calls, and output is deterministic: attribute order is authored order, never hash order, so two runs over the same input produce identical bytes.

What the crate refuses to do

No I/O. No configuration file. No concept of a file, a theme, a template or a plugin. Everything host-specific arrives as data the caller passes in, or through a trait the caller implements.

parse validate transform render the host a CommonMark impl Tokenizer segmenter a trait schemas: a constant, impl SchemaSource YAML, a database, a sandboxed guest a trait your elements, impl TagRenderer your escaping, your void list a trait accent-proust

This is not minimalism for its own sake. It is what makes the crate embeddable in a CMS, a language server and a build tool without any of the three inheriting the others' assumptions -- and it is enforced by a CI job rather than by discipline.

The three traits: Tokenizer, TagRenderer and SchemaSource

Tokenizer segments CommonMark. A default implementation over pulldown-cmark ships behind the pulldown-cmark-tokenizer feature, which is on by default and can be turned off:

accent-proust = { version = "*", default-features = false }

A host that already parses CommonMark supplies its own rather than compiling a second one into the same binary. The seam earns its keep in a specific case: if you pin pulldown-cmark to a git revision, Cargo treats it as a different package, and you would end up rendering some documents through one copy and some through the other.

TagRenderer turns a tag into markup. render::Html is upstream's implementation and what render_all uses; a host that wants different elements, different escaping or a different void-element list implements the trait and calls render_with. The trait is open, close and text, and deliberately not "a tag and a callback for its children": the crate walks the tree on an explicit heap stack because nesting depth is the document's to choose, and a callback would put that depth back on the host's stack one frame per level. The host writes the markup for one tag at a time and never sees the recursion.

SchemaSource answers "what is the schema for this tag or node type?", and Config holds one implementation of it as the only place schemas come from. MapSchemaSource is the implementation a host fills by hand -- two maps, with builtin() starting them at Markdoc's own vocabulary -- and a host whose schemas live somewhere else implements the trait instead. Whether they came from a constant, a YAML file, a database or a sandboxed guest is the implementation's business, and the crate never learns. The trait returns a borrow, which rules out loading a schema on first request by design: a host populates its source, then hands it over.

Invariant 1: the library stands alone

scripts/check-standalone.sh fails the build if the library's manifest gains a path or git dependency, or if a host crate appears anywhere in its tree. A separate CI job builds and tests the crate with --no-default-features, so the no-tokenizer configuration is supported rather than tolerated.

The release script runs the same check before publishing, because a path dependency that reaches crates.io is the class of mistake that cannot be undone once a version is on the index.

Layout mirrors upstream

The module tree matches upstream file-for-file wherever the code is a pure function of its inputs, so a future upstream commit diffs cleanly against its Rust counterpart.

UpstreamHere
src/ast/src/ast/
src/grammar/tag.pegjssrc/grammar/
src/parser.ts, src/tokenizer/src/parse/
src/validator.ts, src/schema.ts, src/schema-types/src/validate/
src/transformer.ts, src/transforms/src/transform/
src/renderers/html.tssrc/render/
src/formatter.tssrc/format/
src/functions/src/functions/
src/tags/src/tags/

Two upstream trees are vendored at the ported revision and never edited:

PathWhatWhy
spec/The conformance corpus and its runnerIt is the test suite, so a fresh clone runs it with cargo test and nothing else
reference/Upstream's TypeScript, its unit tests, and its markdown-it patchA porting pull request shows its source in the same diff, and the yearly upstream refresh is a git diff rather than a second checkout

reference/ is excluded from the published crate -- 392 KB of another language is useful in the repository and dead weight in a download. spec/ is not excluded: a package that cannot run its own tests is the worse trade.

The workspace

The library is the workspace root package and stays at the repository root. Members are hosts:

MemberWhat
crates/accent-proust-wasmWebAssembly bindings for a browser or other JavaScript host. Ships to npm rather than crates.io, so it sets publish = false
crates/accent-proust-cliThe command-line host, a binary named accent-proust: fmt, validate, render, transform and parse. publish = false until it has a release cadence of its own
crates/accent-proust-schema-configThe declarative schema vocabulary both hosts read: the keys, the refusal of an unknown one with the path to it, and the mapping onto Schema. Not a host; what the hosts share

A binding that carries the library across an ABI is a host in the same sense a CMS is, and so is a binary that reads files; each gets a crate beside the library rather than a feature inside it -- the same reasoning that keeps Tokenizer a trait rather than an implementation. The vocabulary crate is not a host but what two hosts share, and it lives beside them so that a schema declared for the browser is accepted by the shell -- one from an object, the other from a file -- and the two cannot drift apart. default-members = ["."] holds a bare cargo build, cargo test and cargo clippy --all-targets to the library alone, so no member can quietly join the standalone, MSRV or conformance lanes; each brings its own CI job.

The two ratchets

Both fail on drift in either direction, and both are the point of the project rather than bookkeeping.

Conformance. conformance-baseline.txt records where the corpus counter stood. cargo test --test conformance compares the run against it and fails on any difference: a drop is a regression and is not mergeable, a rise is a baseline that was not updated in the same commit. It is deliberately not an absolute 105/105 gate, which would leave every pull request failing a required check until the port finished.

conformance: 95 green, 10 annotated, 0 failing (of 105)

Divergences. A case that should stop being green is a divergence, which means an entry in DIVERGENCES.md and a move from green to annotated -- not a smaller number in the baseline. That file is normative, not a changelog.

Upstream error ids are the one place where divergence is disallowed outright rather than declared, because external tooling binds to them.

Panic-freedom

unsafe_code is forbidden and panic, unwrap, expect and indexing_slicing are denied in CI. An #[allow] with a comment saying why the bound is proven is the intended escape hatch; an #[allow] without one is the thing being prevented.

The promise covers values a caller builds as well as documents the crate parses, and it covers every way of touching one. Each public recursive type -- ast::Node, ast::Value, renderable::Tag and renderable::Scalar -- writes out all four of its traversals by hand: Drop, Clone, PartialEq and Debug. A derived implementation of any of them recurses once per level, and a stack overflow aborts rather than panics, so a caller could otherwise kill the process with a value it assembled through the public API. Drop and PartialEq walk a worklist, Clone walks post-order onto a plan and rebuilds bottom-up, and Debug emits from a token stack. None of it is unsafe.

Property tests assert the parser never panics on arbitrary input, and fuzzing precedes publication.