Sex documentation engine survey #32

Open
opened 2026-09-22 09:47:14 +02:00 by pkulev · 0 comments
Collaborator

We will keep in mind that agents need some kind of docs extracted from humans human docs too. Yes?

Sex has single documentation surface: Readme.org. Ground truth is tests/ and example/*.sex. The compiler already exposes the observability a good language site needs:

  • sexc -C — generated C
  • sexc -m — macro-expanded Sex
  • sexc --public-interface — what (import …) pastes (pub forms + docstrings)
  • sextest — compile + run + expected I/O

A docs engine that does not plug into those four is just a prettier Markdown site.

flowchart LR
  sources[".sex + prose"]
  sexc["sexc -C / -m / --public-interface"]
  ir["Sex doc IR"]
  html["HTML site"]
  llms["docs/llms + llms.txt"]
  tests["doctest via sextest"]
  sources --> sexc --> ir
  ir --> html
  ir --> llms
  ir --> tests

What “good PL docs” actually are

Four jobs, usually split across tools (Diátaxis):

  • Tutorial — sequenced, possibly interactive (The Rust Book, Go Tour, Red Blob Games).
  • How-to — tasks (“link two modules”, “call SDL3”).
  • Reference — indexed, cross-linked, generated from pub + docstrings (rustdoc, ExDoc, odoc).
  • Explanation — C mapping, macros vs C preprocessor, why type-last.

Plus two Sex-specific jobs:

  • Observability — same snippet as Sex, expanded Sex, and C, with the compiler as oracle (Compiler Explorer, not a notebook).
  • Agent payload — keep a Markdown/llms.txt tree that cannot drift from the human site.

Red Blob Games hexagons is the right tutorial aesthetic (SVG + Vue/D3, reader mutates inputs, diagram updates) and the wrong engine: Amit wrote custom JS per concept. That scales to a handful of explorable pages, not a language manual. Use it as a widget style inside a real docs compiler.

Landscape, grouped by what they optimize

1. Prose site generators (you write Markdown; they theme, search, nav)

Strong at manuals. Weak at “this fn exists because the compiler said so” unless you add a plugin.

  • Sphinx + MyST — still the most capable documentation compiler: domains, inventories, intersphinx, doctest, PDF, cross-refs that fail the build if a symbol is missing. Custom Sex domain (:sex:fn:, :sex:macro:, …) is the documented extension path. MyST lets current .md stay Markdown. Not pretty by default; themes (Furo, Book, PyData) fix that. Python toolchain next to a Chicken compiler.
  • MkDocs + Material — fastest pretty Markdown site. Plugins exist; no real language domain. Easy to outgrow once you want typed cross-refs and doctest.
  • mdBook — purpose-built for language books (The Rust Book). Playground + mdbook test are the pattern Sex wants, but they are Rust-shaped. Preprocessors are stdin/stdout JSON — implementable in Chicken. Weak as an API reference (Rust uses rustdoc beside the book).
  • Docusaurus, VitePress, Starlight, Fumadocs — modern MDX/islands. Best host if the product is interactive React/Vue widgets (closest to Red Blob). Versioning/i18n come free. Node stack; language semantics are your problem.
  • Antora — AsciiDoc, multi-repo, versioned products. Overkill until Sex has several manuals and release trains.

2. Language-native autodoc (compiler emits the reference)

This is what serious compiled languages eventually grow. Sex would have to write this layer regardless of host.

  • rustdoc — gold standard: intra-doc links, doctests that compile, search, source links. Book is a second tool (mdBook).
  • odoc — accurate cross-refs through fancy module types; .mld for prose.
  • ExDoc — guides + API + search + emits Markdown/llms.txt. Closest existing product to “human site and agent corpus from one source”.
  • Nim nim doc — compiler-built HTML/LaTeX/JSON from ## comments and standalone Markdown/RST, with checked links between them. Architecturally the closest to “fold docs into sexc”.
  • Zig autodoc — WASM frontend over compiler AST (zig std). Interesting long-term (interactive stdlib browser); large investment.
  • Doxygen + Breathe — documents the generated C, not Sex. Useful as a secondary view (“what the C compiler sees”), wrong as the primary language reference.

3. Document-as-program (culturally closest to Sex)

Prose is a program. Tags are functions. Build-time evaluation.

  • Scribble — Racket’s docs: defform, defproc, BNF, examples that run at doc-build, in-source docs. Best PL documentation language in existence. Host is Racket, Sex is Chicken — a second Scheme, not a free lunch.
  • Pollen — “the book is a program”; custom tags, multiple outputs. Authoring-oriented cousin of Scribble.
  • Skribilo — Guile, S-expression documents. Same family, smaller ecosystem.
  • GNU Texinfo — GCC, Guile, Emacs. Superb manuals, terrible modern web UX unless you wrap it.
  • Org-mode (already used for Readme.org) — babel can tangle/run blocks; ox-html/ox-hugo publish. Fine for the project README; a weak multi-page language site unless you commit to a full Org publishing pipeline.

If docs were written in Sex or Chicken S-exprs, this family is the honest design. It fights the existing Markdown-for-agents tree unless you generate docs/llms/ from the Scribble/Org source.

4. Computational / explorable publishing

  • Quarto — Pandoc + {ojs} (Observable). Best off-the-shelf Red-Blob-like diagrams in a docs toolchain. Weak Sex semantics; strong for a “how hex grids / how the type checker” explainer chapter.
  • Jupyter Book 2 / MyST Document Engine — executable Markdown, Thebe live kernels. Built for Python/R/Julia notebooks. Sex is not a kernel language; forcing Jupyter here is a category error.
  • Codapi — drop-in “Run” on fenced blocks; Sphinx/Docusaurus/Markdown adapters. Needs a sandbox server (or WASI). Natural “Run this .sex” widget once a sexc sandbox exists.
  • Compiler Explorer — the observability metaphor Sex actually wants (source | compiler output). Embed via client-state URLs, or copy the three-pane idea. CE does not know Sex today; sexc -C is a tiny Compiler Explorer.
  • Entangled / noweb — literate sync of Markdown ↔ source. Useful if examples live in the book and must stay real files.

Elixir Livebook and the Go Tour are interactive tutorials with a language runtime in the browser or a backend. Sex has no GC’d interpreter; a Tour implies sandboxing sexc + cc (server) or a future WASM sexc (Chicken static + WASI is a research project, not a docs milestone).

Constraints that knock options out

  • Sex is compiled, not a notebook. Build-time sexc (and maybe a small compile API) beats in-browser execution for years. Do not pick Jupyter/Thebe as the core.
  • Docs must not lie. rustdoc doctest / mdbook test / Sphinx doctest: every fenced sex block is a sextest (or at least sexc -C). If the engine cannot fail CI on a stale snippet, it is not the language’s docs.
  • Agents already read Markdown. llms.txt and docs/llms/ are load-bearing. Either Markdown is the source, or the engine emits them (ExDoc-style). Dual-maintaining HTML-only Scribble and hand Markdown will drift.
  • No tree-sitter / highlight.js grammar yet. Any host needs a highlighter (even a crude one) or code blocks stay plain.
  • Do not Doxygen the C as the primary manual. Readers write Sex; C is the observation pane.

What is worth writing ourselves

A full SSG (search, theming, i18n, PDF, mobile nav) is a trap. The Sex-specific core is small:

  1. Inventory — walk .sex, reuse --public-interface + docstrings on pub fn/struct/defmacro/….
  2. IR — JSON or S-expr: symbols, kinds, signatures, source spans, optional -C/-m snapshots.
  3. Doctest — extract fenced sex from prose; run through sextest / sexc.
  4. Widgets (HTML fragments any host can include):
    • three-pane: Sex | -m | -C (build-time first; live compile later)
    • feature-flag toggle documenting #+linux / #+(and unix (not macosx))
  5. Emitters — HTML (via a host) and docs/llms/*.md + llms.txt.

That core is “sexdoc”. The host is a commodity.

Shortlist (when we choose)

  • A. Sphinx + MyST + sexdoc extension — documentation-first, custom domain, build-time -C directives, PDF, inventories. Keep writing Markdown. Interactive diagrams: raw HTML or a few JS islands, not the default. Best if the manual and reference matter more than a slick tutorial landing page.
  • B. mdBook + Chicken preprocessor — language-book UX, preprocessor in the same Scheme world as sexc, mdbook test analogue. Pair with a later autodoc HTML for pub APIs (the rustdoc split). Best if the first artefact is The Sex Book.
  • C. Starlight or Docusaurus + sexdoc as MDX components — best Red-Blob-like tutorials (diagrams as components). Weakest typed cross-refs unless you generate MDX from the IR. Node in the docs toolchain.
  • D. Scribble or Pollen — if we want docs to be an S-expression program and will generate docs/llms/ from it. Highest ceiling, Racket (or Guile) as a third language, slowest to a decent site.
  • E. Fold sexdoc into sexc (Nim/rustdoc path) — right long-term for the reference; still needs a host for the tutorial. Do this after the IR exists, not as the first website.

Not recommended as the primary engine: MkDocs-alone, Antora, Jupyter Book, Org-publish, Doxygen/Breathe-as-front, hand-rolled HTML like Red Blob for the whole manual.

Hybrid that most languages actually ship: book engine (A or B) + compiler-backed reference (E) + 2–3 explorable pages (Quarto {ojs} or MDX diagrams). Codapi/CE-style “Run” only after a sandboxed sexc service exists.

Suggested evaluation, not implementation

If the next step is choosing (not building):

  • Take one existing page (docs/llms/syntax.md) and one module with docstrings (tests/modules/greet.sex).
  • Prototype the same three artefacts in Sphinx-MyST and in mdBook: rendered syntax page, pub API from --public-interface, and a snippet that shows Sex beside sexc -C.
  • Keep docs/llms/ as either source or generated output in both prototypes; reject any pipeline that cannot produce it.
  • Score: CI-failing doctest, cross-ref to greet, three-pane C view, agent Markdown, authoring friction.

That comparison will decide A vs B (and whether C is worth it for diagrams) more honestly than a second survey.

Evaluation checklist

  • Specify sexdoc IR from --public-interface, docstrings, and optional -C/-m snapshots; dual-emit HTML + docs/llms
  • Spike: Sphinx+MyST Sex domain, autodoc from greet.sex, doctest via sexc/sextest, C pane directive
  • Spike: mdBook + Chicken preprocessor for the same three artefacts
  • Score doctest CI, cross-refs, three-pane observability, llms.txt, authoring; pick host A/B/C/D
**We will keep in mind that agents need some kind of docs extracted from ~~humans~~ human docs too. Yes?** Sex has single documentation surface: `Readme.org`. Ground truth is `tests/` and `example/*.sex`. The compiler already exposes the observability a good language site needs: - `sexc -C` — generated C - `sexc -m` — macro-expanded Sex - `sexc --public-interface` — what `(import …)` pastes (`pub` forms + docstrings) - `sextest` — compile + run + expected I/O A docs engine that does not plug into those four is just a prettier Markdown site. ```mermaid flowchart LR sources[".sex + prose"] sexc["sexc -C / -m / --public-interface"] ir["Sex doc IR"] html["HTML site"] llms["docs/llms + llms.txt"] tests["doctest via sextest"] sources --> sexc --> ir ir --> html ir --> llms ir --> tests ``` ## What “good PL docs” actually are Four jobs, usually split across tools (Diátaxis): - **Tutorial** — sequenced, possibly interactive (The Rust Book, Go Tour, Red Blob Games). - **How-to** — tasks (“link two modules”, “call SDL3”). - **Reference** — indexed, cross-linked, generated from `pub` + docstrings (rustdoc, ExDoc, odoc). - **Explanation** — C mapping, macros vs C preprocessor, why type-last. Plus two Sex-specific jobs: - **Observability** — same snippet as Sex, expanded Sex, and C, with the compiler as oracle (Compiler Explorer, not a notebook). - **Agent payload** — keep a Markdown/`llms.txt` tree that cannot drift from the human site. [Red Blob Games hexagons](https://www.redblobgames.com/grids/hexagons/) is the right *tutorial* aesthetic (SVG + Vue/D3, reader mutates inputs, diagram updates) and the wrong *engine*: Amit wrote custom JS per concept. That scales to a handful of explorable pages, not a language manual. Use it as a widget style inside a real docs compiler. ## Landscape, grouped by what they optimize ### 1. Prose site generators (you write Markdown; they theme, search, nav) Strong at manuals. Weak at “this `fn` exists because the compiler said so” unless you add a plugin. - **[Sphinx](https://www.sphinx-doc.org/)** + **[MyST](https://myst-parser.readthedocs.io/)** — still the most capable *documentation compiler*: domains, inventories, intersphinx, `doctest`, PDF, cross-refs that fail the build if a symbol is missing. Custom **Sex domain** (`:sex:fn:`, `:sex:macro:`, …) is the documented extension path. MyST lets current `.md` stay Markdown. Not pretty by default; themes (Furo, Book, PyData) fix that. Python toolchain next to a Chicken compiler. - **[MkDocs](https://www.mkdocs.org/)** + Material — fastest pretty Markdown site. Plugins exist; no real *language domain*. Easy to outgrow once you want typed cross-refs and doctest. - **[mdBook](https://rust-lang.github.io/mdBook/)** — purpose-built for *language books* (The Rust Book). Playground + `mdbook test` are the pattern Sex wants, but they are Rust-shaped. Preprocessors are stdin/stdout JSON — implementable in Chicken. Weak as an API reference (Rust uses rustdoc beside the book). - **[Docusaurus](https://docusaurus.io/)**, **[VitePress](https://vitepress.dev/)**, **[Starlight](https://starlight.astro.build/)**, **[Fumadocs](https://fumadocs.dev/)** — modern MDX/islands. Best host if the *product* is interactive React/Vue widgets (closest to Red Blob). Versioning/i18n come free. Node stack; language semantics are your problem. - **[Antora](https://antora.org/)** — AsciiDoc, multi-repo, versioned products. Overkill until Sex has several manuals and release trains. ### 2. Language-native autodoc (compiler emits the reference) This is what serious compiled languages eventually grow. Sex would have to write this layer regardless of host. - **rustdoc** — gold standard: intra-doc links, doctests that compile, search, source links. Book is a *second* tool (mdBook). - **[odoc](https://github.com/ocaml/odoc)** — accurate cross-refs through fancy module types; `.mld` for prose. - **[ExDoc](https://github.com/elixir-lang/ex_doc)** — guides + API + search + **emits Markdown/`llms.txt`**. Closest existing product to “human site and agent corpus from one source”. - **Nim `nim doc`** — compiler-built HTML/LaTeX/JSON from `##` comments *and* standalone Markdown/RST, with checked links between them. Architecturally the closest to “fold docs into `sexc`”. - **Zig autodoc** — WASM frontend over compiler AST (`zig std`). Interesting long-term (interactive stdlib browser); large investment. - **Doxygen + [Breathe](https://breathe.readthedocs.io/)** — documents the *generated C*, not Sex. Useful as a secondary view (“what the C compiler sees”), wrong as the primary language reference. ### 3. Document-as-program (culturally closest to Sex) Prose is a program. Tags are functions. Build-time evaluation. - **[Scribble](https://docs.racket-lang.org/scribble/)** — Racket’s docs: `defform`, `defproc`, BNF, `examples` that run at doc-build, in-source docs. Best *PL documentation language* in existence. Host is Racket, Sex is Chicken — a second Scheme, not a free lunch. - **[Pollen](https://docs.racket-lang.org/pollen/)** — “the book is a program”; custom tags, multiple outputs. Authoring-oriented cousin of Scribble. - **[Skribilo](https://www.nongnu.org/skribilo/)** — Guile, S-expression documents. Same family, smaller ecosystem. - **GNU Texinfo** — GCC, Guile, Emacs. Superb manuals, terrible modern web UX unless you wrap it. - **Org-mode** (already used for `Readme.org`) — babel can tangle/run blocks; ox-html/ox-hugo publish. Fine for the project README; a weak multi-page language site unless you commit to a full Org publishing pipeline. If docs were *written in Sex or Chicken S-exprs*, this family is the honest design. It fights the existing Markdown-for-agents tree unless you generate `docs/llms/` from the Scribble/Org source. ### 4. Computational / explorable publishing - **[Quarto](https://quarto.org/)** — Pandoc + `{ojs}` (Observable). Best *off-the-shelf* Red-Blob-like diagrams in a docs toolchain. Weak Sex semantics; strong for a “how hex grids / how the type checker” explainer chapter. - **[Jupyter Book 2 / MyST Document Engine](https://jupyterbook.org/)** — executable Markdown, Thebe live kernels. Built for Python/R/Julia notebooks. Sex is not a kernel language; forcing Jupyter here is a category error. - **[Codapi](https://github.com/nalgeon/codapi-js)** — drop-in “Run” on fenced blocks; Sphinx/Docusaurus/Markdown adapters. Needs a sandbox server (or WASI). Natural “Run this `.sex`” widget once a `sexc` sandbox exists. - **[Compiler Explorer](https://godbolt.org/)** — the observability metaphor Sex actually wants (source | compiler output). Embed via client-state URLs, or copy the three-pane idea. CE does not know Sex today; `sexc -C` *is* a tiny Compiler Explorer. - **[Entangled](https://entangled.github.io/)** / noweb — literate sync of Markdown ↔ source. Useful if examples live in the book and must stay real files. Elixir **Livebook** and the **Go Tour** are interactive tutorials with a language runtime in the browser or a backend. Sex has no GC’d interpreter; a Tour implies sandboxing `sexc` + `cc` (server) or a future WASM `sexc` (Chicken static + WASI is a research project, not a docs milestone). ## Constraints that knock options out - **Sex is compiled, not a notebook.** Build-time `sexc` (and maybe a small compile API) beats in-browser execution for years. Do not pick Jupyter/Thebe as the core. - **Docs must not lie.** rustdoc doctest / `mdbook test` / Sphinx doctest: every fenced `sex` block is a sextest (or at least `sexc -C`). If the engine cannot fail CI on a stale snippet, it is not the language’s docs. - **Agents already read Markdown.** `llms.txt` and `docs/llms/` are load-bearing. Either Markdown is the source, or the engine *emits* them (ExDoc-style). Dual-maintaining HTML-only Scribble and hand Markdown will drift. - **No tree-sitter / highlight.js grammar yet.** Any host needs a highlighter (even a crude one) or code blocks stay plain. - **Do not Doxygen the C as the primary manual.** Readers write Sex; C is the observation pane. ## What is worth writing ourselves A full SSG (search, theming, i18n, PDF, mobile nav) is a trap. The Sex-specific core is small: 1. **Inventory** — walk `.sex`, reuse `--public-interface` + docstrings on `pub fn`/`struct`/`defmacro`/…. 2. **IR** — JSON or S-expr: symbols, kinds, signatures, source spans, optional `-C`/`-m` snapshots. 3. **Doctest** — extract fenced `sex` from prose; run through sextest / `sexc`. 4. **Widgets** (HTML fragments any host can include): - three-pane: Sex | `-m` | `-C` (build-time first; live compile later) - feature-flag toggle documenting `#+linux` / `#+(and unix (not macosx))` 5. **Emitters** — HTML (via a host) and `docs/llms/*.md` + `llms.txt`. That core is “sexdoc”. The host is a commodity. ## Shortlist (when we choose) - **A. Sphinx + MyST + sexdoc extension** — documentation-first, custom domain, build-time `-C` directives, PDF, inventories. Keep writing Markdown. Interactive diagrams: raw HTML or a few JS islands, not the default. Best if the manual and reference matter more than a slick tutorial landing page. - **B. mdBook + Chicken preprocessor** — language-book UX, preprocessor in the same Scheme world as `sexc`, `mdbook test` analogue. Pair with a later autodoc HTML for `pub` APIs (the rustdoc split). Best if the first artefact is *The Sex Book*. - **C. Starlight or Docusaurus + sexdoc as MDX components** — best Red-Blob-like tutorials (diagrams as components). Weakest typed cross-refs unless you generate MDX from the IR. Node in the docs toolchain. - **D. Scribble or Pollen** — if we want docs to be an S-expression program and will generate `docs/llms/` from it. Highest ceiling, Racket (or Guile) as a third language, slowest to a decent site. - **E. Fold sexdoc into `sexc` (Nim/rustdoc path)** — right long-term for the *reference*; still needs a host for the tutorial. Do this after the IR exists, not as the first website. **Not recommended as the primary engine:** MkDocs-alone, Antora, Jupyter Book, Org-publish, Doxygen/Breathe-as-front, hand-rolled HTML like Red Blob for the whole manual. Hybrid that most languages actually ship: **book engine (A or B) + compiler-backed reference (E) + 2–3 explorable pages (Quarto `{ojs}` or MDX diagrams)**. Codapi/CE-style “Run” only after a sandboxed `sexc` service exists. ## Suggested evaluation, not implementation If the next step is choosing (not building): - Take one existing page (`docs/llms/syntax.md`) and one module with docstrings (`tests/modules/greet.sex`). - Prototype the same three artefacts in Sphinx-MyST and in mdBook: rendered syntax page, `pub` API from `--public-interface`, and a snippet that shows Sex beside `sexc -C`. - Keep `docs/llms/` as either source or generated output in both prototypes; reject any pipeline that cannot produce it. - Score: CI-failing doctest, cross-ref to `greet`, three-pane C view, agent Markdown, authoring friction. That comparison will decide A vs B (and whether C is worth it for diagrams) more honestly than a second survey. ## Evaluation checklist - [ ] Specify sexdoc IR from `--public-interface`, docstrings, and optional `-C`/`-m` snapshots; dual-emit HTML + `docs/llms` - [ ] Spike: Sphinx+MyST Sex domain, autodoc from `greet.sex`, doctest via `sexc`/`sextest`, C pane directive - [ ] Spike: mdBook + Chicken preprocessor for the same three artefacts - [ ] Score doctest CI, cross-refs, three-pane observability, `llms.txt`, authoring; pick host A/B/C/D
pkulev added the feature label 2026-09-22 09:47:14 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alex-eg/sex#32