Table of Contents

Playbook Format

Note

Status: implemented. The playbook is the versioned JSON document a compile emits and a runtime loads; this note owns its wire format, the writer, and the reader with its full list of checks. The cross-cutting decisions it applies are in the Dialogue Runtime Architecture.

Table of contents

Goal and scope

The compiler ends at an internal DialogueGraph. The playbook is its portable form: a versioned JSON document that a runtime in any language can load.

In scope:

  • the playbook types and their JSON encoding;
  • the writer — DialogueGraph to playbook;
  • the reader — JSON to playbook, refusing anything it cannot play correctly;
  • the compatibility header, and the capability names version 0 defines;
  • a hand-written JSON Schema that specifies the format;
  • ddown compile --output.

Out of scope, each owned elsewhere: playing a playbook, the [compatibility] and [features] configuration, exporters, source maps, and binary encoding.

This note assumes the vocabulary and decisions of the architecture note — playbook, capability, effect, query — and does not restate them.

Where the types live

The playbook is the contract between a compiler and a runtime, so it belongs to neither. It gets its own assembly, which references no other project:

flowchart BT
    PB["DialogueDown.Playbook<br/>types + reader"]
    C["DialogueDown<br/>compiler + writer"] --> PB
    R["DialogueDown.Runtime"] --> PB
    CLI["DialogueDown.Cli"] --> C

The writer lives in DialogueDown, not beside the reader, and that placement is forced rather than chosen. CompilationSuccess.Graph and the entire node and edge model are internal; a writer in DialogueDown.Playbook would therefore have to reference DialogueDown, which would make Runtime → Playbook → DialogueDown and drag Markdig, Tomlyn, and the diagnostics engine into every shipped game. Keeping the writer beside the graph it reads is what lets the runtime stay small.

The writer still emits only public playbook types, so the mapping from internal graph to public contract stays explicit and reviewable — the property this design asked for.

DialogueDown.Playbook is an assembly a game ends up referencing, through the runner that reads its playbooks. It therefore multi-targets net8.0;net10.0 like the other shipped libraries, so a Godot export keeps loading on Godot's bundled runtime — see Target Frameworks. It carries two dependencies. System.Text.Json, scoped to net8.0, supplies an attribute the framework gained in .NET 9; on net10.0 it costs nothing, and it is dropped when net8.0 goes. Generator.Equals.Runtime is the small comparer assembly the generated equality calls, pulled in because the records compare by value (see P12); the generator itself is a private, build-time package that never ships.

The writer sits in DialogueDown.Emission, the stage between the graph and a runtime. Keeping it out of DialogueDown.Playbook keeps the boundary legible: one namespace writes the exchange format, the other reads it. Within the format assembly the types are grouped by role — Speech, Edges, Nodes, Conditions, Weights, Speakers, Checking — matching how the compiler already lays out its script and graph layers, and leaving the document, its header, and the reader at the root.

Two architecture tests guard the shape:

Test Asserts
Playbook_DependsOnlyOn_TheFrameworkAndGenerator DialogueDown.Playbook references no other project and, outside System, only the equality generator
Runtime_DependsOnlyOn_ThePlaybook DialogueDown.Runtime references DialogueDown.Playbook and nothing else

The document

{
  "$schema": "https://pengzhengyi.github.io/dialoguedown/schema/playbook-0.schema.json",
  "format": { "version": 0, "requires": ["core"], "uses": [] },
  "script": "chapter-01.dialogue.md",
  "entry": 0,
  "anchors": { "the-inn": 4 },
  "speakers": [
    { "name": "Alice", "tags": [{ "name": "mood", "value": "warm" }] },
    { "default": true, "tags": [] }
  ],
  "nodes": [
    {
      "id": 0,
      "kind": "line",
      "speaker": 0,
      "speech": [
        { "kind": "text", "text": "My favorite color is " },
        { "kind": "query", "key": "Alice.FavoriteColor" },
        { "kind": "text", "text": "." }
      ],
      "out": [{ "kind": "succession", "target": 1 }]
    },
    {
      "id": 1,
      "kind": "choice",
      "ordered": false,
      "out": [
        {
          "kind": "option",
          "target": 2,
          "label": [{ "kind": "text", "text": "Ask about the inn" }],
          "condition": { "kind": "key", "key": "IsCurious" }
        },
        { "kind": "option", "target": 3, "label": [{ "kind": "text", "text": "Say nothing" }] }
      ]
    }
  ]
}

The format header

The header answers one question — should this be loaded at all? — so it is grouped rather than scattered across the root. That makes the loader's first check visible in both the document and the schema, and lets new format-level fields land without cluttering the top level.

Following glTF, which groups version information under asset while leaving scenes and nodes at the root, only the header is nested; entry, anchors, speakers, and nodes are all content in different shapes, so wrapping them again would add a level of nesting to every access and buy nothing. script stays at the root because provenance is not compatibility.

entry and anchors

A game does not play a file top to bottom — it starts a specific conversation. Ink addresses one with ChoosePathString("knot.stitch") and Yarn Spinner with SetNode("Start"); both resolve a name against the table of named things they already keep. Here that table is anchors, which maps every scene's slug to its node, and it is what a host uses to begin at a named scene or to follow a jump.

entry adds the one fact anchors cannot hold: where a playthrough begins when nothing says otherwise. That is the document's top, which has no heading and so has no slug.

It is a single field rather than a table of named ways in. A table would have exactly one member, under a name the compiler never wrote, and would promise a multiplicity that does not exist — while buying nothing, since a new field is additive anyway. Starting somewhere else is already expressible through anchors, and a debugger's "begin here" is an argument to the runner rather than a fact about the document.

Speakers are hoisted, and addressed by index

A speaker's name and tags are hoisted out of the lines that quote them, so a host has one place to bind a portrait, a voice, or a color, and a diff stays local when a speaker changes.

A line names its speaker by index, the way every other reference in a playbook works. The alternative — synthesizing a string key from the name — would invent an identifier nobody wrote, and would spend the word id, which the script language already uses for the writer's own @id. A host reading a speaker's id therefore always knows the writer typed it.

Both the @id and the name are optional, because the anonymous default speaker has neither and still says lines.

Mapping the graph

The writer's whole job is this mapping. Every row is one test.

Nodes

Graph kind Carries
LineNode line speaker, speech, condition?
ChoiceNode choice ordered
RandomChoiceNode random-choice —
BranchNode branch —
ControlNode control effects, condition?
EndNode end — (no outgoing edges)

Edges

Graph kind Carries
SuccessionEdge succession target
OptionEdge option target, label, condition?
RandomOptionEdge random-option target, weight, condition?
BranchEdge branch target, order, condition?
DivertEdge divert target, label, condition?

order on a branch edge preserves if/elseif/else evaluation order, which is otherwise lost in a JSON array a reader may not be required to keep ordered.

Both label-bearing edges carry their own text, rather than deriving it from the node they lead to. For an option that is a correctness matter as much as a convenience: an arm with an empty body leads straight past the choice, so a label read back off the target would be another line's words. Only ChoicePass, which still holds the arm's body, can know. DivertPass carries a jump's label for a different reason — nothing else keeps it, since a jump is written inside a line but is no part of what that line says.

Speech fragments

AST kind Carries
Text text text
StyledText styled style (italic, bold, strikethrough), children
Link link target, label
Image image source, alt
LineBreak break —
Query query key
DefaultCommand default-command action
CustomCommand custom-command name, args
ReservedTag, CustomTag tag name, value?, reserved

The two command kinds are spelled the way the guide names them, rather than shortened to command and call — the second is a word the project does not use with writers, and the pair reads better matched.

A fragment's enum values — style's italic, bold, and strikethrough — are pinned by hand-written converters, so a C# rename cannot change what a playbook says; see Enum wire names.

Fragments nest — StyledText.Children and a link or image label are themselves fragment lists — so the encoding is recursive. Nothing is flattened to a string, because a host re-renders it: Godot as BBCode, the report as HTML, the CLI as ANSI.

Tags stay in the fragment list, in position, rather than being hoisted beside the speaker. A tag is a plausible hook — a host may merely recognize one, or act on it — and both cases need to know where in the line it attached. Hoisting would discard that for a small saving.

A command inside a line's speech is an effect (the graph's LineNode.Effects). The playbook keeps it in place for the same reason, so a runtime knows where in the line it fires. How a runner orders speaking and performing it is not settled: the runner reports it inside Said and does not ask the host to perform it. A command on its own line is a control node, which the runner does perform.

Conditions and weights

Both are objects with a kind, never bare scalars, so negation and expressions stay additive:

"condition": { "kind": "key", "key": "IsCurious" }
"weight":    { "kind": "auto" }
"weight":    { "kind": "number", "percentage": 25.0 }
"weight":    { "kind": "query",  "key": "Bob.Affection" }

The wire name is condition, matching the Condition type in the AST and the graph, and the word the guide uses with writers.

Reading a playbook

The reader is a gatekeeper before it is a parser. It parses, then hands the result to a checker — one rule, one class — and returns only what passes. PlaybookCheckerFactory.CreateDefault() wires the standard set, in this order:

Checker Refuses
FormatChecker → VersionChecker a format.version outside the range this build reads
FormatChecker → CapabilityChecker a name in requires this build does not support
NodePositionChecker a node whose id is not its position (nodes[i].id != i)
ReferenceChecker an entry, anchor, edge target, or speaker index that lands nowhere
OutwardShapeChecker a node whose ways out break the outward-shape rule — see Playbook Reader Rules
BranchArmOrderChecker a branch whose arms are out of order, lack a gated arm, or put the else before the end — see Playbook Reader Rules

The version settles before the capabilities, because a document of an unknown shape may describe its capabilities in terms an older build would misread. Each later checker assumes what the ones before it proved.

The rules are injected, not inherited. PlaybookReader takes an IPlaybookChecker; PlaybookCheckerFactory.CreateDefault() wires the standard set, and a runtime that reads fewer constructs than this build supplies its own. Each rule takes the policy it applies — a version range, a capability set — rather than reading a constant, so nothing has to be subclassed to be narrowed.

A checker refuses at the first thing it finds wrong rather than gathering a list. The compiler reports every diagnostic because a writer wants the whole set; a malformed playbook is compiler output, so there is nothing for a reader to work through and fix. Every failure names the offending value and the expectation. Nothing degrades, because a skipped condition does not error — it silently tells the wrong story.

The reader is also where the format's forward compatibility lives: unknown object properties are ignored, so a newer compiler may add optional metadata without breaking an older reader, while an unknown entry in requires is a hard refusal.

Version 0 defines exactly two capability names:

Capability Meaning
core Everything the compiler emits. Always present in requires
cross-file-jump Reserved, and neither emitted nor read in version 0. A node reference is a plain index; a reference into another script waits for the linker to settle what a script identity is. Naming it keeps it unclaimed and makes the widening additive

Key design decisions

P1 — The playbook is a designed contract, not a graph dump

The writer maps each internal type onto a public shape by hand. Serializing DialogueGraph directly would be faster to build and would inherit SourceSpan, SpeakerSymbol, and every future refactor of compiler internals as a breaking format change — the coupling this design exists to avoid. An explicit mapping is also a place to put a test per construct.

P2 — System.Text.Json polymorphism with a kind discriminator

[JsonPolymorphic(TypeDiscriminatorPropertyName = "kind")] plus [JsonDerivedType(typeof(LinePlaybookNode), "line")] round-trips the node, edge, fragment, condition, and weight unions with no custom converters. The discriminator is spelled kind rather than the library default $type because the format is a public contract, not a .NET serialization detail; a TypeScript reader switches on the same word.

No hand-written converter is needed. A node reference is a plain integer index, so every value in the format is either a primitive or a tagged object — see P10.

P3 — The playbook speaks the project's own words

Every wire name is the word the project already uses: target for an edge's destination (Edge(NodeId Target), Link.Target, Jump.Target), condition for what must hold (the Condition type, and three notes titled Conditional …), and query for a pure read of the world.

Inventing wire names is how a format drifts from the language its users speak.

P4 — What a runner can derive is not stored

A node does not carry the query keys needed to leave it. They are derivable by walking the node — its condition, its speech, its effects, and every out-edge's condition and weight — so storing them would put one fact in two places, where the two can disagree.

Storing them buys nothing measurable either. A runner walks each node once at load and caches the result, so there is no play-time difference. Resolving a node's reads in one round trip comes from the protocol's Resolve/Supply pair, not from a field in the artifact.

The walk is subtler than it looks — a query can hide inside an option's label — and that is exactly what the conformance corpus exists to keep honest across runtimes.

P5 — Node ids are dense indices

A playbook node's id is its position in the nodes array, and the reader refuses a document where it is not.

This is a renumbering, not a passthrough. The compiler's NodeId is by contract an opaque handle — DialogueGraph resolves it through an id-keyed dictionary and its documentation says outright not to treat the value as a list index — and ids are minted in the order blocks are encountered, which is not the document order the node list is in. So the writer must translate either way; the only question is what it translates into.

A dense index is chosen because it makes the document verifiable. One comparison per node proves the ids are unique, gapless, and correctly ordered at once:

Failure Carrying the opaque ids Dense indices
Duplicate id Needs a separate uniqueness pass Caught by nodes[i].id == i
Gap in the numbering Legal, and undetectable Caught by nodes[i].id == i
Node list silently reordered Undetectable Caught by nodes[i].id == i

That last row is the one that matters: a reordered array is a valid playbook that tells a different story, exactly the failure this format exists to prevent. The explicit id costs a few bytes and buys a checksum — as well as a document that reads well and answers jq '.nodes[] | select(.id == 42)'.

Resolving a reference then costs an array index rather than a dictionary lookup, but that is a bonus, not the reason. A dialogue graph is hundreds of nodes; the lookup was never the problem.

The usual argument for preserving original ids — correlating a runtime error with a compiler diagnostic — does not apply, because NodeId is internal and never surfaces. DialogueDown's diagnostics address source positions, not nodes.

P6 — Anchors, not the region tree

Play needs to answer "which node opens #the-inn", which is a flat anchors table. The full RegionTree — nesting, OwnNodes, scene labels — serves analysis and presentation, so emitting it would ship data with no consumer. Adding it is additive under an advisory uses entry.

P7 — The schema is the specification; conformance is proven, not generated

A hand-written JSON Schema 2020-12 is the format's normative structural specification. It is what a porter reads, so every field carries a description — which a generated schema cannot provide.

The C# types are hand-written too, as sealed records with [JsonPolymorphic]. Neither direction of code generation is worth its cost here: generating C# from the schema produces mutable POCOs, and generating the schema from C# needs .NET 9 (the shipped libraries still target net8.0) and yields an undocumented schema.

The two are kept honest by validating real output: every golden playbook is checked against the schema in CI, so a drift in either direction fails the build. A TypeScript runner would generate its types from the schema — the payoff that makes a hand-written spec worth writing.

This splits responsibility cleanly:

The schema is normative for structure. The conformance corpus is normative for behavior.

P8 — No validator ships with the reader

Schema validation stays in authoring, editor tooling, and CI; the reader relies on typed deserialization plus the explicit semantic checks below. That is what every comparable format does — glTF ships a separate validator tool and its loaders (Three.js, Babylon, Unity, Godot) never schema-validate; Yarn Spinner's schemas serve its VS Code extension and CI, not its runtime.

It is also what a schema cannot do that decides it. Structure it handles well, and the shipped schema proves it: kind enums, per-kind required fields, recursion through nested speech, and even "an end leads nowhere" as a maxItems. Relational integrity it cannot express at all — nodes[i].id == i, references landing in range, or "an unknown requires refuses while an unknown uses does not". Those need code regardless, so a validator dependency would add weight without removing work.

Keeping it out also sidesteps a licensing trap worth recording: JsonSchema.Net attaches a EULA to its binaries from v9.0.0 that asks revenue-generating users to pay, and Newtonsoft.Json.Schema is AGPL below a paid tier with a ten-validations-per-hour cap. Neither is acceptable to inherit into a game.

P9 — Human-readable by default

The writer pretty-prints. A playbook is meant to be opened, jq-ed, and reasoned about, and readability is worth more than bytes for a file that measures in kilobytes. A compact mode — and, if it ever proves its value, compression or a binary encoding — stays available behind a CLI flag, because the writer is a seam.

P10 — Cross-file references wait for the linker

A node reference is an integer index and nothing else. A reference into another script would have to spell out what a script identity is, whether a bare script means its root scene, and how an anchor is written — three answers owned by the linker, which is explored rather than implemented. Encoding guesses about them into a public contract would make the linker inherit them.

Deferring costs nothing, because the widening is already additive: a playbook that uses cross-file references will declare the cross-file-jump capability, and a version-0 runner refuses the whole document before parsing a single node. It cannot misread a reference shape it never reaches. The capability manifest is what makes cross-file additive — not the shape of the reference field.

P11 — Absent is absent; the writer emits the shortest true document

A field whose value is the default is omitted: no null, and no false. A tag without a value is { "name": "aside" }, and a reader treats a missing field exactly as it treats a missing requires — as the default case.

Writing both spellings would let two documents mean the same thing, which doubles what a schema, a reader, and a golden file each have to say.

Warning

Omitting defaults is applied per flag, never as a blanket serializer setting. The blanket condition also drops value types equal to zero — which would silently erase a node's id, an edge's target, and the first branch arm's order.

P12 — Records compare by value

The playbook types are records, which advertise value equality, but every collection property is an ImmutableArray<T> or an ImmutableSortedDictionary<TKey,TValue>, whose own Equals is reference equality. Left alone, two structurally identical playbooks compare unequal, so == on a public contract means something narrower than it appears to.

Generator.Equals supplies the equality instead. Each collection-owning record is partial and marked [Equatable], and each collection property carries [OrderedEquality] or [UnorderedEquality]; the generator emits Equals and GetHashCode at build time. Its analyzer fails the build when an [Equatable] type declares a collection property without one of those attributes, so a property added later cannot be forgotten.

This is the contract's one exception to "nothing outside System". The generated code calls comparer types from Generator.Equals.Runtime, which every consumer and embedded game therefore carries — about 18 KB, plus Microsoft.Bcl.HashCode (about 20 KB) that the package's netstandard2.0 target pulls in. The architecture test allows exactly that dependency. Hand-writing Equals and GetHashCode per record would keep the assembly reference-free, but it repeats the same code across a dozen records and guards the forgotten property only with a test; the generator's compile-time check is the stronger guard, and the dependency is small.

Error and boundary cases

Case Behavior
format.version newer than the reader Refuse, naming both versions
format.version older than the reader's floor Refuse
Unknown name in requires Refuse, naming the capability
Unknown name in uses Accept — advisory by definition
Unknown object property Ignore — forward compatibility
nodes[i].id != i Refuse
Node reference out of range Refuse
A node reference that is not a number Refuse — the deserializer types it as an integer, so a string is not a reference at all
An entry pointing nowhere Refuse — a playbook nothing can start is not playable
Duplicate speaker id Refuse; the writer asserts uniqueness before emitting
Duplicate anchor Cannot occur; the compiler already rejects it (DLG2001)
A script that compiles with errors No playbook is written; --output is untouched
A script that compiles with warnings A playbook is written. Warnings are a smell a compiler tolerates; anything intolerable belongs in the error tier
A script with no dialogue A valid playbook with an entry that reaches end
Empty speech on a line Cannot occur; the AST rejects empty styled content

Integration

Seam Change
CompilationSuccess Unchanged. The writer consumes its internal graph inside the same assembly
IPlaybookWriter New public seam in DialogueDown, registered in AddDialogueDown and the CLI composition root, following the IDialogueGraphBuilder pattern
CompileCommand A playbook is what compile emits unless told otherwise: -o names where it goes, and --emit dot asks for the stage graphs instead.
CLI presentation Nothing in the CLI reads a playbook, so there is no load-time failure to render. A command that plays one renders it outside the DLG code space, in the style of CLI Diagnostic Rendering
DialogueDown.csproj References DialogueDown.Playbook; the package ships both
Central package management DialogueDown.Playbook inherits Directory.Packages.props: System.Text.Json (net8.0) for a framework attribute, Generator.Equals (private, build-time) to generate equality, and the Generator.Equals.Runtime it calls
CI A check-jsonschema step validates every golden playbook against the local schema file, so validation never depends on the network
Editors Emitted playbooks carry a versioned $schema URL published with the existing GitHub Pages site, so VS Code validates a playbook wherever it lands. See open questions for zero-config registration

Testability

Level What it covers
Unit — writer One test per row of Mapping the graph: each node, edge, fragment, condition, and weight kind
Unit — reader Every refusal in Error and boundary cases, each asserting the message names the offending value
Round-trip Compile, write, read, and assert the JSON comes back identical — the primary safety net, and cheap because both directions land here
Equality Records compare by value across every collection; guarded by the analyzer, a reflection test, and a corpus round-trip
Exhaustive Reflection over each closed union, so a construct added to the AST fails here rather than at whatever runtime reaches it first
Golden A committed playbook per compiling examples/*.dialogue.md, so a format change is a reviewable diff
Schema Every golden playbook validates against the schema in CI

Round-trip tests live in DialogueDown.Tests, which already sees internals and can reference both assemblies. Playbook fixtures are built through a shared factory so a shape change touches one file.

A round trip is asserted as text, not as objects: the writer's JSON must read back into a document that writes the same JSON again, so a change in how a field is spelled or ordered is caught rather than hidden. Where a test asks a question about the model, the records compare by value (see P12) and can be asserted directly.

Goldens use Verify, which supplies the matching and the accept workflow. It was measured against this suite before being taken: it compares text rather than reparsing the JSON — which is what makes a formatting change visible at all — and one golden serves both target frameworks.

Golden playbooks churn when node positions shift, which is expected: they are a build artifact nobody hand-edits. The semantic regression asset is the hand-authored conformance corpus, which pins what a playbook means rather than what it serializes to.

The exhaustive tests are the ones that repaid the most. Asking reflection for every member of InlineFragment turned up three the mapping table never listed — a jump, a condition, and a jump indicator — and so caught, before any of it shipped, that a jump survives into a line's speech and would have thrown on every script containing one.

Open questions and deferred work

  • Zero-config editor support. A versioned $schema URL works wherever a playbook lands, but a reader must still know to look. Registering *.playbook.json with SchemaStore would make VS Code validate a playbook with no $schema key at all — the best experience for a non-developer. It needs the URL to be stable first.
  • Line identity stays deferred. Nothing needs reserving: the schema allows properties it does not name, and a reader ignores them, so adding lineId is a field to populate rather than a shape to change.
  • The schema has no negative tests. Seven malformed playbooks were checked by hand and each was refused, but nothing stops an edit turning the schema into a rubber stamp. CI runs --check-metaschema, which catches structural breakage and not a weakened rule. Committed counter-examples would close that.
  • A divert's label is not drawn on the graph stage of the report. The words survive on the edge and in the Dialogue AST stage, but the graph view would read better with them.