Playbook Format
Note
Status: implemented. The playbook is the versioned JSON document a compile emits and a runtime loads; this note owns its wire format, the writer, and the reader with its full list of checks. The cross-cutting decisions it applies are in the Dialogue Runtime Architecture.
Table of contents
- Goal and scope
- Where the types live
- The document
- Mapping the graph
- Reading a playbook
- Key design decisions
- Error and boundary cases
- Integration
- Testability
- Open questions and deferred work
Goal and scope
The compiler ends at an internal DialogueGraph. The playbook is its portable
form: a versioned JSON document that a runtime in any language can load.
In scope:
- the playbook types and their JSON encoding;
- the writer —
DialogueGraphto playbook; - the reader — JSON to playbook, refusing anything it cannot play correctly;
- the compatibility header, and the capability names version 0 defines;
- a hand-written JSON Schema that specifies the format;
ddown compile --output.
Out of scope, each owned elsewhere: playing a playbook, the
[compatibility] and [features] configuration, exporters, source maps, and
binary encoding.
This note assumes the vocabulary and decisions of the architecture note — playbook, capability, effect, query — and does not restate them.
Where the types live
The playbook is the contract between a compiler and a runtime, so it belongs to neither. It gets its own assembly, which references no other project:
flowchart BT
PB["DialogueDown.Playbook<br/>types + reader"]
C["DialogueDown<br/>compiler + writer"] --> PB
R["DialogueDown.Runtime"] --> PB
CLI["DialogueDown.Cli"] --> C
The writer lives in DialogueDown, not beside the reader, and that placement
is forced rather than chosen. CompilationSuccess.Graph and the entire node and
edge model are internal; a writer in DialogueDown.Playbook would therefore
have to reference DialogueDown, which would make
Runtime → Playbook → DialogueDown and drag Markdig, Tomlyn, and the diagnostics
engine into every shipped game. Keeping the writer beside the graph it reads is
what lets the runtime stay small.
The writer still emits only public playbook types, so the mapping from internal graph to public contract stays explicit and reviewable — the property this design asked for.
DialogueDown.Playbook is an assembly a game ends up referencing, through the
runner that reads its playbooks. It therefore multi-targets net8.0;net10.0 like
the other shipped libraries, so a Godot export keeps loading on Godot's bundled
runtime — see Target Frameworks. It carries two
dependencies. System.Text.Json, scoped to net8.0, supplies an attribute the
framework gained in .NET 9; on net10.0 it costs nothing, and it is dropped when
net8.0 goes. Generator.Equals.Runtime is the small comparer assembly the
generated equality calls, pulled in because the records compare by value (see
P12); the generator itself is a private,
build-time package that never ships.
The writer sits in DialogueDown.Emission, the stage between the graph and a
runtime. Keeping it out
of DialogueDown.Playbook keeps the boundary legible: one namespace writes the
exchange format, the other reads it. Within the format assembly the types are
grouped by role — Speech, Edges, Nodes, Conditions, Weights, Speakers,
Checking — matching how the compiler already lays out its script and graph
layers, and leaving the document, its header, and the reader at the root.
Two architecture tests guard the shape:
| Test | Asserts |
|---|---|
Playbook_DependsOnlyOn_TheFrameworkAndGenerator |
DialogueDown.Playbook references no other project and, outside System, only the equality generator |
Runtime_DependsOnlyOn_ThePlaybook |
DialogueDown.Runtime references DialogueDown.Playbook and nothing else |
The document
{
"$schema": "https://pengzhengyi.github.io/dialoguedown/schema/playbook-0.schema.json",
"format": { "version": 0, "requires": ["core"], "uses": [] },
"script": "chapter-01.dialogue.md",
"entry": 0,
"anchors": { "the-inn": 4 },
"speakers": [
{ "name": "Alice", "tags": [{ "name": "mood", "value": "warm" }] },
{ "default": true, "tags": [] }
],
"nodes": [
{
"id": 0,
"kind": "line",
"speaker": 0,
"speech": [
{ "kind": "text", "text": "My favorite color is " },
{ "kind": "query", "key": "Alice.FavoriteColor" },
{ "kind": "text", "text": "." }
],
"out": [{ "kind": "succession", "target": 1 }]
},
{
"id": 1,
"kind": "choice",
"ordered": false,
"out": [
{
"kind": "option",
"target": 2,
"label": [{ "kind": "text", "text": "Ask about the inn" }],
"condition": { "kind": "key", "key": "IsCurious" }
},
{ "kind": "option", "target": 3, "label": [{ "kind": "text", "text": "Say nothing" }] }
]
}
]
}
The format header
The header answers one question — should this be loaded at all? — so it is grouped rather than scattered across the root. That makes the loader's first check visible in both the document and the schema, and lets new format-level fields land without cluttering the top level.
Following glTF, which groups version information under asset while leaving
scenes and nodes at the root, only the header is nested; entry, anchors,
speakers, and nodes are all content in different shapes, so wrapping them again
would add a level of nesting to every access and buy nothing. script stays at the
root because provenance is not compatibility.
entry and anchors
A game does not play a file top to bottom — it starts a specific conversation.
Ink addresses one with ChoosePathString("knot.stitch") and Yarn Spinner with
SetNode("Start"); both resolve a name against the table of named things they
already keep. Here that table is anchors, which maps every scene's slug to its
node, and it is what a host uses to begin at a named scene or to follow a jump.
entry adds the one fact anchors cannot hold: where a playthrough begins when
nothing says otherwise. That is the document's top, which has no heading and so has
no slug.
It is a single field rather than a table of named ways in. A table would have
exactly one member, under a name the compiler never wrote, and would promise a
multiplicity that does not exist — while buying nothing, since a new field is
additive anyway. Starting somewhere else is already expressible through
anchors, and a debugger's "begin here" is an argument to the runner rather than a
fact about the document.
Speakers are hoisted, and addressed by index
A speaker's name and tags are hoisted out of the lines that quote them, so a host has one place to bind a portrait, a voice, or a color, and a diff stays local when a speaker changes.
A line names its speaker by index, the way every other reference in a playbook
works. The alternative — synthesizing a string key from the name — would invent an
identifier nobody wrote, and would spend the word id, which the script language
already uses for the writer's own @id. A host reading a speaker's id therefore
always knows the writer typed it.
Both the @id and the name are optional, because the anonymous default speaker
has neither and still says lines.
Mapping the graph
The writer's whole job is this mapping. Every row is one test.
Nodes
| Graph | kind |
Carries |
|---|---|---|
LineNode |
line |
speaker, speech, condition? |
ChoiceNode |
choice |
ordered |
RandomChoiceNode |
random-choice |
— |
BranchNode |
branch |
— |
ControlNode |
control |
effects, condition? |
EndNode |
end |
— (no outgoing edges) |
Edges
| Graph | kind |
Carries |
|---|---|---|
SuccessionEdge |
succession |
target |
OptionEdge |
option |
target, label, condition? |
RandomOptionEdge |
random-option |
target, weight, condition? |
BranchEdge |
branch |
target, order, condition? |
DivertEdge |
divert |
target, label, condition? |
order on a branch edge preserves if/elseif/else evaluation order, which is
otherwise lost in a JSON array a reader may not be required to keep ordered.
Both label-bearing edges carry their own text, rather than deriving it from the
node they lead to. For an option that is a correctness matter as much as a
convenience: an arm with an empty body leads straight past the choice, so a label
read back off the target would be another line's words. Only ChoicePass, which
still holds the arm's body, can know. DivertPass carries a jump's label for a
different reason — nothing else keeps it, since a jump is written inside a line but
is no part of what that line says.
Speech fragments
| AST | kind |
Carries |
|---|---|---|
Text |
text |
text |
StyledText |
styled |
style (italic, bold, strikethrough), children |
Link |
link |
target, label |
Image |
image |
source, alt |
LineBreak |
break |
— |
Query |
query |
key |
DefaultCommand |
default-command |
action |
CustomCommand |
custom-command |
name, args |
ReservedTag, CustomTag |
tag |
name, value?, reserved |
The two command kinds are spelled the way the guide names them, rather than shortened to command and call — the second is a word the project does not use with writers, and the pair reads better matched.
A fragment's enum values — style's italic, bold, and strikethrough — are
pinned by hand-written converters, so a C# rename cannot change what a playbook
says; see Enum wire names.
Fragments nest — StyledText.Children and a link or image label are themselves
fragment lists — so the encoding is recursive. Nothing is flattened to a string,
because a host re-renders it: Godot as BBCode, the report as HTML, the CLI as
ANSI.
Tags stay in the fragment list, in position, rather than being hoisted beside the speaker. A tag is a plausible hook — a host may merely recognize one, or act on it — and both cases need to know where in the line it attached. Hoisting would discard that for a small saving.
A command inside a line's speech is an effect (the graph's LineNode.Effects).
The playbook keeps it in place for the same reason, so a runtime knows where in the
line it fires. How a runner orders speaking and performing it is not settled: the
runner reports it inside Said and does not ask the host to perform
it. A command on its own line is a control node, which the runner does perform.
Conditions and weights
Both are objects with a kind, never bare scalars, so negation and expressions
stay additive:
"condition": { "kind": "key", "key": "IsCurious" }
"weight": { "kind": "auto" }
"weight": { "kind": "number", "percentage": 25.0 }
"weight": { "kind": "query", "key": "Bob.Affection" }
The wire name is condition, matching the Condition type in the AST and the
graph, and the word the guide uses with
writers.
Reading a playbook
The reader is a gatekeeper before it is a parser. It parses, then hands the result
to a checker — one rule, one class — and returns only what passes.
PlaybookCheckerFactory.CreateDefault() wires the standard set, in this order:
| Checker | Refuses |
|---|---|
FormatChecker → VersionChecker |
a format.version outside the range this build reads |
FormatChecker → CapabilityChecker |
a name in requires this build does not support |
NodePositionChecker |
a node whose id is not its position (nodes[i].id != i) |
ReferenceChecker |
an entry, anchor, edge target, or speaker index that lands nowhere |
OutwardShapeChecker |
a node whose ways out break the outward-shape rule — see Playbook Reader Rules |
BranchArmOrderChecker |
a branch whose arms are out of order, lack a gated arm, or put the else before the end — see Playbook Reader Rules |
The version settles before the capabilities, because a document of an unknown shape may describe its capabilities in terms an older build would misread. Each later checker assumes what the ones before it proved.
The rules are injected, not inherited. PlaybookReader takes an
IPlaybookChecker; PlaybookCheckerFactory.CreateDefault() wires the standard
set, and a runtime that reads fewer constructs than this build supplies its own.
Each rule takes the policy it applies — a version range, a capability set —
rather than reading a constant, so nothing has to be subclassed to be narrowed.
A checker refuses at the first thing it finds wrong rather than gathering a list. The compiler reports every diagnostic because a writer wants the whole set; a malformed playbook is compiler output, so there is nothing for a reader to work through and fix. Every failure names the offending value and the expectation. Nothing degrades, because a skipped condition does not error — it silently tells the wrong story.
The reader is also where the format's forward compatibility lives: unknown
object properties are ignored, so a newer compiler may add optional metadata
without breaking an older reader, while an unknown entry in requires is a hard
refusal.
Version 0 defines exactly two capability names:
| Capability | Meaning |
|---|---|
core |
Everything the compiler emits. Always present in requires |
cross-file-jump |
Reserved, and neither emitted nor read in version 0. A node reference is a plain index; a reference into another script waits for the linker to settle what a script identity is. Naming it keeps it unclaimed and makes the widening additive |
Key design decisions
P1 — The playbook is a designed contract, not a graph dump
The writer maps each internal type onto a public shape by hand. Serializing
DialogueGraph directly would be faster to build and would inherit SourceSpan,
SpeakerSymbol, and every future refactor of compiler internals as a breaking
format change — the coupling
this design exists to avoid.
An explicit mapping is also a place to put a test per construct.
P2 — System.Text.Json polymorphism with a kind discriminator
[JsonPolymorphic(TypeDiscriminatorPropertyName = "kind")] plus
[JsonDerivedType(typeof(LinePlaybookNode), "line")] round-trips the node, edge,
fragment, condition, and weight unions with no custom converters. The discriminator
is spelled kind rather than the library default $type because the format is a
public contract, not a .NET serialization detail; a TypeScript reader switches on
the same word.
No hand-written converter is needed. A node reference is a plain integer index, so every value in the format is either a primitive or a tagged object — see P10.
P3 — The playbook speaks the project's own words
Every wire name is the word the project already uses: target for an edge's
destination (Edge(NodeId Target), Link.Target, Jump.Target), condition for
what must hold (the Condition type, and three notes titled Conditional …), and
query for a pure read of the world.
Inventing wire names is how a format drifts from the language its users speak.
P4 — What a runner can derive is not stored
A node does not carry the query keys needed to leave it. They are derivable by walking the node — its condition, its speech, its effects, and every out-edge's condition and weight — so storing them would put one fact in two places, where the two can disagree.
Storing them buys nothing measurable either. A runner walks each node once at load
and caches the result, so there is no play-time difference. Resolving a node's
reads in one round trip comes from the protocol's Resolve/Supply pair, not
from a field in the artifact.
The walk is subtler than it looks — a query can hide inside an option's label — and that is exactly what the conformance corpus exists to keep honest across runtimes.
P5 — Node ids are dense indices
A playbook node's id is its position in the nodes array, and the reader refuses
a document where it is not.
This is a renumbering, not a passthrough. The compiler's NodeId is by contract
an opaque handle — DialogueGraph resolves it through an id-keyed dictionary and
its documentation says outright not to treat the value as a list index — and ids
are minted in the order blocks are encountered, which is not the document order
the node list is in. So the writer must translate either way; the only question is
what it translates into.
A dense index is chosen because it makes the document verifiable. One comparison per node proves the ids are unique, gapless, and correctly ordered at once:
| Failure | Carrying the opaque ids | Dense indices |
|---|---|---|
| Duplicate id | Needs a separate uniqueness pass | Caught by nodes[i].id == i |
| Gap in the numbering | Legal, and undetectable | Caught by nodes[i].id == i |
| Node list silently reordered | Undetectable | Caught by nodes[i].id == i |
That last row is the one that matters: a reordered array is a valid playbook that
tells a different story, exactly the failure this format exists to prevent. The
explicit id costs a few bytes and buys a checksum — as well as a document that
reads well and answers jq '.nodes[] | select(.id == 42)'.
Resolving a reference then costs an array index rather than a dictionary lookup, but that is a bonus, not the reason. A dialogue graph is hundreds of nodes; the lookup was never the problem.
The usual argument for preserving original ids — correlating a runtime error with a
compiler diagnostic — does not apply, because NodeId is internal and never
surfaces. DialogueDown's diagnostics address source positions, not nodes.
P6 — Anchors, not the region tree
Play needs to answer "which node opens #the-inn", which is a flat anchors
table. The full RegionTree — nesting, OwnNodes, scene labels — serves analysis
and presentation, so emitting it would ship data with no consumer. Adding it is
additive under an advisory uses entry.
P7 — The schema is the specification; conformance is proven, not generated
A hand-written JSON Schema 2020-12 is the format's normative structural
specification. It is what a porter reads, so every field carries a description —
which a generated schema cannot provide.
The C# types are hand-written too, as sealed records with
[JsonPolymorphic]. Neither direction of code generation is worth its cost here:
generating C# from the schema produces mutable POCOs, and generating the schema
from C# needs .NET 9 (the shipped libraries still target net8.0) and yields an
undocumented schema.
The two are kept honest by validating real output: every golden playbook is checked against the schema in CI, so a drift in either direction fails the build. A TypeScript runner would generate its types from the schema — the payoff that makes a hand-written spec worth writing.
This splits responsibility cleanly:
The schema is normative for structure. The conformance corpus is normative for behavior.
P8 — No validator ships with the reader
Schema validation stays in authoring, editor tooling, and CI; the reader relies on typed deserialization plus the explicit semantic checks below. That is what every comparable format does — glTF ships a separate validator tool and its loaders (Three.js, Babylon, Unity, Godot) never schema-validate; Yarn Spinner's schemas serve its VS Code extension and CI, not its runtime.
It is also what a schema cannot do that decides it. Structure it handles well,
and the shipped schema proves it: kind enums, per-kind required fields,
recursion through nested speech, and even "an end leads nowhere" as a maxItems.
Relational integrity it cannot express at all — nodes[i].id == i, references
landing in range, or "an unknown requires refuses while an unknown uses does
not". Those need code regardless, so a validator dependency would add weight
without removing work.
Keeping it out also sidesteps a licensing trap worth recording: JsonSchema.Net
attaches a EULA to its binaries from v9.0.0 that asks revenue-generating users to
pay, and Newtonsoft.Json.Schema is AGPL below a paid tier with a
ten-validations-per-hour cap. Neither is acceptable to inherit into a game.
P9 — Human-readable by default
The writer pretty-prints. A playbook is meant to be opened, jq-ed, and reasoned
about, and readability is worth more than bytes for a file that measures in
kilobytes. A compact mode — and, if it ever proves its value, compression or a
binary encoding — stays available behind a CLI flag, because
the writer is a seam.
P10 — Cross-file references wait for the linker
A node reference is an integer index and nothing else. A reference into another script would have to spell out what a script identity is, whether a bare script means its root scene, and how an anchor is written — three answers owned by the linker, which is explored rather than implemented. Encoding guesses about them into a public contract would make the linker inherit them.
Deferring costs nothing, because the widening is already additive: a playbook that
uses cross-file references will declare the cross-file-jump capability, and a
version-0 runner refuses the whole document before parsing a single node. It cannot
misread a reference shape it never reaches. The capability manifest is what makes
cross-file additive — not the shape of the reference field.
P11 — Absent is absent; the writer emits the shortest true document
A field whose value is the default is omitted: no null, and no false. A
tag without a value is { "name": "aside" }, and a reader treats a missing field
exactly as it treats a missing requires — as the default case.
Writing both spellings would let two documents mean the same thing, which doubles what a schema, a reader, and a golden file each have to say.
Warning
Omitting defaults is applied per flag, never as a blanket serializer setting.
The blanket condition also drops value types equal to zero — which would silently
erase a node's id, an edge's target, and the first branch arm's order.
P12 — Records compare by value
The playbook types are records, which advertise value equality, but every collection
property is an ImmutableArray<T> or an ImmutableSortedDictionary<TKey,TValue>,
whose own Equals is reference equality. Left alone, two structurally identical
playbooks compare unequal, so == on a public contract means something narrower
than it appears to.
Generator.Equals supplies the equality instead. Each collection-owning record is
partial and marked [Equatable], and each collection property carries
[OrderedEquality] or [UnorderedEquality]; the generator emits Equals and
GetHashCode at build time. Its analyzer fails the build when an [Equatable] type
declares a collection property without one of those attributes, so a property added
later cannot be forgotten.
This is the contract's one exception to "nothing outside System". The generated
code calls comparer types from Generator.Equals.Runtime, which every consumer and
embedded game therefore carries — about 18 KB, plus Microsoft.Bcl.HashCode (about
20 KB) that the package's netstandard2.0 target pulls in. The architecture test
allows exactly that dependency. Hand-writing Equals and GetHashCode per record
would keep the assembly reference-free, but it repeats the same code across a dozen
records and guards the forgotten property only with a test; the generator's
compile-time check is the stronger guard, and the dependency is small.
Error and boundary cases
| Case | Behavior |
|---|---|
format.version newer than the reader |
Refuse, naming both versions |
format.version older than the reader's floor |
Refuse |
Unknown name in requires |
Refuse, naming the capability |
Unknown name in uses |
Accept — advisory by definition |
| Unknown object property | Ignore — forward compatibility |
nodes[i].id != i |
Refuse |
| Node reference out of range | Refuse |
| A node reference that is not a number | Refuse — the deserializer types it as an integer, so a string is not a reference at all |
An entry pointing nowhere |
Refuse — a playbook nothing can start is not playable |
| Duplicate speaker id | Refuse; the writer asserts uniqueness before emitting |
| Duplicate anchor | Cannot occur; the compiler already rejects it (DLG2001) |
| A script that compiles with errors | No playbook is written; --output is untouched |
| A script that compiles with warnings | A playbook is written. Warnings are a smell a compiler tolerates; anything intolerable belongs in the error tier |
| A script with no dialogue | A valid playbook with an entry that reaches end |
Empty speech on a line |
Cannot occur; the AST rejects empty styled content |
Integration
| Seam | Change |
|---|---|
CompilationSuccess |
Unchanged. The writer consumes its internal graph inside the same assembly |
IPlaybookWriter |
New public seam in DialogueDown, registered in AddDialogueDown and the CLI composition root, following the IDialogueGraphBuilder pattern |
CompileCommand |
A playbook is what compile emits unless told otherwise: -o names where it goes, and --emit dot asks for the stage graphs instead. |
| CLI presentation | Nothing in the CLI reads a playbook, so there is no load-time failure to render. A command that plays one renders it outside the DLG code space, in the style of CLI Diagnostic Rendering |
DialogueDown.csproj |
References DialogueDown.Playbook; the package ships both |
| Central package management | DialogueDown.Playbook inherits Directory.Packages.props: System.Text.Json (net8.0) for a framework attribute, Generator.Equals (private, build-time) to generate equality, and the Generator.Equals.Runtime it calls |
| CI | A check-jsonschema step validates every golden playbook against the local schema file, so validation never depends on the network |
| Editors | Emitted playbooks carry a versioned $schema URL published with the existing GitHub Pages site, so VS Code validates a playbook wherever it lands. See open questions for zero-config registration |
Testability
| Level | What it covers |
|---|---|
| Unit — writer | One test per row of Mapping the graph: each node, edge, fragment, condition, and weight kind |
| Unit — reader | Every refusal in Error and boundary cases, each asserting the message names the offending value |
| Round-trip | Compile, write, read, and assert the JSON comes back identical — the primary safety net, and cheap because both directions land here |
| Equality | Records compare by value across every collection; guarded by the analyzer, a reflection test, and a corpus round-trip |
| Exhaustive | Reflection over each closed union, so a construct added to the AST fails here rather than at whatever runtime reaches it first |
| Golden | A committed playbook per compiling examples/*.dialogue.md, so a format change is a reviewable diff |
| Schema | Every golden playbook validates against the schema in CI |
Round-trip tests live in DialogueDown.Tests, which already sees internals and can
reference both assemblies. Playbook fixtures are built through a shared factory so a
shape change touches one file.
A round trip is asserted as text, not as objects: the writer's JSON must read back into a document that writes the same JSON again, so a change in how a field is spelled or ordered is caught rather than hidden. Where a test asks a question about the model, the records compare by value (see P12) and can be asserted directly.
Goldens use Verify, which supplies the matching and the accept workflow. It was measured against this suite before being taken: it compares text rather than reparsing the JSON — which is what makes a formatting change visible at all — and one golden serves both target frameworks.
Golden playbooks churn when node positions shift, which is expected: they are a build artifact nobody hand-edits. The semantic regression asset is the hand-authored conformance corpus, which pins what a playbook means rather than what it serializes to.
The exhaustive tests are the ones that repaid the most. Asking reflection for every
member of InlineFragment turned up three the mapping table never listed — a jump,
a condition, and a jump indicator — and so caught, before any of it shipped, that a
jump survives into a line's speech and would have thrown on every script containing
one.
Open questions and deferred work
- Zero-config editor support. A versioned
$schemaURL works wherever a playbook lands, but a reader must still know to look. Registering*.playbook.jsonwith SchemaStore would make VS Code validate a playbook with no$schemakey at all — the best experience for a non-developer. It needs the URL to be stable first. - Line identity stays deferred. Nothing needs
reserving: the schema allows properties it does not name, and a reader ignores
them, so adding
lineIdis a field to populate rather than a shape to change. - The schema has no negative tests. Seven malformed playbooks were checked by
hand and each was refused, but nothing stops an edit turning the schema into a
rubber stamp. CI runs
--check-metaschema, which catches structural breakage and not a weakened rule. Committed counter-examples would close that. - A divert's label is not drawn on the graph stage of the report. The words survive on the edge and in the Dialogue AST stage, but the graph view would read better with them.