AI and LLMs
An STXT document generated by a model is validated against a template:
the output is verifiable, and every error goes back to the model with its line and code.
A language model that generates a document, or extracts one from free text, produces text. If that text is Markdown, there is nothing to check it against: any output passes, with invented fields, missing sections or badly written dates. If it is STXT with a template, the output is validated like any other document, and whatever does not comply with the template is an error with a line and a code.
What the validator checks
- The closed content model rejects undeclared node names
(
CHILD_NOT_DECLARED,NODE_NOT_DEFINED_IN_SCHEMA). - Cardinalities reject missing or surplus nodes (
TOO_FEW_CHILDREN,TOO_MANY_CHILDREN). - Types (
DATE,EMAIL,NUMBER...) reject wrong formats, andENUMs reject values outside the list (INVALID_VALUE).
The template of this page's executable example, as it is in its .stxt/:
Template (@stxt.template): com.acme.reports
Structure >>
Report (com.acme.reports):
Title: (1)
Date: (1) DATE
Author: (1)
Status: (1) ENUM [Draft, Final]
Summary: (1) TEXT
Section: (*)
Title: (1) @Title
Content: (1) TEXT
Action: (*)
Owner: (1)
Due: (?) DATE
Description: (1) TEXT
Description >>
Report: A short internal report extracted from free text (notes, emails, minutes)
Date: Date of the source material, YYYY-MM-DD
Status: Draft until a person has reviewed it
Summary: Three sentences at most
Section: One per topic discussed
Action: One per agreed follow-upAn output with four errors, all of them caught:
Report (com.acme.reports): Market analysis
Title: Market analysis
# ERROR: DATE requires YYYY-MM-DD
Date: tomorrow
Author: Ana García
# ERROR: not a value of the ENUM
Status: Pending
# ERROR: undeclared child (closed model)
Subtitle: Short version
# ERROR: Summary, which is mandatory, is missingThe loop: generate, validate, fix
- The model generates the STXT document.
- The validator checks it against the template.
- If there are errors, they go back to the model as they are, with line and code.
- The model fixes them and the loop repeats, until the document validates or the attempts run out.
Step 2 is stxt validate -: the CLI reads the document from standard input,
validates it against the grammars of the current directory and prints each finding
as <stdin>:line: [CODE] message (error), with exit code 1 if there are errors
(The command line). With the library, it is UnifiedSchemaProvider to
load the template and Parser with SchemaValidator to get the same errors.
The complete loop is in
examples/llm/
of the specification repository: a Node program that turns free-text meeting notes
into a valid Report (com.acme.reports). Alongside it are the template above, a
correct example document, the prompt with the rules of the language and the source
notes. The prompt carries the STXT rules, the whole template, the valid example and
the text to convert, and asks for the document alone.
$ node generate.mjs > report.stxt
--- attempt 1: 3 error(s)
line 3: [INVALID_VALUE] Date: Invalid date (12 March 2026) (schema)
line 5: [INVALID_VALUE] The value 'draft' is not one of the allowed values of Status (schema)
line 1: [TOO_FEW_CHILDREN] 0 nodes of 'com.acme.reports:summary' and min is 1 (schema)
--- attempt 2: valid
The document goes to standard output and the dialogue to standard error, with exit
code 0 only if the last version validates. The example uses the Anthropic API, but
nothing in the loop depends on the provider: the call to the model is a function
that returns text. The directory's README summarizes which model mistake produces
which code.
What makes generation easier
- Few line forms. Every line is
Name: value,Name:,Name >>or the text of a block, and there are no alternative syntaxes for the same element. The document is generated top to bottom, with no backward references. - No escaping. A
>>block is literal text: the model only has to indent it. In JSON, the same text requires escaping quotes and line breaks:
{"summary": "Line 1\nLine 2 with \"quotes\" and more text..."}
- Names are compared by their canonical form.
TITLE:,Title:andtitle:are the same node (STXT-SPEC §4.3); the values of anENUM, on the other hand, are compared exactly, case included.
Prompting practices
- Include the complete template of the namespace: it declares which nodes exist, their cardinalities and their allowed values.
- Add one example document, complete and valid.
- Ask for a fixed indentation: tabs, or four spaces per level.
- Always validate the output with the parser, never by inspection.
- Return the validator's messages to the model as they are, with line and code.
The cases where this flow fits are developed in CMS and publishing, where the build rejects generated pages that do not validate, and in Corporate documents, for extracting data from free text.