skip to the page

Connecting a model

Everything on this site so far has been the runtime being sure. It reads a plan, weighs it, judges it, applies it, records it — and none of that involves guessing about anything.

There is exactly one step that guesses, and it is the first one: turning "add a closing line to the end of this page" into a list of operations. This page is about putting a real language model in that slot — what it gets shown, what it is allowed to say back, what that costs, and what to do when it fails.

The smallest tree that renderslive · rendered through the runtime

Hello from a tree

Nothing here was written as markup.

Propose a change:

Every one of these goes through the same pipeline a model’s answer would: interpreted, analysed, weighed, judged, applied, appended. The log is real and it is in this tab — reload the page and the example is back as it was.

show the tree
{
  "kind": "element",
  "id": "n_firsttree5",
  "type": "loom.page",
  "props": {
    "fills": true,
    "width": "readable",
    "loom:theme": {
      "palette": "minimal",
      "fontPack": "minimal-sans",
      "stylePreset": "precise"
    }
  },
  "children": [
    {
      "kind": "element",
      "id": "n_firsttree2",
      "type": "loom.heading",
      "props": {
        "level": 1
      },
      "children": [
        {
          "kind": "text",
          "id": "n_firsttree1",
          "value": "Hello from a tree"
        }
      ]
    },
    {
      "kind": "element",
      "id": "n_firsttree4",
      "type": "loom.prose",
      "props": {},
      "children": [
        {
          "kind": "text",
          "id": "n_firsttree3",
          "value": "Nothing here was written as markup."
        }
      ]
    }
  ]
}
A page with a heading and a sentence. Every node names a registered primitive and carries props that primitive declared.

The box above is doing every step except the guess. Read on and you can see the exact request the guess would have been.

The slot, and what goes in it

An interpreter is an interface with one method:

type ChangeInterpreter = {
  interpret: (intent: EditIntent, tree: LoomTree) => Promise<Result<ProposedChange, InterpretationError>>
}

A sentence and a page go in; a proposal or a reason comes out. Nothing else in the runtime knows what is behind it. That is not a small property — it is what lets the Gate, the log and every test in this repository run without a network.

Loom ships one implementation, modelInterpreter, and it is deliberately not wired to a vendor. It takes a ModelClient: one method, one request, one reply. The Anthropic adapter is the only file in Loom that knows a specific vendor exists, and it lives behind its own entry point so a host that brings its own model never loads it.

import Anthropic from "@anthropic-ai/sdk"

import { modelInterpreter, randomIdFactory, systemClock } from "@jam-overture/loom"
import { anthropicModelClient } from "@jam-overture/loom/anthropic"
import { catalogueOf } from "@jam-overture/loom/sdk"

const anthropic = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY })

export const interpreter = modelInterpreter({
  client: anthropicModelClient(anthropic.messages),
  idFactory: randomIdFactory,
  clock: systemClock,
  catalogue: catalogueOf(registry),
  themeCatalogue: themes.catalogue(),
})

That object is a ChangeInterpreter. It goes into the runtime you already have, in the slot the preset was in, and nothing downstream changes.

Two of those arguments are the interesting ones

catalogue and themeCatalogue are optional, and leaving them out is legal and usually wrong.

Without a catalogue the model is told nothing about what your deployment can build with, so the only primitives it knows are the ones already on the page. Without a theme catalogue, "make it warmer" has nowhere to go: the tree names one palette and nothing says another exists, so the model either leaves it alone or invents an id that fails to resolve.

Give it both, and it is choosing from a list you wrote.

What the model is actually shown

Not a screenshot, and not your JSX. Five blocks of plain text, in this order — generated below from the same functions the runtime calls, against the same registry the example above renders through, for the sentence on its first chip:

The standing instructions

The same words on every request, so a provider's cache can hold them and so what Loom asks for is one string somebody can read.

2,816 chars
You translate a request about a user interface into discrete edits to a component tree.

The tree is given as an indented outline. Every line starts with a node id. There are three node kinds:
- element — an instance of a primitive, with props and ordered children
- text — a leaf string
- slot — a named region whose children are fallback content
… 24 more lines

What it may build with

Your registry, one line each. A model cannot invent a primitive because it is never shown one it does not have.

19,477 chars
- loom.page — The root of a page. Mounts the theme and stacks its children in one column. props: fills?, width?
- loom.nav — The bar across the top of a page: a brand region, loom.link children as the menu, and an actions region. props: align?, position?, tone? slots: brand, actions
- loom.menu — A run of loom.link behind one button, dropped over the page — the nested menu in a header or the folded column in a footer. props: align?, placement?
- loom.banner — The announcement strip above a page: a short message as children, and a region for the one thing to do about it. props: align?, label?, tone? slots: action
- loom.section — A band of the page: an optional heading region above its content. props: align?, anchor?, eyebrow?, tone?, width? slots: heading
- loom.split — Two side-by-side regions, start and end, that stack when the page is narrow. props: align?, ratio?, reverse? slots: start, end
… 100 more primitives

What it may theme with

Palettes, font packs and style presets by id and description — never a color, which is the whole of the theming bargain.

6,155 chars
Palettes:
- minimal — Minimal. White paper, black ink, one green. Components are outlined rather than filled; the green appears only as rules, rings and tinted marks.
- editorial — Editorial. Restrained, high-contrast magazine palette with a single muted accent.
- bold — Bold. Black canvas with bright yellow and red accents. High-energy, graphic.
- paper — Paper. Warm ivory with a terracotta accent. A printed page rather than a screen.
- slate — Slate. Cool grey with an indigo accent. Restrained, technical, quiet.
… 48 more lines

The page as it stands

The tree above, as an indented outline. Every line begins with the node id an operation would have to name.

367 chars
tree t_firsttree1 revision 0
n_firsttree5 element loom.page fills=true loom:theme={"palette":"minimal","fontPack":"minimal-sans","stylePreset":"precise"} width="readable"
  n_firsttree2 element loom.heading level=1
    n_firsttree1 text "Hello from a tree"
  n_firsttree4 element loom.prose
    n_firsttree3 text "Nothing here was written as markup."

What was asked

The sentence somebody typed, and who was doing the typing.

71 chars
Request (user-instruction): Add a closing line to the end of this page.

Four things are worth noticing before the detail.

The page is an outline, and every line starts with an id. That is the whole addressing scheme. An operation says remove n_firsttree2, and it can only do that because the model was handed that id and told to copy it exactly.

The instructions are a constant. The same words on every request for every tree, which is why they sit at the front where a provider's cache can hold them, and why "what are we asking the model to do" is a string you can read rather than a template you have to reconstruct.

Your registry is the vocabulary, stated in full. A model cannot propose marquee.blink because it has never been shown one. This is the bounded vocabulary bargain doing its work in the most literal way it ever does.

No colors appear anywhere. Themes are given as ids and the sentence their author wrote about each. A model choosing between your registered palettes is the deal; a model choosing hex is the thing that deal rules out.

What it is allowed to say back

The reply is constrained to a JSON schema, and the schema is narrower than the tree model in three ways that all matter:

  • An inserted node carries no id. Identity is minted by the runtime, always. A model that could name a new node could collide with an existing one, and the provenance of an id would stop meaning anything.
  • Props arrive as a JSON-encoded string rather than a typed structure — and nothing is trusted about the contents. They are parsed, required to be an object, and checked value by value before anything reaches a tree.
  • Only elements and text can be inserted. A slot is a region a primitive declares, not content an edit adds.

The reply is also a three-way outcome rather than a delta: change, no-change, or not-understood. "I understood you and nothing needs changing" is an answer, not a failure to answer, and conflating the two would make every satisfied request look like a broken one.

Why the props are a string, which looks like a mistake

Structured output compiles a schema into a grammar, and the grammar has a size ceiling. A typed prop union repeated at every element at every level put the schema four times over that ceiling, which meant no live call ever succeeded. Encoding props as one string costs the grammar a single production however many props there are, and buys back the nesting depth.

The margin is not generous, and it is measured rather than assumed. The reply schema is capped at a nesting depth of 4, which compiles to 3,381 bytes against a budget of 3,500. Depth 5 would be 3,818, and over.

What one request costs

measurePrompt answers this without sending anything, which makes "what did registering another twenty primitives cost me" a number instead of a feeling:

BlockCharactersShare
The standing instructions2,81610%
What it may build with19,47767%
What it may theme with6,15521%
The page as it stands3671%
What was asked710%
One request, with 106 primitives and 51 theme ids registered28,886100%
The same request, with nothing registered3,25411%

The reader's question — the page and the sentence they typed — is about two per cent of the request. Everything else is the vocabulary this deployment chose to offer, and it is re-sent on every proposal and again on every repair.

That is not an argument for registering less. It is an argument for knowing the number, and for putting the stable blocks at the front where a cache can hold them, which is where they already are.

When it fails, and who has to do something

A guess can fail in seven ways. A host that switches on seven codes has to hold all of them in its head and can still get the grouping wrong, so the runtime answers the question that actually matters first: who would have to do something for the next attempt to go differently.

asker
Nothing failed except the asking. Say it differently.
model
The answer that came back. The same question may work on a second try.
provider
The service, this time. Waiting is a reasonable response.
deployment
Your deployment, until somebody changes it. Waiting never fixes this one.
runtime
Loom, for what it sent. Neither waiting nor configuration helps.
CodeWhoseWhat the runtime says
not-understoodaskerthe request was not understood: "make it pop" could mean the heading, the accent or the spacing
no-change-neededaskernothing needed to change: the heading is already the largest thing on the page
malformed-proposalmodelthe model's answer could not be used: operations.0.parentId: not a node id
refusedmodelthe model declined to answer: the model's own safety policy declined
interpreter-unavailableproviderthe model could not be reached, and may answer later: the request timed out
interpreter-misconfigureddeploymentthis deployment cannot reach a model until an operator changes that: no API key was configured
interpreter-request-rejectedruntimethe model service rejected the request, and would reject it again unchanged: the output schema exceeded the grammar budget

Two of those clear on their own. Three of them never do, however many times you retry. Notice what is deliberately not here: there is no retryable flag. Whether to resample a bad answer is a policy, and a host that resamples malformed replies and refuses to resample anything else is making a reasonable choice the runtime should not have made for it.

A second go

A model can be told its proposal was refused and asked to try again more conservatively. That is a separate interface, ChangeRepairer, and a runtime handed no repairer cannot repair — so "this deployment lets AI have a second attempt" is a visible choice where you compose things, rather than a property of whichever interpreter you happened to wire in.

modelInterpreter implements both, so you can offer it or not. A repair sends the whole request again plus the refused delta and the Gate's objection, so it costs about twice what the table above reports.

The one rule the repair prompt insists on is worth reading even if you never turn it on: a revision must be a more conservative way to satisfy the same request — smaller in what it destroys, or narrower in what it touches. It must not be the same change split into a piece small enough to slip through, with the remainder asked for again.

What this site is running, and why it matters

Every chip in every box on this site is a deterministic interpreter. It computes its operations from the tree in front of it rather than guessing them.

That is not a mock. It is a real implementation of the same interface, it re-plans against the page as it stands, it can decline, and everything after it — the analysis, the stakes, the Gate, the log — cannot tell the difference and is not told. Its provenance says authoredBy: "runtime" and a confidence of 1, because a computed delta is not a guess and a confidence nobody graded must not walk into calibration as a model's perfect record.

It is what makes this site possible: dozens of examples that must behave identically on every visit is not a thing to point at a model, and a documentation site whose examples only worked when somebody had configured an API key would be broken for everyone who cloned the repository.

The swap is one argument. Everything you have read on this page, and every verdict you have seen on the others, is what happens on the other side of it.