skip to the page

What every ask leaves behind

The pages before this one were about the runtime deciding things, and about what your own code does with the answer. This one is about what it remembers.

Most of what happens when somebody asks for a change never reaches the page. A sentence becomes a plan, the plan gets measured, the Gate weighs it — and very often nothing is applied at all. The page's history has no room for any of that, because a history of a page is a list of what happened to the page.

So there is a second record, and it is the one that keeps the asks.

The smallest tree that renderslive · rendered through the runtime

Hello from a tree

Nothing here was written as markup.

Propose a change:

Every one of these goes through the same pipeline a model’s answer would: interpreted, analysed, weighed, judged, applied, appended. The log is real and it is in this tab — reload the page and the example is back as it was.

show the tree
{
  "kind": "element",
  "id": "n_firsttree5",
  "type": "loom.page",
  "props": {
    "fills": true,
    "width": "readable",
    "loom:theme": {
      "palette": "minimal",
      "fontPack": "minimal-sans",
      "stylePreset": "precise"
    }
  },
  "children": [
    {
      "kind": "element",
      "id": "n_firsttree2",
      "type": "loom.heading",
      "props": {
        "level": 1
      },
      "children": [
        {
          "kind": "text",
          "id": "n_firsttree1",
          "value": "Hello from a tree"
        }
      ]
    },
    {
      "kind": "element",
      "id": "n_firsttree4",
      "type": "loom.prose",
      "props": {},
      "children": [
        {
          "kind": "text",
          "id": "n_firsttree3",
          "value": "Nothing here was written as markup."
        }
      ]
    }
  ]
}
A page with a heading and a sentence. Every node names a registered primitive and carries props that primitive declared.

Every number on this page comes from that tree. The site opened a journal, asked it for ten changes over four weeks, and everything below is read back out with the runtime's own functions.

Two records, and only one of them is about the page

They are easy to confuse and they answer completely different questions.

The logThe journal
Kept by@jam-overture/loom/store@jam-overture/loom/telemetry
One entry perchange that was appliedthing that was asked
Holds a refusalneveralways
What it is forrebuilding the pageunderstanding the asking
May be deletednoyes, and that is the point

The history of a page is the first column. This page is the second.

The line that matters is the third row. A refused change never touches the log — the log is the truth about the page, and a change that was not applied is not part of that truth. If nothing else wrote it down, nobody could ever tell you how often the Gate says no, which is the first question anyone asks after turning it on.

Turning it on is six lines and one decision

import { commitIntent } from "@jam-overture/loom/write"
import { collectTelemetry } from "@jam-overture/loom/telemetry"
import { postgresTelemetryJournal } from "@jam-overture/loom/telemetry/postgres"

const telemetry = collectTelemetry(postgresTelemetryJournal(db))

const runtime = { interpreter, policySource, clock, idFactory, events: telemetry.sink }

const outcome = await commitIntent({ store, holds, runtime }, intent)

await telemetry.flush()

telemetry.sink goes in the events slot the runtime already has. That is the whole integration — there is no second place telemetry is wired in, and nothing between the intent and the outcome knows it is being watched.

The decision is the last line. Emitting does no I/O. Every stage of the request puts a record in a list and returns immediately; the writing happens once, at flush, at a point you choose and can await. Both halves of that are deliberate:

  • An accepted change should not wait on a database to say it was accepted.
  • A write nobody awaits is a promise a serverless platform cancels the moment the response ends — which loses telemetry silently, on exactly the requests that finish fastest.

What a record keeps

Here is one, printed exactly as the journal holds it — the moment the first of the ten asks arrived:

One record, exactly as the journal holds it
{
  "treeId": "t_firsttree1",
  "occurredAt": "2026-07-01T09:00:01.000Z",
  "event": {
    "type": "intent-received",
    "intent": {
      "intentId": "i_corpusaddasentenced01",
      "origin": "user-instruction",
      "actor": "the reader",
      "baseRevision": 0,
      "utteranceLength": 43,
      "observedAt": "2026-07-01T09:00:00.000Z"
    }
  },
  "seq": 1,
  "recordedAt": "2026-07-01T09:00:10.000Z"
}

Read it once more and notice what is not there.

The person typed "Add a closing line to the end of this page." The record says utteranceLength: 43. The sentence is not kept. Telemetry is aggregated and retained for weeks; what somebody typed is content, and content does not belong in a table you keep in order to count things.

Three rules decide what crosses from the runtime's internal narration into a record, and each of them is one sentence:

  • A proposal's delta is kept. For a change that was refused, this journal is the only place it ever existed.
  • Anything the log already holds is dropped. Storing an applied change's inverse twice invites two stores to disagree about it.
  • The utterance is not kept. A length is enough to tell a one-word ask from three paragraphs.

A record is not the story. An episode is.

Records arrive one stage at a time, and no single one of them answers a question anybody has. Here is one ask — the one the Gate held and a person then allowed — as the twelve records it wrote:

12 records, in the order they were written
#26intent-receiveduser-instruction from the reader, against revision 2
#27policy-resolvedjudged under "loom-docs"
#28change-proposed1 operation, confidence 1, authored by the runtime
#29change-assessedstakes high, reversible, touching loom.heading
#30disposition-decidedrequires-confirmation — touches protected loom.heading
#31proposal-heldtaken into custody, waiting for a person
#32hold-confirmedallowed by the maintainer
#33policy-resolvedjudged under "loom-docs"
#34change-assessedstakes high, reversible, touching loom.heading
#35disposition-decidedrequires-confirmation — touches protected loom.heading
#36change-appliedapplied in memory at revision 3
#37change-committedwritten to the log at revision 3
The same 12 records, folded into one episode
resolution.kind
committed
policyId
loom-docs
intent.utteranceLength
49 characters
assessment.stakes
high
disposition.kind
requires-confirmation
held
true
answeredBy
the maintainer
committedRevision
3

Two things in that list are worth stopping on.

It is two requests, not one. Records 26 to 31 are somebody clicking; record 32 is a different person, minutes or days later, saying yes. The fold does not care that they were separate — it is one ask, with one ending.

The Gate ran again. Look at records 33, 34 and 35: after the confirmation, the policy is resolved again and the change is assessed again, against the page as it now stands. A person saying yes is permission to proceed, not permission to skip the check. A change the Gate would refuse today stays refused however enthusiastically it was approved yesterday.

episodesOf does that fold, and it is a derivation, not a second store. The journal is what is written; episodes are computed from it when somebody asks. A stored episode table would be a second copy of the same facts, and the two would eventually disagree — silently, because nothing would be checking.

Ten asks, and eight ways they can end

This is the table a deployment actually reads. Every ending the runtime can reach, with what this site's corpus did:

10 asks over four weeks, by how they ended
committedThe change was allowed and is in the page's history. This is the only ending that changed anything.4
refusedThe Gate said no. Nothing was applied, and this journal is the only place the attempt exists.2
awaiting-answerThe Gate would not decide alone and put the change in front of a person. Nobody has answered yet.1
discardedA person was asked, and said no.1
not-interpretedNothing could be planned from the sentence. The Gate never saw a change, because there was not one.1
not-writableThe ask named a revision the page has moved past. It was turned away before anything was planned.1
failedSomething broke on the way — a store, a hold, a commit. Not a verdict about the change.0
openThis window does not contain the ending. The ask is still running, or the page of records stops mid-story.0
8 proposals across 10 asks — 3 of them held for a person, 0 of them a second attempt at something refused. The two numbers differ because an ask that was turned away, or that nothing could be planned from, never produced a proposal at all. A refusal does: the Gate has to see a change before it can say no to one, and this journal is where that change is kept.

Four of those eight are not verdicts, and folding them into "it failed" is the most common mistake a host makes here:

  • not-interpreted — nothing could be planned from the sentence. The Gate never saw a change, because there was not one. A rising count is a signal about the model, not about your users.
  • not-writable — the ask named a revision the page has moved past. Somebody's tab was open too long. Nobody did anything wrong.
  • failed — a store, a hold or a commit broke. It says nothing at all about whether the change was a good idea.
  • open — the window you read does not contain the ending.

The zeroes are printed rather than hidden. An ending that has never happened and an ending this table forgot are different claims, and a reader looking at six rows would have no way to tell which they were being shown.

Was the confidence worth anything?

Connecting a model ends on a promise it does not keep there. Of the interpreter this site runs on, it says a confidence nobody graded

must not walk into calibration as a model's perfect record.

Calibration is this report, and that sentence is a rule about what may enter it. A model's own confidence is worth something on one condition — that it be measured against what actually happened — and this is the measuring.

The idea is one sentence. Take every proposal that claimed about 0.9, and see what fraction of them actually survived. If a hundred claimed 0.9 and seventy survived, the model's 0.9 is really a 0.7 and the gap is 0.2. Do it for ten bands and you know whether a number means anything at all.

Now here is that report over this site's own journal:

Calibration over this site's own journal

overall.judged

0

Claims with a verdict to score them against.

runtimeAuthored

8

Proposals set aside because nobody graded them.

overall.gap

null

No claims, so no distance between claimed and observed.

Zero out of zero is null, not 0. A report that rounded an absent measurement down to a perfect one would be the single most misleading number this package could produce.

Nothing was judged, and that is the correct answer. This site's interpreter does not guess — it computes its operations from the tree, stamps authoredBy: "runtime" and a confidence of 1, and eight proposals went through the Gate that way. calibrationOf sets every one of them aside and reports the count, because scoring them would measure a constant the runtime stamps rather than a claim anybody made.

That is the whole discipline in one number. A confidence that nobody graded is not a claim, and a report that let eight of them in would show a model with a flawless record that does not exist.

Two things a real report splits that a first reading would pool. Unjudged proposals are counted separately — a commit that failed on the database says nothing about whether the model was right, and folding those into the rejections makes a runtime look overconfident every time infrastructure breaks. And claims are grouped by the policy that judged them, because survival is not a property of the model alone: a host that tightened its Gate moves the observed rate without the model having changed at all.

Forgetting is a feature, and it has rules

A journal only grows. Retention is how it is allowed to shrink, and it is the only destructive operation anywhere in Loom's storage.

import { applyRetention } from "@jam-overture/loom/telemetry"

const pruned = await applyRetention(journal, { policy: { maxAgeMs: 30 * 24 * 60 * 60 * 1000 } })

Here is what one run against this site's corpus would do, under a fortnight's policy:

60 records, 14-day policy, planned on 2026-08-01
horizon
2026-07-18 — the age the policy sets. Nothing recorded after this is a candidate, whatever else is true.
forgets
14 — records this run would drop, everything below position 15
keptUnsettled
6 — old enough to go, kept because their own ask is still waiting on a person
keptBehind
24 — old enough to go and finished, kept only because the unsettled ask sits in front of them. Forgetting is a prefix, so this is the price of never cutting a story in half.
unsettledEpisodes
1 — asks inside the candidate window that nothing has settled

Two rules produced those numbers and neither is obvious.

Forgetting is a prefix. A journal drops its oldest records or none; it never punches holes. A hole would make a record's position start meaning something a reader could misread, and it would make "is this window complete?" unanswerable.

An episode is never cut in half. A window that kept a commit but forgot the proposal it committed reads as a change nobody proposed — which is precisely the fault episodesOf reports as unattributed. A journal that manufactured the exact fault its own fold exists to detect would not be worth trusting.

Put together, they explain keptBehind. Those 24 records are old enough to go and completely finished; they are still here because one ask from three weeks ago is still waiting for somebody to answer it, and it sits in front of them. That number is a symptom, and a host watching it grow has learned something useful about their review queue rather than about their database.

What this page is not showing you

The journal here is memoryTelemetryJournal, which is a real implementation and lives for as long as this page took to build. A deployment swaps one line for postgresTelemetryJournal(db) — the same database the store uses — and nothing above that line changes. Both satisfy one interface and one shared contract test, which is what makes the swap boring.

The corpus is ten asks. A real one is thousands, and the two numbers that get interesting at that size are the ones this page can only describe: the calibration gap per confidence band, and how the refusal rate moves when you change your policy.

And 12 of 18 event types appear in it. Ten ordinary asks against a three-node page exercise two thirds of what the runtime can narrate; the rest are failures nothing here provoked. The full list is in the reference, generated from the schema that validates them.