Document Parts

Every job, agent round and CI run produces text: progress logs, transcripts, token streams, downloaded CI logs. This page is the platform's one shape for keeping that text as durable mesh nodes, searchable while it is still arriving, without flooding mesh_nodes and without touching the original.

The shape

{collection}/_Documents/{slug}                               Document            (mesh_nodes)
{collection}/_Documents/{slug}/_DocumentPart/000000          DocumentPart        (document_parts)
{collection}/_Documents/{slug}/_DocumentPart/000001          DocumentPart        (document_parts)
{collection}/_Documents/{slug}/_DocumentPart/000001/_PartAnnotation/{id}
                                                             DocumentPartAnnotation (document_part_annotations)

Written as the text arrives

The document's own hub is the one writer of the document. A producer — any hub — appends:

// open once (idempotent), then append as text arrives, then seal
hub.OpenDocumentLog(new DocumentLogTarget("Admin/content", "logs/jobs/build-42.log"))
   .SelectMany(path => hub.AppendToDocument(path, chunkOfText, offset))   // offset = producer's char offset
   ...
hub.CompleteDocument(path);                                               // writes the trailing part

DocumentLogTarget fixes the document's windows when it opens (ChunkSize/ChunkOverlap, null = 1000/150) and says whether the producer already stored the original (ProducerStoresOriginal). 🚨 Every member's default is its CLR default on purpose: the mesh serialiser omits default-valued members, so a positional parameter whose default differs from the CLR default arrives flipped at the document's hub (a StoreOriginal: false once arrived as true and rewrote a producer's original). A best-effort producer checks hub.SupportsDocumentLogs() first — it is true only where the content-indexing module (which owns the Document hub) is installed.

As soon as a window is complete (the document holds Start + 1000 characters) the hub writes that part node, embeds it and upserts it into the vector index — right then, not at the end. The unsealed remainder (Tail, under one window) is held on the document node itself, so nothing lives only in memory. Completion writes the trailing partial part, seals the document, records the original's hash and stores the original.

Idempotent by (document, index). A part's path and text are a pure function of the original, so a retried or concurrent flush rewrites the same node with the same content, and the trim that follows a part write is conditional on the document not having advanced past it. An append that carries the producer's Offset skips whatever the document already holds, so redelivery never duplicates text; an offset past the end is a gap and is refused.

Equivalence. However the text is split into appends, the parts are exactly TextChunker.Chunk(fullText, 1000, 150) — the same windows an upload of the finished file produces.

Attaching further node types to a part

A node about ONE chunk — a person's note, an agent's label (root-cause, flake, noise), a finding the bug-fix learning loop consumes — is a DocumentPartAnnotation at {partPath}/_PartAnnotation/{id}. A further attached type needs exactly one more SatelliteTableMapping entry: a segment longer than _DocumentPart (placement picks the longest mapped segment in the path) and the table it should live in. Existing tables are untouched; the provisioning proc adds the new table to every partition.

Placement rule, and the guard

A storage adapter places a satellite by the longest mapped segment anywhere in its path. _DocumentPart outranks _Activity, _Thread and _Comment, so a document filed under a job activity or a thread keeps its parts in document_parts. It does not outrank _ThreadMessage, and it ties _Notification / _UserActivity; DocumentPartPaths.PartPlacementProblem refuses a document filed there, because its parts would be written into another table and never be found again.

The vector index is the existing content chunk index (content_chunks, the store behind search_chunks / get_chunk). A part is upserted under its (collectionPath, filePath, index), so search_chunks finds a running job's log while it is still being written and get_chunk steps through it like any indexed file.

Consumers

Producer Document Original
A job / queue run (the queue work) DocumentPaths.For("{partition}/content", "logs/jobs/{job}.log"); Admin/content for instance-wide jobs stored on completion
An agent round (RoundTranscript) DocumentPaths.For("{threadPath}/transcripts", "{responseMessageId}.md") the response cell is the record (ProducerStoresOriginal)
A CI job log the PR babysitter downloads DocumentPaths.For("Admin/content", "logs/ci/{owner}/{repo}/{jobId}.log") stored on completion
An uploaded file (the content indexer) DocumentPaths.For(collection, file) — parts written after the summary; pruned on a shorter re-upload, removed on delete the uploaded file

A round transcript is best effort: it never fails or holds the round, and an instance turns it off with DocumentLogs:ThreadTranscripts=false.