Design¶
The document format¶
One topic, one source, one file:
# <Title>
Source: <the URL the text was checked against>
One paragraph per claim, blank lines between them. Only what the source supports.
- Line 1 is the title and line 2 the
Source:URL: the app labels every retrieved passage with both, and an answer's citation is the source.scripts/corpus-check.shfails a file without them. - One source per document. A second source means a second document, so a citation always names the page the sentence came from.
knowledge/safety/holds safety topics (MIP-0022). The app appends an emergency footer to any answer grounded in one; never write that footer into the Markdown.
Chunking¶
marola-app's
Corpus
turns a document into retrieval chunks:
- It reads every
.mdfile directly underknowledge/and directly underknowledge/safety/. Any other subdirectory is ignored. - It takes the first
#line as the title and the firstSource:line as the source, and drops both from the body. - It splits the body on blank lines, then merges neighbouring paragraphs while the merged chunk stays within 700 characters. A paragraph longer than that stays one chunk; it is never cut.
- Each chunk carries the title, the source, and whether the file is under
safety/.
So a paragraph is the smallest unit retrieval can return: keep each one about one claim, readable on its own.
knowledge/README.md is loaded too: it has a title and no Source: line, so its chunks carry an
empty source. Giving it a Source: line would make it citable.
Which model embeds the chunks is the app's setting: Choosing the embedder.