MIP-0076: Agent routing tooling — a generated wiring table for cross-repo questions, a pinned offline graph for symbols¶
| Status | Accepted (2026-10-03, maintainer review on #635) |
| Author | Claude (Opus 5.5), from the maintainer's graphify question and two offline spikes (2026-10-02, 2026-10-03) |
| Created | 2026-10-03 |
| Phase | None: dev-loop tooling, outside docs/PHASES.md's sequence. No Phase 1 prerequisite, no paid resource |
| Related | MIP-0074 §11 (the follow-up note this answers) and its REPOS.md routing table, MIP-0017 §9 (rejected graphify for the monorepo; §2 says what changed), MIP-0070 §5.4 (the contracts the wiring table describes), MIP-0063 (named graphify as a later harness layer), MIP-0060 (precedent: a third-party agent tool pinned through Nix) |
| Effort | M — one stdlib+PyYAML extractor with fixtures (~200 lines, prototyped), one devkit tool wrapping a nixpkgs package, one plugin skill, a generated block in REPOS.md, one umbrella CI job, and a one-line notify-umbrella change in five repos; no Scala |
| Gain | infra/dev-loop — "who produces, consumes or triggers X" gets a parsed answer in one ~650-token table, and that table cannot drift; "where is symbol X" in an unfamiliar repo gets one bounded query instead of reading trees |
| Effort vs Gain | do when MIP-0074 lands — the wiring table goes into the REPOS.md MIP-0074 creates (task 27, marola-dev/marola#617), and the graph indexes the docs MIP-0074 rewrites |
| Depends on | MIP-0074: REPOS.md (task 27) and the restructured docs. Uses the devkit's existing nixpkgs lock, which already carries graphify 0.9.66. No Phase 1 gate. Creates no paid resource, and §5.2 rules out a model call on the tool's own path |
| Blocked by | 0074 |
| Risk | The graph half goes unused: it answers only symbol questions, 3 of 8 in both spikes, and loses to git grep whenever the agent knows the keyword. The wiring table carries most of the value |
| Cost so far | — |
1. Summary¶
Two tools, split by the kind of question. Cross-repo questions get a wiring table: a devkit
extractor parses every repo's workflows, dispatches, release uploads and pin files into a generated
block in the umbrella's REPOS.md, checked for staleness like the MIP graph. Symbol questions get
graph: a devkit tool running a pinned graphify offline and keyless, with its output outside the
checkout and a small default answer budget. A plugin skill tells an agent which to use, and when
plain git grep is cheaper than both.
2. Motivation¶
The maintainer's question (2026-10-02): as the umbrella and its submodules grow, so does what an agent must read before acting. Could a graph act as a routing table that sends it to the right repo and file, and so avoid context rot?
MIP-0017 §9 rejected graphify because marola was "small enough that Explore/Grep … already gives
an agent working context in seconds". MIP-0070 removed that premise. There are now six repos plus
the devkit, and an agent started in one sees only that tree. A question like "which image does the
map's board use" crosses three repos and two pin files.
Two offline spikes measured it on the real workspace (Appendix), with 0 tokens and no network:
- graphify graphs the whole umbrella in one run: 4,921 nodes and 8,447 edges in 3.7 s on 0.9.74.
- It answers symbol questions only: 3 hits, 3 partials and 2 misses out of 8 real questions, in
both spikes. Every miss or partial was about the wiring between repos. That wiring is invisible to
it:
.ymlis a "document" for its LLM pass, and pin files, justfiles and Dockerfiles are skipped. - Injecting parsed wiring into graphify's graph barely helps
query. One of the four wiring questions moved from miss to partial, because matching is lexical over labels.updatealso evicts injected nodes outside the scan root. - A generated Markdown table of the same facts answers 5 of 8 questions, all the wiring ones, in 2.5 KB read once. graphify answers 3 of 8 at about 1.8 KB per question.
- The extractor found a real gap that grep and the graph both missed:
pointer-sync.ymllistens forsubmodule-updated, but nothing sends it. Every repo'snotify-umbrellasends onlysubmodule-docs-updated, on docs pushes. Confirmed against each repo'smainon 2026-10-03.
3. User-visible change¶
None for marola's users. For an agent or contributor:
$ just wiring # umbrella: regenerate REPOS.md's wiring block
$ just wiring --check # quality-other: fail if the block is stale
$ just graph build # any repo: ~3 s, output in ~/.cache/marola-graph/<repo>/
$ just graph query "where is the swimability score computed?"
NODE Swimability src=marola-app/core/src/main/scala/marola/scoring/Swimability.scala …
A sample of the generated block, from the spike 2 prototype on 2026-10-03:
| Artifact | Published by | Pinned in | Read by |
|---|---|---|---|
ghcr.io/marola-dev/marola-app image |
marola-app/.github/workflows/docker.yml |
marola-ml/marola-image, marola-oods/marola-image, marola-site/marola-image |
marola-ml/scripts/app-image.sh, …, marola-site/scripts/board-schema.sh |
4. Data sources and dependencies reviewed¶
- graphify (
Graphify-Labs/graphify, PyPIgraphifyy, Apache-2.0; the wheel also ships aLICENSE-MITfor code from before the relicensing). It parses code locally with tree-sitter, and sends docs, YAML, PDFs and images to an LLM pass. Upstream releases about daily: 0.9.74 came out on 2026-10-02. Pick: the nixpkgs packagepkgs/by-name/gr/graphify, 0.9.66 both onnixos-unstableand at the devkit's locked nixpkgs (b4fd65b). It includestree-sitter-scalaand hasGRAPHIFY_OUT(the output-dir override, honoured end to end since 0.8.45). Updates arrive through the devkit's normalnix flake update, with no second pinning mechanism. Lagging upstream by a few weeks is a feature for a tool this young. - graphify's MCP server (the
mcpextra) is the same engine as the CLI. Its node lists matched the CLI's for all 8 questions, and it adds about 6.6 KB of tool schemas. Rejected. - graphify's installers, hooks and
--strictwrite assistant config and hooks, and strict mode blocks an agent's first file read. Since 0.9.72, any CLI command also refreshes installed skills. Rejected: they collide with the devkit'score.hooksPathand its plugin. - graphify's LLM pass: about 0.6–0.7M input tokens per full pass, and it would return the wiring as inferred edges. Rejected (§9).
- Claude Code and generated files (code.claude.com docs, 2026-10-03):
.claudeignorehas no effect; the supported mechanism isRead(...)rules inpermissions.deny.- Editing repo files keeps the prompt cache.
- An untracked
graphify-out/still changes the git-status snapshot each session starts from, and invites agents to read 5–10 MB of JSON. - So the output goes outside the checkout.
- The wiring sources are marola's own files:
.github/workflows/*.ymlin every repo,.gitmodules,flake.lock, the pin files (marola-image,corpus.version,resources.version) and whatever scripts, justfiles, Dockerfiles orbuild.sbtname them.
5. Design¶
5.1 The wiring extractor and REPOS.md's generated block¶
scripts/wiring.py in marola-devkit (stdlib plus PyYAML, added to the devkit's python), on PATH
as wiring, with a wiring recipe in devkit.just. It runs in the umbrella, where every submodule
is checked out, and reads the devkit's tree from the flake's store path. It parses:
| Source | Yields |
|---|---|
on: per workflow (repository_dispatch.types, workflow_run, workflow_call, push paths hitting a pin) |
what triggers each workflow |
event_type= / …/dispatches in run:, and uses: marola-dev/<repo>/.github/workflows/X@ref with its event-type |
who sends each dispatch, and each reusable-workflow ref |
gh release upload/create, IMAGE: ghcr.io/… + build-push-action, actions/deploy-pages |
what each workflow publishes or deploys |
pin files, and any script/justfile/Dockerfile/build.sbt naming one or its releases/download URL |
where each artifact is pinned, and who reads the pin |
It writes four tables, between <!-- wiring:start -->/<!-- wiring:end --> markers in
docs/2-Building-marola/REPOS.md:
- artifact → published by → pinned in → read by;
- dispatch → sent by → triggers;
- pin bump → workflow;
- deploy → site.
wiring --check regenerates the block and fails on any difference, the same pattern as
mip_graph.py's block in docs/MIPs/README.md. It also fails on two more conditions:
- a dispatch type that is listened for but never sent, or sent but never listened for;
- an artifact with no reader. An artifact the regex can't resolve is excepted with a one-line
wiring.allow entry giving a reason. Today that is api-docs.tar.gz, fetched through variables
in fetch-api-docs.sh, and ghcr.io/marola-dev/marola-ml, which the maintainer keeps (decided
2026-10-03) although no workflow or script in the org reads it.
Decided 2026-10-03: the wiring lives as a generated block inside REPOS.md. REPOS.md's per-repo
rows (what each repo owns, MIP-0074 §5.1) stay hand-written. The artifact and dispatch facts become
generated, which removes the drift MIP-0074 §8 accepts for them. REPOS.md stays the one copy, as
the maintainer decided for MIP-0074. That amends MIP-0074 task 27: its contracts table of artifacts
and pins is written as an empty wiring:start/wiring:end block, or as a hand-written table that
step 3 below replaces.
5.2 graph in marola-devkit¶
scripts/graph.sh on PATH, with pkgs.graphify added to runtimeDeps, plus a graph *args recipe.
Every repo gets it with its next devkit bump.
graph buildrunsgraphify extract . --code-only --no-labeland thengraphify update .(the local Markdown pass) in place, withGRAPHIFY_OUT=${XDG_CACHE_HOME:-$HOME/.cache}/marola-graph/<repo>. It recordsHEADandgit submodule statusinbuild.jsonbeside the graph. graphify honours.gitignoreby default, so.tmp/wt-*worktrees and.envstay out. An uninitialised submodule is an error, not a silent gap. A repo with noise to drop commits a.graphifyignore: the umbrella drops**/vendor/**(leaflet.js was 81 nodes of noise) and the non-canonicalboard.schema.json.- No model call on this path. Every graphify call runs under
env -i, passing onlyPATH,HOME(the cache dir),GRAPHIFY_OUT,GRAPHIFY_NO_AUTO_REFRESH=1andOLLAMA_BASE_URL=http://127.0.0.1:9. Where unprivileged user namespaces exist, it also runs underunshare -rn, with a notice when it can't. The tool never passes--backend, and refusesextractwithout--code-only. graph query|path|explainpass--graph <cache>/graph.json, andquerydefaults to--budget 400. Spike 2 found--budgetonly truncates: start nodes and order were identical at 2000, 800 and 400, so all hits survive at about 1.8 KB instead of 6.6 KB. They print a staleness line whenHEADor a submodule SHA differs frombuild.json; a rebuild takes about 3 s.- Never: graphify's installers or hooks; committing any of its output; reading
GRAPH_REPORT.md(about 21k tokens) into context.
5.3 The plugin skill¶
plugins/marola-devkit/skills/routing/SKILL.md, a routing rule rather than a tutorial:
- "What produces, consumes, pins or triggers X?" Read the wiring block in
REPOS.md: the umbrella's copy, or docs.marola.dev from a submodule. - "Where is symbol or concept X?" when the keyword is unknown or the repo unfamiliar. Use
just graph query, then read only the files it names. - The keyword is known. Use
git grep(with--recurse-submodulesin the umbrella). - Never read
GRAPH_REPORT.mdorgraph.jsonwhole.
5.4 The umbrella¶
AGENTS.md: one paragraph under "The repos" (MIP-0074 task 27 turns that section into a link toREPOS.md), naming the wiring block,just graphand the skill.quality-otherrunswiring --check, so a PR that changes a workflow or pin fails until the block is regenerated. pointer-sync's PR gets it too: it moves submodules, so it regenerates the block in the same commit.- A
graphjob inci.yml, path-filtered to submodule pointers and the devkit pin. It runsjust graph buildon the PR's pinned commits as a smoke test of each devkit/graphify bump, and uploadsgraph.jsonplusbuild.jsonas an artifact with 7-day retention for a session without Nix. It is never committed and never deployed.
5.5 Rollout¶
- marola-devkit:
wiringwith fixtures andwiring.allowsupport. The reusablenotify-umbrella.ymlsendssubmodule-updatedon every push tomain, and alsosubmodule-docs-updatedwhen the push touchedREADME.mdordocs/**(agit diff --name-onlyover the pushed range, instead of the caller'spaths:filter). Release. - marola-devkit:
graphand the routing skill. Release. - each code repo (app, site, corpus, ml, oods): bump the devkit, and drop
paths:from thenotify-umbrella.ymlcaller so it runs on every push tomain. One small PR per repo. - marola: devkit bump; the wiring block in
REPOS.md;wiring --checkinquality-other; theAGENTS.mdparagraph; the CI job; pointer-sync regenerating the block. The routing re-test (§7) goes in the PR description.
Decided 2026-10-03: every code repo sends submodule-updated on each push to main, as MIP-0070
described. Steps 1 and 3 close the gap the extractor found, so step 4's wiring --check gates with no
allow entry for it. Until step 3 lands everywhere, pointer-sync's daily cron keeps pointers moving, as
it does today.
6. Scoring / safety impact¶
None. No code path in marola-app changes; this is dev tooling.
7. Verification plan¶
wiring.py, pytest with fixture repos:dispatch_types_listened,dispatch_send_step,reusable_workflow_ref_and_event_type;release_upload_assets,image_publish,deploy_pages_site,pin_file_and_reader;check_fails_on_stale_block,check_fails_on_orphan_dispatch,check_fails_on_unread_artifact,allow_entry_needs_reason.graph.sh --self-test, with a fakegraphifyonPATHthat records its argv and environment:env_scrubbed(no*_API_KEY,ANTHROPIC_*orOPENAI_*gets through);no_auto_refresh_set;refuses_extract_without_code_only;output_outside_checkout(git status --porcelain --ignoredis empty after a build);uninitialised_submodule_fails;stale_notice_on_new_head;query_default_budget_400.- Live: after MIP-0074, on a fresh
--recurse-submodulesclone, re-run the 8 questions using the skill's rule and record the result in the umbrella PR. Done means at least 7 of 8 answered right first (spike 2: 5 from the table, 3 from the graph), with 0 tokens in graphify's report footer. - CI:
quality-otherturns red on a test branch that adds an unread pin or an unsent dispatch type.
8. Risks, limitations, and honest caveats¶
- The extractor knows marola's idioms: its pin names,
gh release, dispatch steps. Anything fetched through variables is invisible to it (fetch-api-docs.shtoday), and the INFERRED edges come from regexes over shell text.wiring.allowmakes each blind spot explicit instead of silent. - Accidental LLM spend if someone runs bare
graphifyinstead ofgraph. The skill names onlyjust graph. - Network isolation is best effort.
unshare -rnneeds unprivileged user namespaces, which Ubuntu 24.04 restricts (AGENTS.md's ai-jail note). The guarantee isenv -i,--code-onlyand--no-label; the namespace is defence in depth. - Graph answers are lexical. Stop words pull in noise: "point" ranked
SamplingPointCoordinatesabove the Telegram entry point. When every start node is reached, the budget is ignored and output grew to 13 KB once. - Staleness: the graph answers for the tree it was built from, and the staleness line is the only
defence. The wiring block cannot go stale on
main, because--checkgates it. - Upstream churn: graphify changes daily. The nixpkgs pin and the CI smoke job catch a breaking change at bump time, not in a session.
9. Alternatives considered¶
- Do nothing: MIP-0017's call, which was right for the monorepo. After the split, the cross-repo questions are exactly the ones an agent can no longer grep for.
- Only a hand-written
REPOS.md(MIP-0074 as accepted): it answers the wiring questions, but it drifts, and it missed the unsentsubmodule-updated. - Wiring injected into graphify's graph (prototyped in spike 2):
pathis exact (0.2–0.35 KB) once you know the labels, butqueryrarely reaches the injected nodes. It also has to be re-injected after everyupdate, which evicts nodes outside the scan root, and it ties us to graphify's private eviction rules. - graphify's LLM pass over YAML and docs: about 0.6–0.7M tokens per pass, for edges that come back as inferred, when a ~200-line parser gets them exactly and for free.
- graphify's MCP server: the same answers plus about 6.6 KB of schemas.
- Per-repo graphs merged with
merge-graphs: no repo-to-repo edges, and it drops about 900 doc→code links. - Committing graph output: 5.9 MB that churns on every change.
- A PyPI/uv pin instead of nixpkgs: a second pinning mechanism beside the devkit's flake.
- The wiring as a separate generated file instead of a block in
REPOS.md: a second place for the same facts, against the maintainer's one-copy decision on MIP-0074.
11. Open questions¶
None. The maintainer decided the three questions the spikes raised (2026-10-03):
- The wiring block lives in
REPOS.md, generated (§5.1), which amends MIP-0074 task 27. - Every code repo sends
submodule-updatedon each push tomain(§5.5 steps 1 and 3). ghcr.io/marola-dev/marola-mlstays. It is awiring.allowentry, not a retirement (§5.1).
Appendix¶
Spike 1 report (2026-10-02)¶
Measured on a fresh clone in a scratch directory, not in this repo. Kept verbatim apart from heading levels; "below" and "Proposal" refer to this report, which §5 supersedes.
Verdict: go with conditions for a follow-up MIP. Treat it as an optional, local, on-demand symbol and doc-mention index for in-repo routing, not as the cross-repo routing table: the graph does not contain the cross-repo wiring (workflows, image pins, tarball pins, dispatch events).
Versions and setup¶
- graphify 0.9.73 (PyPI
graphifyy), installed into a scratchpad venv withuv venv+uv pip install. Nouv tool install, nographify install, nographify <platform> install, nohook install. - Umbrella
marola-dev/marolaat8387c25, cloned with--recurse-submodules. Submodules: marola-app16ab0c4, marola-corpusc62001b, marola-ml50d08b3, marola-oodsdad1e0a, marola-site568199a. marola-devkit cloned separately atd1872de. - No-LLM guarantee, enforced rather than trusted. Every graphify call went through a wrapper:
unshare -rn env -i HOME=<scratchpad>/fakehome PATH=<venv>/bin:/usr/bin:/bin OLLAMA_BASE_URL=http://127.0.0.1:9 graphify …That means no network namespace, no API keys in the environment, and a fake HOME, so nothing reaches~/.cacheor~/.graphify. This matters because graphify auto-detects a backend from whatever keys are set (Gemini → Kimi → Claude → OpenAI → … → Ollama), andcluster-only/labelname communities with that backend unless you pass--no-label. The report footer saidToken cost: 0 input · 0 output. The fake HOME was still empty at the end, and the clone'sgit statusshowed only?? graphify-out/: no.gitignore,CLAUDE.mdor.claude/was written.
Commands¶
gfy extract . --code-only --no-label --timing # per repo dir, and once at the umbrella root
gfy update . # incremental; no LLM ("Code graph updated")
gfy query "<question>" # default --budget 2000 tokens
gfy explain "pointer-sync.sh"; gfy path "corpus-fetch.sh" "Corpus"
gfy merge-graphs <7 graph.json> --out merged.json
1. Multi-repo¶
One run at the umbrella root walks all five checked-out submodules into one graph. No flags needed: the 202 code files it found equal the sum of the per-repo runs (16+139+23+4+17+3). It works because the submodule worktrees are plain directories on disk. Nothing in graphify is submodule-aware, so an un-initialised submodule just drops out silently.
| Graph | Files (code) | Nodes / edges, extract --code-only |
Nodes / edges, after update (adds .md) |
Cold build | graph.json |
|---|---|---|---|---|---|
| umbrella only | 16 | 195 / 343 | 1889 / 2058 | 0.7 s | 0.18 MB (code-only) |
| marola-app | 139 | 1655 / 3401 | 1771 / 3683 | 1.5 s | 2.0 MB |
| marola-site | 23 | 504 / 767 | 534 / 793 | 1.2 s | 0.40 MB |
| marola-corpus | 4 | 8 / 12 | 60 / 51 | 0.3 s | 7 KB |
| marola-ml | 17 | 213 / 426 | 270 / 478 | 0.6 s | 0.22 MB |
| marola-oods | 3 | 9 / 14 | 31 / 32 | 0.3 s | 8 KB |
| marola-devkit | 52 | 536 / 1121 | 688 / 1256 | 1.1 s | 0.58 MB |
| workspace (one run) | 202 | 2556 / 4950 | 4530 / 7978 | 2.6 s (3.4 s incl. timing) | 3.0 MB → 5.5 MB |
Whole graphify-out/ after update: 17 MB in total. That is graph.json 5.5 MB, graph.html 4.5 MB,
GRAPH_REPORT.md 84 KB (~21k tokens, too big to load as context), manifest.json 100 KB, plus the AST
cache. Peak RSS is about 105 MB.
Surprise: update is not the same as extract --code-only. update also runs graphify's local,
deterministic Markdown extractor (headings, plus inline code-span mentions resolved to code symbols as
references edges, from markdown_resolution.py). That pass made no LLM calls, which the netns
enforces. It nearly doubles the graph. A CI recipe has to pick one of the two on purpose.
Merging: merge-graphs works (7 graphs → 5209 nodes / 8351 edges in 0.5 s, 6.5 MB). It prefixes
ids with repo:: and adds a repo attribute. It is a disjoint union, though: zero repo-to-repo
edges, only edges into shared external nodes (kyo, java, …). Its cross-repo pass
(cross_repo_calls.py) only binds typed member calls across repos, and marola's repos never call
each other. A merge also loses the ~900 doc→code edges that the single workspace run finds.
Conclusion: graph the umbrella in one run; don't merge per-repo graphs.
2. Cross-repo edges¶
- Code-only workspace graph: 3 repo-to-repo edges, all spurious. They are
rationale_foredges from the umbrella'sscripts/strip_external_scripts.pydocstrings to marola-app's copy of the same file, a dedup artefact. - After
update(Markdown mentions): ~900 repo-to-reporeferencesedges, almost all umbrella docs/MIPs → marola-app Scala symbols (818), plus umbrella → site (57) and a few ml/corpus → app. A 15-edge sample was mostly right (MIP-0055 §5.6 →KnowledgeStore, MIP-0031 →.waterVerdict()) with some noise (a nix-tooling plan →TrailFinderSpec.find()). These are "which doc talks about this code" edges, not wiring. - Missing, and these are exactly the edges a cross-repo agent needs.
.yml/.yamlare in graphify'sDOC_EXTENSIONS, so they get only the LLM semantic pass: all 85 workflow/YAML files are absent from a no-LLM graph. That covers pointer-sync.yml, docs.yml, notify-umbrella.yml, site.yml, release.yml and everyrepository_dispatchtype.marola-image,corpus.version, justfiles, Dockerfiles and.nixare unclassified and skipped (80 files). So there is no edge for: app image → site/ml/oods (marola-imagepins), corpus tarball → app/ml (corpus.version,corpus-fetch.sh), ml-resources produced by apprelease.yml→ consumed by mlresources-fetch.sh, notify-umbrella →repository_dispatch→ pointer-sync/docs. - The board-schema contract is in the graph twice:
marola-app/cli/src/main/resources/board.schema.jsonandmarola-site/site/board.schema.json, 215 nodes each. Nothing links the two copies.
3. Routing quality (workspace graph after update, 4530 nodes)¶
Output cost: each query printed 47–75 lines, 6.4–8.0 KB (~1.6–2.5k tokens), because the default
budget is 2000 tokens and most queries hit the "TRUNCATED: showing 43 of 301 nodes" path. For the
baseline, git grep --recurse-submodules -l -i -E <kw> -- ':!*.md' printed 62–1234 bytes. Caveat in
grep's favour: I picked the keywords knowing the answers, while graphify took the plain-English
question. "Right first" means the right repo + file in the first NODE src= lines.
| # | Question | graphify query | grep baseline |
|---|---|---|---|
| 1 | Where is the swimability score computed? | Hit, 1st: marola-app/.../scoring/Swimability.scala (6.6 KB) |
hit within 15 files (852 B) |
| 2 | What builds the map's boards, which image? | Partial: BoardSpec/Main/SiteBuilder (app) found; the first src is the wrong repo's marola-ml/scripts/app-image.sh; marola-site/.github/workflows/site.yml + marola-image absent (6.4 KB) |
marola-image → site.yml, board-schema.sh, site justfile (297 B) |
| 3 | Where does the RAG corpus come from, how pinned? | Partial: Corpus.scala 1st; the pin (corpus.version, corpus-fetch.sh, corpus release.yml) absent (6.7 KB) |
corpus.version → 10 files across app/corpus/ml (311 B) |
| 4 | What consumes ml-resources? | Hit (consumer): marola-ml/scripts/resources-fetch.sh 1st, full answer in 36 nodes; the producer (app release.yml) absent (8.0 KB) |
4 files, producer + consumer (144 B) |
| 5 | Which workflow moves the submodule pointers? | Miss: AGENTS.md, then MIP-0040/Reviewer noise; pointer-sync.sh is a node but wasn't reached; the workflow isn't in the graph (6.6 KB) |
.github/workflows/pointer-sync.yml 2nd (423 B) |
| 6 | Where is the Telegram bot's entry point? | Partial: Main.scala 5th, after SamplingPointCoordinates (the word "point") and TELEGRAM-SETUP.md (6.5 KB) |
27 files, mostly issue templates (1234 B) |
| 7 | How does a marola-site docs change reach docs.marola.dev? | Miss: marola-site/docs/index.md, .mcp.json, DEV-FLOW.md; notify-umbrella → docs.yml → prepare-docs.sh is absent (6.6 KB) | submodule-docs-updated → docs.yml, pointer-sync.yml (62 B) |
| 8 | Where is the Open-Meteo forecast fetched? | Hit, 2nd: OpenMeteoClient.scala (6.5 KB) |
3 files incl. it (191 B) |
Score: 3 hits, 3 partial, 2 misses. Every hit is a "where is symbol X" question about Scala/Python in
one repo. Every miss or partial is about the artefact/workflow wiring between repos. explain
"pointer-sync.sh" is cheap (523 B) and accurate, but you have to know the node name already, and
it only shows the script's own functions, not who runs it. path "corpus-fetch.sh" "Corpus" and
path "SiteBuilder" "board-schema.sh": "No directed path found." Matching is lexical, so stop words
like "point" or "entry" pull in noise.
4. Freshness and cost of keeping it current¶
- A full rebuild of the whole workspace takes ~2.6 s on a laptop;
updatetakes 3.3–4.5 s (it re-renders the report and graph.html each time). Staleness costs nothing to fix, so incremental support (update,watch,check-update, manifest-based) is a nice-to-have, not a requirement. - The graph records
built_at_commit, the umbrella HEAD only. It doesn't capture submodule SHAs, so after pointer-sync it can't tell whether a submodule's slice is stale. A recipe would have to recordgit submodule statusnext to it. - Regeneration points: local on demand (cheapest, always fresh), or a CI job after pointer-sync.
graphify's own hooks (
hook install, which writes git hooks and a merge driver into git config) break this repo'score.hooksPath= devkit model and are excluded. - If committed: 5.5 MB graph.json that churns on every code change, plus
.graphify_root/cache files that are machine-local. Don't commit it to any working tree. The README's prompt-cache warning says graphify writes into the workspace and every write invalidates Claude Code's prompt cache unless ignored. Its suggested.claudeignoreis unverified: I don't know that Claude Code honours that file. The safe form is an output path outside the checkout, read with--graph <path>. graphify always writesgraphify-out/under the scanned root, so a recipe has to scan from a scratch copy or move the outputs afterwards.
5. The docs semantic pass (not run)¶
- What it adds over the local Markdown pass: LLM-extracted concepts/relations from .md and, crucially, from .yml (the workflows), plus PDFs/images. Of everything graphify offers, only this path could capture the cross-repo wiring, and even then it would come back as LLM-inferred edges, not parsed ones.
- Size: 261 doc-class files (85 of them YAML), ~351k words / 2.6 MB ≈ ~0.6–0.7M input tokens per full
pass, plus JSON output and retries. Truncated chunks are bisected up to 3 levels (up to 8× calls
per chunk) unless
GRAPHIFY_MAX_RETRY_DEPTH=0. Backend: auto-detected from env keys, default Claude modelclaude-sonnet-4-6;--backend claude-clispends the subscription,--backend ollamais free and local but slow.--updatere-extracts only changed docs. - Worth it? Not now. MIP-0074 is restructuring the docs, which would invalidate the pass, and the
inferred workflow edges are what we'd least want to trust. A ~50-line deterministic parser of
.github/workflows/*.yml+ pin files (below) gets the wiring exactly, for free.
Proposal: umbrella vs devkit vs each repo¶
marola-devkit (shared tooling, versioned with the devkit tag):
- graphify pinned (uv lock in a tools/graphify/ dir, or a flake package), never uv tool install.
- A graph recipe/script that runs netns-isolated with env scrubbed (the wrapper above, or at
minimum env -i + --no-label + --backend unset), uses a chosen mode (extract --code-only then
update, i.e. code + local Markdown), writes to ${XDG_CACHE_HOME}/marola-graph/<repo>/ (outside
the checkout), and records git submodule status alongside.
- .graphifyignore defaults: **/vendor/** (leaflet.js is 81 nodes of noise), graphify-out/,
.tmp/, generated board.schema.json copies except one canonical copy.
- A small wiring extractor (the gap graphify doesn't fill): parse workflows for
repository_dispatch types/sends, uses: marola-dev/marola-devkit/...@, and pin files
(marola-image, corpus.version, release asset names) into nodes/edges in graphify's graph.json
schema, merged into the umbrella graph. This is where the cross-repo routing value is.
- The agent skill/snippet: "for 'where is symbol X', try graph query/explain first; for
'what consumes/produces/triggers X', read the wiring table". Shipped as a plugin skill, not via
graphify claude install (hooks; strict mode blocks the first raw Read).
- Later, optionally: a reusable CI workflow that builds the graph and uploads it as an artifact.
umbrella (marola):
- The workspace graph is built in one run at the root (multi-repo works; merge doesn't add edges).
Publish it as a CI artifact after pointer-sync / on docs.yml's schedule, or just build it locally
in under 5 s. No branch, no docs-site page (5.5 MB, low human value).
- AGENTS.md routing guidance: a hand-written, deterministic repo routing table (repo → owns →
artefact produced → consumers → workflow that moves it). This answers Q2–Q7 better than any graph
query today, in about 30 lines.
- The just graph entry point, delegating to the devkit recipe.
each repo: nothing committed. A repo-local just graph (devkit recipe) for single-repo sessions is
fine, but the per-repo graphs are small enough that grep wins there too (marola-corpus 8 nodes,
marola-oods 9).
Would NOT do: graphify install / claude install / hook install / --strict; commit
graph.json; run the semantic pass in CI with an API key; rely on merge-graphs for cross-repo
structure; load GRAPH_REPORT.md (~21k tokens) into agent context.
Risks¶
- Accidental LLM spend: any env with an API key + a bare
graphify extract/cluster-only(no--code-only/--no-label) calls a model. The devkit wrapper has to scrub env and kill the network. - Prompt-cache invalidation if outputs land in the working tree.
- Fast-moving upstream (0.9.73, a YC startup steering toward a hosted platform): CLI semantics have already drifted between versions (README notes ghost-duplicate fixes, a Terraform id regen), so pin it.
- Query output is ~2k tokens per question, versus 0.1–1 KB for a good grep. It only beats grep when the agent doesn't know the keyword.
- Lexical matching produces confident-looking wrong first hits (Q2 wrong repo, Q6 "point").
Open questions¶
- Does Claude Code honour
.claudeignore, or must outputs live outside the repo? (I assumed outside.) - Does graphify's MCP server (
[mcp]extra) give better routing than the CLI query? Not tested. - Would
--budget 800keep hits while cutting cost? Not measured. - Is a wiring extractor emitting graphify-schema JSON maintainable in devkit, or should the routing
table just be generated Markdown (
just routing-table) and skip graphify entirely?
Spike 2 report (2026-10-03)¶
Measured on fresh clones in a scratch directory, not in this repo. Kept verbatim apart from heading levels; §5 supersedes its proposal.
Verdict: use graphify only for "where is symbol X", at --budget 400. Generate the cross-repo wiring
as a Markdown table from a small deterministic extractor, not as nodes injected into graphify. The
extractor works and its path answers are exact, but query almost never surfaces the injected nodes,
and update drops some of them.
Setup, the same as spike 1. Umbrella 273bd11 (app 0b2f982, corpus c62001b, ml 50d08b3, oods
dad1e0a, site e25d12b) and devkit 141541b, freshly cloned into the scratchpad. Every graphify call
ran under unshare -rn env -i HOME=<scratch>/fakehome PATH=<venv>/bin:/usr/bin:/bin OLLAMA_BASE_URL=http://127.0.0.1:9
GRAPHIFY_NO_AUTO_REFRESH=1 [GRAPHIFY_OUT=<abs>] graphify …, with --code-only --no-label on extract.
At the end the fake HOME was empty, and both clones had no untracked or ignored files.
1. Upstream drift¶
- Version: 0.9.74, published to PyPI 2026-10-02 16:21 UTC, two days after 0.9.73. Release cadence is about daily (0.9.67 through 0.9.74 in about a week).
- Licence. Still Apache-2.0 (
License-Expression, GitHubspdx_id). The wheel now also shipsLICENSE-MITand aNOTICE: code contributed before the relicensing stays available under MIT. No problem for us. - YAML: unchanged.
.yml/.yamlare still inDOC_EXTENSIONS, so they get only the LLM doc pass. The one exception is package manifests (apm.ymletc.), which go through the code path. Workflows still never enter a no-LLM graph. - Output dir: it exists, and spike 1 missed it. The
GRAPHIFY_OUTenv var takes an absolute path (paths.py, #686). WithGRAPHIFY_OUT=<scratch>/out-ws,extractandupdatewrote everything there, and the clone stayed clean (git status --porcelain --ignoredwas empty). No scratch copy or move step is needed.query/path/explain --graph <abs>read the graph from there. - Multi-repo/submodules: nothing new. The only mention of submodules is gitlink handling in
detect.py.merge-graphsandcross_repo_callsare unchanged. - New safety item. Since 0.9.72, any non-install CLI command auto-refreshes stale installed skills.
It writes
SKILL.mdand makes.bakcopies.GRAPHIFY_NO_AUTO_REFRESH=1turns this off, and it now belongs in the devkit wrapper. It was a no-op here because the fake HOME has no skills. - Did the spike-1 numbers move? Two things changed at once: the repos and the version. I measured both, on the same commit:
| Workspace (umbrella + 5 submodules, 209 code files) | 0.9.73 | 0.9.74 | spike 1 (0.9.73, old commit, 202 files) |
|---|---|---|---|
extract --code-only nodes / edges |
2604 / 5021 | 2640 / 5099 | 2556 / 4950 |
update nodes / edges |
4885 / 8354 | 4921 / 8447 | 4530 / 7978 |
| time, extract / update | – | 2.7 s / 3.7 s (RSS 105 / 173 MB) | 2.6 s / 3.3–4.5 s |
| graph.json after update | – | 5.9 MB (graph.html 4.9 MB, report 92 KB) | 5.5 MB |
The version change accounts for +36 nodes / +78 edges (Scala abstract defs, from changelog #3962).
The rest is repo growth. Timings are flat.
2. Query budget (CLI query, workspace graph, 0.9.74)¶
The default budget is 2000. --budget only truncates: the start nodes and their order were identical
at all three budgets for all 8 questions. So the hit/partial/miss result cannot change, and it matches
spike 1.
| # | Question | Result (any budget) | bytes @2000 | @800 | @400 |
|---|---|---|---|---|---|
| 1 | swimability score | hit, Swimability.scala 1st | 6654 | 3034 | 1816 |
| 2 | map boards + image | partial (ml app-image.sh 1st, BoardSpec/Main; no site.yml/pin) |
6569 | 2928 | 1797 |
| 3 | RAG corpus + pin | partial (Corpus.scala 1st; no pin/release) | 6714 | 3080 | 1943 |
| 4 | consumes ml-resources | hit, resources-fetch.sh 1st | 8110 | 3028 | 1784 |
| 5 | workflow moving pointers | miss (break-ui SKILL.md, MIP-0040, AGENTS.md) | 6679 | 3059 | 1809 |
| 6 | Telegram entry point | partial (Main 5th, after SamplingPointCoordinates "Entry") | 6614 | 3037 | 1851 |
| 7 | site docs → docs.marola.dev | miss (site docs/index.md, MIP-0074, .mcp.json) | 6767 | 3091 | 1891 |
| 8 | Open-Meteo fetch | hit, OpenMeteoClient 2nd | 6689 | 2960 | 1770 |
At --budget 400 all 3 hits survive in the first 2 NODE lines, at about 1.8 KB per answer: ~27% of
the default output (8–15 nodes instead of 36–58). The cost is that the "TRUNCATED: answer may be among
the N cut nodes" banner is always present. Q4 at 400 still names fetch()/release() in
resources-fetch.sh. One trap: if every start node is reached, the budget is ignored ("complete answer
over budget"). Q4 on the wired graph printed 13 KB at --budget 2000.
3. MCP server ([mcp] extra)¶
- Start it with
python -m graphify.serve <graph.json>. Transports arestdio(the default) andhttp(Streamable HTTP, optional--api-key). Tools: query_graph(question, mode=bfs|dfs, depth=3, token_budget=2000, context_filter[])get_node(label|node_id)get_neighbors(label, relation_filter, token_budget)get_community(community_id, token_budget)god_nodes(top_n, exclude_hubs_percentile)graph_stats()shortest_path(source, target, max_hops=8, undirected)list_prs,get_pr_impact(pr_number),triage_prs. These three need GitHub, so they are useless offline.
It also exposes 6 resources (graphify://report, stats, god-nodes, surprises, audit, questions).
- It runs offline. I drove it under the netns wrapper with a 60-line stdio JSON-RPC client:
initialize, tools/list, then query_graph × 8 questions × 3 budgets, shortest_path, get_node.
The server identified as graphify 0.9.74.
- No better than the CLI. It is the same engine. The NODE lists were identical to the CLI's for all
8 questions at 800, and the bytes were within ±3. Its default depth=3 (CLI: 2) finds more nodes, but
they are then truncated to the same output. On top of that, the tool schemas are 6.6 KB (~1.6k tokens)
of definitions. Claude Code's tool search defers MCP tools by default, which softens this.
shortest_path corpus-fetch.sh → Corpus: "No path found", the same as spike 1.
4. Wiring extractor prototype¶
wiring_extract.py: 203 lines (176 non-blank, non-comment), stdlib + PyYAML, 0.17 s. Its inputs:
- .gitmodules and flake.lock, for the devkit ref.
- Each repo's workflows: on.repository_dispatch.types, workflow_run, workflow_call input
defaults, push/PR paths that hit a pin file, event_type=/dispatches in run:, and
uses: marola-dev/<repo>/.github/workflows/X@ref with its with: event-type (or the reusable
workflow's default). Also gh release upload|create asset names, IMAGE: ghcr.io/… +
build-push-action, actions/deploy-pages (domain from the CNAME echo or mkdocs site_url), and
the scripts/* that a step runs.
- The pin files marola-image, corpus.version, resources.version, plus any script, justfile,
Dockerfile or build.sbt that names a pin, and its releases/download URL.
Edges run along the flow (producer → artifact → pin → consumer), so path can follow them. They
reuse graphify's own file-node ids, so the wiring joins the code graph. The schema came from
validate.py: file_type must be one of the 6 valid types, confidence one of
EXTRACTED/INFERRED/AMBIGUOUS. Every injected edge is tagged context: "wiring".
Output: 83 nodes, 179 edges. That is 53 workflows, 6 pins, 4 assets, 2 images, 3 events, 2 sites, 7 repos, and 6 files graphify skips (justfile, Dockerfile, build.sbt). Confidence: - EXTRACTED: contains 53, uses_reusable 41, triggers 10, submodule 5, pinned_in 3, deploys 2. - INFERRED, i.e. regexed from shell text: runs 32, read_by 12, dispatches 8, publishes 6, downloaded_by 3, pinned_in 3.
Merged into graph.json: 5004 nodes.
query after injection (Q2, Q3, Q5, Q7): start-node selection is lexical over labels, and the
existing doc/code nodes outrank the new ones.
- Q2, Q3: still partial, 0 wiring nodes in the answer.
- Q5: still a miss. The pointer-sync.yml workflow node never matches "workflow … pointers".
- Q7: miss → partial. The 1st node is now docs.marola.dev (marola GitHub Pages site) with
src=.github/workflows/docs.yml, but the notify-umbrella → dispatch half of the chain is not in the
output.
- Bytes are unchanged (±100).
path after injection is exact and cheap (197–344 B), once you know a label:
release.yml workflow (marola-corpus) --publishes--> marola-corpus-<tag>.tar.gz --downloaded_by--> marola-app/scripts/corpus-fetch.sh
notify-umbrella.yml (marola-site) --dispatches--> submodule-docs-updated --triggers--> docs.yml --deploys--> docs.marola.dev
docker.yml (marola-app) --publishes--> ghcr.io/marola-dev/marola-app --pinned_in--> marola-image (marola-site) --triggers--> site.yml
release.yml (marola-app) --publishes--> ml-resources-<tag>.tar.gz --downloaded_by--> resources-fetch.sh
The first of these printed "target match was ambiguous" (two corpus-fetch.sh, app and ml). path
"…tar.gz" "Corpus" (the Scala class) fails: nothing links the fetched tarball to Corpus.scala,
because the build-time handoff runs through Dockerfile/build.sbt resources.
Does update keep injected nodes? Partly. Injected nodes whose source_file exists under the scan
root survive (_origin: "semantic"). Anything with a path outside it is evicted as deleted: all 14
devkit nodes and 55 edges, including all 41 uses_reusable edges. Before I fixed the file-id join,
injected nodes on re-extracted code files were evicted too. The rule is to re-run the extractor after
every update. It is idempotent (it replaces everything tagged context=wiring), but it ties us to
graphify's private eviction rules (watch.py, _origin), which have changed in recent releases.
The Markdown form (--md, generated from the same scan) is 2,555 bytes (~650 tokens) for the whole
workspace. It has four tables:
- artifact → published by → pinned in → pin read by
- dispatch event → sent by → triggers
- pin bump → workflow
- deploy workflow → site
A sample row: | ghcr.io/marola-dev/marola-app container image | marola-app/.github/workflows/docker.yml |
marola-ml/marola-image, marola-oods/marola-image, marola-site/marola-image | marola-ml/scripts/app-image.sh, …,
marola-site/scripts/board-schema.sh |.
It answers Q2, Q3, Q4, Q5 (pointer-sync.yml as the receiver of submodule-*-updated) and Q7. That is
5/8 for 2.5 KB loaded once, against graphify query's 3/8 at ~1.8 KB per question. Q1, Q6 and Q8 are
symbol questions, out of its scope.
| Form | Answers | Bytes | Keeping it correct |
|---|---|---|---|
graphify query @400 |
3/8 (symbol questions) | ~1.8 KB per question | rebuild ~3 s; lexical noise |
graph + wiring, query |
3 hits + Q7 partial | ~1.8 KB per question | extractor + re-inject after each update; depends on graphify internals |
graph + wiring, path |
Q2/3/4/7 exact | 0.2–0.35 KB, but needs labels | same |
| generated routing.md | 5/8 (all the wiring) | 2.5 KB once | extractor only; a diffable table that a CI check can regenerate and compare |
Findings the extractor surfaced that grep and graphify both missed:
- Nothing sends submodule-updated. pointer-sync.yml listens for it, and MIP-0070 says every code
repo dispatches it on each push to main. Only submodule-docs-updated is ever sent.
- Some assets have no consumer:
- app api-docs.tar.gz: the umbrella's fetch-api-docs.sh downloads through $repo/$ASSET_NAME
variables, which the regex can't resolve. This is a known extractor miss.
- ml api-docs.tar.gz: the same miss.
- ghcr.io/marola-dev/marola-ml: no consumer found.
- Reusable-workflow refs are skewed: the umbrella is at @v0.2.4, every submodule at @v0.2.2.
5. Claude Code and ignore files¶
.claudeignoredoes not exist. The permissions doc says verbatim: "If your project has a.claudeignorefile, it has no effect, so move its entries intoReaddeny rules" (code.claude.com/docs/en/permissions, §Read and Edit). The supported mechanism ispermissions.deny: ["Read(./graphify-out/**)"], gitignore-style patterns, applied best-effort to Grep/Glob/@-mentions.respectGitignoreonly affects the@file picker (settings-reference).- Prompt cache. "Editing files in your repository" is listed under actions that keep the cache. File contents enter context only when read, and a changed file that was already read only gets an appended system-reminder (code.claude.com/docs/en/prompt-caching). So graphify's README warning overstates the effect for Claude Code. Writing graph files into the tree does not invalidate the cached prefix by itself.
The real costs are different:
- an untracked graphify-out/ changes the git-status snapshot at session start. That snapshot is
part of the prefix, so sequential sessions stop sharing a cache.
- agents may Read or Grep 5–10 MB of JSON/HTML.
- git status noise.
Writing outside the checkout through GRAPHIFY_OUT=<abs> is not the only safe option, but it is
the simplest, and a Read deny rule would still be needed for the in-tree form.
Not checked¶
- Whether better node labels (workflow descriptions from the header comments) would get wiring nodes
ranked by
query. That would be tuning to the question set. get_neighbors/context_filter=['wiring']as a routing call.- Running the extractor in CI.
- Windows/macOS.
- The 0.9.74 HTTP transport.
- The private-name scan from spike 1. Its list lives under the main checkout's
.tmp/, which this spike was told not to touch.
Checked live (2026-10-03)¶
- https://pypi.org/pypi/graphifyy/json (0.9.74, uploaded 2026-10-02T16:21:07)
- https://github.com/Graphify-Labs/graphify (
gh api: licence Apache-2.0, releases v0.9.74 2026-10-02 and v0.9.73 2026-09-30, CHANGELOG.md 0.9.69–0.9.74) - https://code.claude.com/docs/en/permissions (the
.claudeignoresentence, Read deny rules) - https://code.claude.com/docs/en/settings-reference (
respectGitignore) - https://code.claude.com/docs/en/prompt-caching (actions that keep or invalidate the cache, cache scope)
- https://code.claude.com/docs/en/settings (a
permissions.denyRead(./.env)example)
Hygiene: the clones, both venvs (0.9.74 and the 0.9.73 comparison), the fake HOME and all outputs were
deleted. What remains is this report and wiring_extract.py. Nothing was written outside the scratchpad.
Checked live¶
gh api repos/NixOS/nixpkgs/contents/pkgs/by-name/gr/graphify/package.nix, 2026-10-03, onnixos-unstableand at the devkit's lockedb4fd65b:version = "0.9.66",buildPythonApplication,tree-sitter-scalaamong the dependencies, optionalmcp = [ mcp starlette ], licenceasl20 mit. At the umbrella's own lock (0968519): 0.9.53. Commit history: 0.9.48 → 0.9.53 (09-02) → 0.9.61 (09-13) → 0.9.66 (09-23).Graphify-Labs/graphifyCHANGELOG.md(master) andgraphify/paths.pyatv0.9.66, 2026-10-03:GRAPHIFY_OUThonoured end to end since 0.8.45 (#1423) and present at 0.9.66; skill auto-refresh andGRAPHIFY_NO_AUTO_REFRESHsince 0.9.72.notify-umbrella.ymlonmainof marola-app, -site, -corpus, -ml, -oods, 2026-10-03: each calls the devkit's reusable workflow@v0.2.4on pushes touchingREADME.md/docs/**, with noevent-type, so the defaultsubmodule-docs-updatedis sent. The umbrella'spointer-sync.yml:9listens for[submodule-updated, submodule-docs-updated], and its comment says the former is not sent today. Spike 2 saw the submodules at@v0.2.2: that was the umbrella's pinned commits, notmain.marola-dev/marola-devkitat141541b(shallow clone, 2026-10-03):flake.nix'stoolsmap,runtimeDeps,consumerTools;devkit.just;plugins/marola-devkit/{skills,hooks}.docs/MIPs/MIP-0017-agentic-tooling-survey.md:87,169(the earlier rejection),MIP-0074-…md:90,275,331, andMIP-0074.tasks.mdrow 27 (marola-dev/marola#617, open on 2026-10-03).- Everything in the two spike reports was measured in those sessions; their own "Checked live" lists apply.
Not checked¶
- Whether GitHub-hosted
ubuntu-latestrunners allowunshare -rn. - nixpkgs' 0.9.66 build itself: both spikes ran PyPI wheels (0.9.73, 0.9.74). The 0.9.73 → 0.9.74 difference was +36 nodes, so 0.9.66 is expected to be close, not verified.
- That graphify honours
.gitignoreduringextractby default: from its README (--no-gitignoredisables it). Both spikes ran on fresh clones, where nothing was ignored. - Any of the new code:
wiring.pyexists as a 203-line prototype in spike 2's scratch area, not in any repo.graph.sh, the skill and the CI job are designs only.