Integrations¶
Each pluggable backend: the trait in core, the local (free) implementation, how AppConfig wires
it, and what has been verified. A cloud backend would be one more implementation of the same trait,
opt-in per integration (GCP is the path under discussion, MIP-0057). The endpoints and their terms
are in Data sources, the variables in the
Configuration reference.
| Integration | Trait (core) |
Local implementation | AppConfig |
Status |
|---|---|---|---|---|
| Query synthesis | LlmClient |
LocalLlmClient (Ollama) |
llmClient, tracedLlmClient |
Run live |
| Agentic tool access | — | SwimConditionsMcpServer (stdio) |
reads fromEnv per call |
Run live over raw JSON-RPC |
| Chat server | — | ChatServer (JDK HTTP server) |
knowledgeStore, llmClient |
Tested offline (ChatServerSpec) |
| Sighting reports | SightingStore |
LocalFileSightingStore (JSON lines) |
sightingStore |
Run live |
| Photo analysis | VisionClient |
LocalVisionClient (multimodal Ollama) |
visionClient |
Request path run live; no description yet |
| Observability | Tracing, RunLedger |
MlflowTracing, MlflowRunLedger |
tracing, runLedger |
Tested offline |
| Bathing-water quality | WaterQualityClient |
IMA/SC, INEA/RJ, INEMA/BA clients | waterQualityClient(origin) |
Run live |
| Facilities | AccessibilityClient |
OverpassAccessibilityClient, NoopAccessibilityClient |
accessibilityClient |
Run live |
| Ocean knowledge | Embedder, KnowledgeStore |
OllamaEmbedder, FileKnowledgeStore |
knowledgeStore |
Run live; default embedder broken on current Ollama (#12) |
Callers take the trait, never the implementation class: that is what keeps core free of any
backend reference, and the review style guide flags the opposite.
Query synthesis¶
The ranked list is useful without a model: every number comes from data and a deterministic rule. The LLM's job is narrower, turning the top row into a sentence or two, and the ranking stays outside it on purpose.
LlmClientis the trait (completeover chat messages).LocalLlmClientcalls any OpenAI-compatible/chat/completions: Ollama by default, or LM Studio or a llama.cpp server.CompiledPromptloads a DSPy-compiled artifact,recommendation_prompt.json, and replays it as messages: the instructions as the system message, each demo as a user/assistant pair, the real input last. It mirrors DSPy'sChatAdapterin good faith, not byte for byte. The JSON shape is the real output ofdspy.Predict(...).save()(dspy 3.3.1);outputFieldnames the output,summaryhere.Reviewerreplays a second artifact,review_prompt.json(outputFieldreview_json), over the draft and the same facts. It checks the jellyfish and whale mention policy and that the draft asserts nothing outside the facts, and returns a 0–100 score, a verdict (approve/revise) and afinal_summary, its own correction onrevise. The output is one JSON string rather than three fields so it reuses the single-output replay; the first{…}in the reply is parsed, because small models wrap JSON in prose.
Both artifacts are compiled offline by marola-ml's
dspy/compile_recommendation_prompt.py
(BootstrapFewShot, with a metric that rewards mentioning jellyfish risk, and whale likelihood less,
when Moderate or High) and land here as a bot PR. --summarize always runs both passes.
Verified end to end against local Ollama models: the compile with dolphin-mixtral:8x7b
(its demos carry "augmented": true), and --summarize replaying the result. The reviewer returned
well-formed JSON from both that model and llama3.2:1b, and in one bootstrap run caught a planted
draft missing its jellyfish mention. Small models follow instructions imperfectly: one mentioned a
whale at Low, and the 1B reviewer's corrections fixated on whales over jellyfish. The daily
docker-smoke.yml runs --summarize with llama3.2:1b and fails when the reviewer does not
answer (Development).
Agentic tool access¶
SwimConditionsMcpServer
exposes the pipeline as four MCP tools (find_nearby_beaches, get_swim_recommendation,
get_water_quality, ask_ocean_question; arguments and results in the
CLI reference), so an agent decides when to call what instead of
Recommender fixing the order. It speaks stdio (StdioServerTransportProvider), which needs no
network exposure; .mcp.json registers it for Claude Code as just mcp-server. A remote client
would need the SDK's servlet transports and a public endpoint, which is a deployment and not wired.
The SDK's tool handlers are plain synchronous Java callbacks, so the server runs each Kyo effect
with Sync.Unsafe.evalOrThrow under AllowUnsafe.embrace.danger, Kyo's documented escape hatch
for a foreign-callback boundary
(Effects map).
Each call reads AppConfig.fromEnv afresh.
Two build traps it left behind: the assembly concatenates META-INF/services/* before discarding
the rest of META-INF, or the SDK's ServiceLoader finds no JSON-schema validator; and
Compile / run / mainClass is pinned to marola.Main, so the server runs as
cli/runMain marola.agent.SwimConditionsMcpServer.
Verified by piping raw JSON-RPC (initialize, notifications/initialized, tools/list,
tools/call) into the assembled jar, with live Overpass and Open-Meteo answers, when it had its
first two tools. A session from a real MCP client is not recorded.
Chat server¶
ChatServer (--serve-chat, MIP-0033)
wraps OceanQa.answer and the safety footer behind /health and /ask, so the map's chat widget,
reached through a tunnel to the machine running it, gets a grounded answer with the footer instead
of talking to Ollama directly. Its endpoints are in the
CLI reference; running and exposing it is
Chat and MCP.
Sighting reports¶
SightingStore (record,
recentFor) with
LocalFileSightingStore,
one JSON object per line in data/sightings.jsonl. It is the collection half of the heuristics'
calibration loop (Heuristics). A user would
report through the Telegram bot, which is Phase 1 and not built; --report-sighting is the local
stand-in.
Verified: --report-sighting jellyfish Arpoador "note" wrote a correctly shaped line, read
back from the file.
Photo analysis¶
VisionClient (describe(imageBytes))
with LocalVisionClient: a
multimodal Ollama model (llava, moondream) over the same /chat/completions endpoint, the image
as a base64 image_url content part. Photos would also arrive through the bot; --analyze-photo
<path> is the stand-in.
Verified up to the model: the encoding, the request shape, the round trip, and Ollama's
model 'llava' not found surfacing through Abort instead of a crash. No multimodal model was
installed, so no description has been checked.
Observability¶
Infra tracing and one span per LLM call behind a vendor-free trait (MIP-0010).
Tracing has withSpan and llmSpan;
Tracing.Noop is the default. MAROLA_TRACES=mlflow picks
MlflowTracing in
AppConfig.tracing, which Main resolves once per run; a server that is down prints a warning and
traces nothing, so observability never fails a recommendation. A recommendation is one trace:
marola.recommend, with bestPerBeachTomorrow, trails and the draft and review llm.<model>
spans under it. --ask is traced too.
MlflowTracingexports OTLP/HTTP to<MAROLA_MLFLOW_TRACKING_URI>/v1/traces, with thex-mlflow-experiment-idheader MLflow requires; the id of<prefix>/tracesis resolved over REST byMlflowApi. Each span is exported as it ends (SimpleSpanProcessor): a short-lived CLI has no moment to flush a batch. Nesting is anAtomicReferenceto the current span rather than OpenTelemetry's thread-local context, which a Kyo effect may not stay on: exact for the CLI's one linear pipeline, wrong for concurrent ones. A failing effect ends its span withERRORand rethrows.TracedLlmClientwraps the LLM client (AppConfig.tracedLlmClient):gen_ai.operation.name,gen_ai.request.model, message count and prompt and completion sizes in characters. No token counts:LlmClient.completedrops the response'susage. The prompt and completion text are recorded only withMAROLA_TRACE_CONTENT, because the prompt carries the swimmer's coordinates.RunLedgeris the experiment-tracking half:MlflowRunLedgerwhen the tracking URI is set,RunLedger.Noop(no call) otherwise.--benchmarklogs to it throughBenchmarkLedger.
Tested offline against OpenTelemetry's in-memory exporter (MlflowTracingSpec: names,
attributes, nesting, error status, endpoint and header) and scripted MLflow responses
(MlflowRunLedgerSpec). A trace into a live just mlflow-up server is not recorded; the run to
check it is in Development.
Bathing-water quality¶
Each Brazilian state's agency publishes its own samples, so the provider follows the origin
(MIP-0001, MIP-0031). AppConfig.waterQualityClient(origin) picks IMA/SC, INEMA/BA or INEA/RJ by
bounding box (or as MAROLA_WATER_QUALITY_PROVIDER forces) and wraps it: for IMA/SC first in
FallbackWaterQualityClient (the feed, then the bulletin PDF when the feed gives nothing), then for
every agency in CachedWaterQualityClient (the last good fetch, under data/water-cache/).
ImaScWaterQualityClient: one emptyPOSTto the JSON feed the portal's map uses, about 260 points with coordinates and their last five samples, parsed tolerantly.ImaScPdfWaterQualityClient: the newest weekly bulletin, found from the portal's index rather than pinned.ImaScPdfParserreads its dates and verdicts, which carry no coordinates, so rows are joined by beach and point name to the HTTP feed's point list. It covers the feed being unreachable while the bulletin is up.IneaRjWaterQualityClientandInemaBaWaterQualityClient: bulletin PDFs only, read byIneaPdfParserandInemaPdfParser, placed with hand-curated coordinate tables (SamplingPointCoordinates). INEA's PDFs are found on its city pages; INEMA's URL pins one campaign.PdfLineskeeps each line's position: read in natural order these tables pair a verdict with the wrong beach, which for a safety verdict is a wrong answer, not a formatting bug.CachedWaterQualityClientwrites only a non-empty fetch. Serving the cache is not serving stale data: freshness is judged on the agency's sample dates, which do not move because a download failed.
Recommender makes one call per run and matches the points to beaches; the matching and the
verdict are pure and on Heuristics. The tide turns
(Tides) and sea lore (SeaLore) that MIP-0001 added beside it are there too.
Verified live from Campeche on 2026-09-05: point 73 (Riozinho) IMPRÓPRIA at 749
enterococci/100 mL, the other four PRÓPRIA, Campeche −20 with the spot named. Each parser has a
real bulletin as its fixture, and PipelineGoldenSpec replays a recorded IMA response
(Development). The feed is undocumented and samples go stale
off-season (MIP-0001 §8).
Facilities and trails¶
Two more Overpass queries per run, on the same public instance as BeachFinder:
OverpassAccessibilityClient
counts parking, toilets, showers and lifeguards within 300 m of each beach (MIP-0021;
MAROLA_FACILITIES=off swaps in NoopAccessibilityClient, so the call site never special-cases
"off"), and TrailFinder finds named
paths and tracks near the beaches, with OSM's sac_scale and surface copied verbatim or left
empty, never guessed (MIP-0030). Both follow BeachFinder's single-source shape, not the water
clients' one per region: OSM's coverage is global.
Ocean knowledge: retrieval¶
Grounded Q&A over marola-corpus's Markdown documents
at the release corpus.version pins (.tmp/knowledge locally, /app/knowledge in the image).
How the corpus reaches this repo and marola-ml is the umbrella's
Architecture; this is the in-app half.
Corpusreads every.mddirectly under the directory and under itssafety/, and merges paragraphs into chunks of at most 700 characters, each keeping its document's title andSource:URL.OllamaEmbedderembeds through Ollama's native/api/embed, not the/v1the chat uses. Its default,llama3.2, is a chat model, which current Ollama refuses to embed with (#12); setMAROLA_LOCAL_EMBED_MODEL(Choosing the embedder).FileKnowledgeStorekeeps the vectors as JSON underdata/, re-embeds when a fingerprint of the corpus files (names, sizes, mtimes) and the model changes, and ranks by cosine.OceanQaasks the model to answer only from the top four passages, citing[n], or to replyNO_ANSWER_IN_PASSAGES. With nothing relevant (no passage atMAROLA_ASK_MIN_SCOREor above, or that reply),strictabstains, never calling the model when no passage qualifies, andgeneralanswers from the model's own knowledge under an "unsourced" label.--askdefaults togeneral; the chat server and MCP usestrict.SafetyFooterappends the emergency footer to any answer drawn from asafety/document (MIP-0022).
OceanBenchmark (--benchmark, just
benchmark) puts 22 questions, ten inside the corpus and twelve outside, through three arms on the
same model: the plain prompt, strict and general. Its scores are deterministic (keyword coverage,
citation, abstention, latency) and its report ends with a computed verdict. marola-ml's gate keeps
the runs; the
2026-09-05 baseline
had general beating the plain prompt 0.84 to 0.75 overall and 0.92 to 0.55 inside the corpus, after
the NO_ANSWER_IN_PASSAGES stage was added, with llama3.2 embeddings. Fine-tuning targets format
and tone, not facts, and lives in marola-ml's
fine-tune.