Skip to content

Design

marola-ml does three jobs, and none of them is on the request path. When the app answers a question it runs no Python: it replays two JSON prompts and calls a model Ollama serves. This repo produces those prompts and that model, and decides whether a model is good enough to ship.

Job Code Produces Runs when
The prompt compile dspy/compile_recommendation_prompt.py recommendation_prompt.json, review_prompt.json, as a PR to marola-app by hand: just compile-prompt or a compile-prompt.yml dispatch
The fine-tune finetune/ Tier 1: marola-llama3.2, an Ollama Modelfile, no weights changed. Tier 2: marola-sea, a QLoRA + DPO model published to Hugging Face Tier 1: just finetune-model, and every Dockerfile.local build; Tier 2: by hand, a marola-sea-publish.yml dispatch
The benchmark gate scripts/benchmark.sh, scripts/benchmark_gate.py, docs/benchmarks/ the decision to move ghcr.io/marola-dev/marola-ml:local docker-local.yml, on a push to main touching the model, the pins or the gate

Kroki

Where each job's details live

  • The prompt compile: the two signatures, the cost of an LLM in the loop, Langfuse and MLflow.
  • The fine-tune: the tiers, the preset ladder, preflight, merge, publish and the licences that follow a base model.
  • The benchmark gate: what it compares, the embedder gap, re-baselining.

Patterns

  • Pinned releases, never another repo's tree. The app image (marola-image, a tag and its digest), the app's resources tarball (resources.version) and the corpus (corpus.version). scripts/corpus-fetch.sh and scripts/resources-fetch.sh unpack the last two into .tmp/knowledge and .tmp/resources. The benchmark is the pinned app image's --benchmark; nothing here builds the app.
  • The standard library by default. Torch, transformers, peft and trl sit behind lazy imports in finetune/, so every script's --self-test, the dataset builders and the gate run on a bare Python 3.12. DSPy is the compile's own dependency (dspy/requirements.txt).
  • A self-test per script. Each one has --self-test, and just quality runs them all. A dataset builder or the gate reads its input after fetching it and fails on an empty one.
  • One run directory per base model. Training writes under finetune/out/<preset>/, and each script refuses an output or an adapter that belongs to another base, before importing torch.
  • Irreversible steps are a human's. A Hugging Face upload, from just finetune-publish or a dispatched marola-sea-publish.yml run (which also tags marola-sea-v<N>), is started by a person (who may run what).

Module map

Path What it is
dspy/compile_recommendation_prompt.py the two DSPy signatures, their trainsets and metrics, the compile
finetune/Modelfile, finetune/Modelfile.adapter Tier 1's persona on llama3.2; Tier 2's adapter on its base
finetune/build_dataset.py, finetune/build_dpo_dataset.py the SFT set and the DPO pairs, from the pinned resources and corpus
finetune/preflight.py VRAM, RAM, disk and a rough ETA for a preset on this machine
finetune/train_lora.py, finetune/train_dpo.py QLoRA SFT, then DPO on top of it; train_lora.py holds PRESETS
finetune/merge_export.py, finetune/publish_hf.py merge into the base, GGUF export and quantization; the Hugging Face upload
scripts/benchmark.sh, scripts/benchmark_gate.py, scripts/app-image.sh run the pinned app's --benchmark, gate the report, read the pin
scripts/corpus-fetch.sh, scripts/resources-fetch.sh unpack the pinned releases into .tmp/
scripts/analyze_training.py, scripts/marola-sea-pull.sh a training run's report; pull the published model into Ollama
scripts/api-docs.sh pdoc of finetune/ and scripts/
Dockerfile.local Ollama with marola-llama3.2 and nomic-embed-text:v1.5 in its store
docs/benchmarks/ the kept benchmark runs the gate compares against