Skip to content

The prompt compile (DSPy)

Python, run offline only: this never runs in production and marola's Scala/Kyo runtime never imports Python. It exists to produce two artifacts, both loaded and replayed at request time by marola.llm.CompiledPrompt/marola.llm.Reviewer via a plain OpenAI-compatible chat-completions call, with no Python in the runtime path. Both live in marola-app's core/src/main/resources/:

  • recommendation_prompt.json: the summarizer (turns a BestHour into a sentence).
  • review_prompt.json: the reviewer/critic pass that grades the summarizer's own output and can replace it (marola.llm.Reviewer).

See the system architecture §5a (query synthesis) for why this exists and compile_recommendation_prompt.py for the actual program (both dspy.Signatures, both trainsets, both compile calls); this page is just the "how to run it" instructions.

DSPy is Python-only ("Declarative Self-improving Python"; no JVM port exists; see FUTURE-WORK §10 for the Scala-ecosystem gap this leaves and the proposed ds4s port), and its optimizer is a compile-time step, not a runtime dependency, so it doesn't need to run in the deployed service.

Why this needs a real LLM, and can cost a little money

DSPy's optimizers (BootstrapFewShot, MIPROv2, ...) work by actually calling a language model repeatedly: running the draft program against training examples, checking outputs against a metric, and keeping/refining what worked. There's no way to "compile" a prompt without a real model in the loop. Against a paid endpoint, budget a few cents to a few dollars depending on the optimizer and trainset size (BootstrapFewShot with 3 examples, as configured here, is cheap; MIPROv2 costs more). Per AGENTS.md's cost-safety rule, a run against a paid endpoint needs a human's go-ahead first. It's a deliberate, small, one-time spend, not something to run repeatedly or in CI.

Setup

cd dspy
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Running it

Point MAROLA_DSPY_MODEL (a LiteLLM model string, since that's what DSPy uses under the hood) and the matching provider credentials at whichever model marola will actually run at request time, so the optimized prompt matches the model that'll replay it. Unset, it is ollama_chat/smollm2:360m; MAROLA_DSPY_API_BASE points it at an Ollama server (http://localhost:11434). A paid model, for example:

export MAROLA_DSPY_MODEL=openai/gpt-4o-mini
export OPENAI_API_KEY=sk-...
python compile_recommendation_prompt.py

just compile-prompt runs the same script from the repo root. This writes both recommendation_prompt.json and review_prompt.json in one run, to .tmp/compiled/ (--out overrides it). Re-run it whenever TRAINSET/REVIEW_TRAINSET grow (these double as hand-labeled eval sets; see FUTURE-WORK §4.1 for the gap between "doubles as an eval set" and an actual held-out dspy.Evaluate loop, which doesn't exist yet) or the target model changes. There's no watch mode, it's a manual step. The compile prompt workflow (compile-prompt.yml, dispatch only, a local Ollama model, so free) compiles and opens a PR that commits both into the app's resources directory above (MIP-0070 §5.4).

Optional: tracing the compile run with Langfuse

BootstrapFewShot/MIPROv2 call the model many times per run (once per training example per bootstrap attempt, more for MIPROv2's Bayesian search) to figure out which few-shot demos and instructions actually work; that process is otherwise a black box. Setting these three env vars turns on Langfuse tracing for the whole run, via the OTEL-based Python SDK v3 and Langfuse's own DSPy integration:

export MAROLA_LANGFUSE_PUBLIC_KEY=pk-lf-...
export MAROLA_LANGFUSE_SECRET_KEY=sk-lf-...
export MAROLA_LANGFUSE_BASE_URL=https://cloud.langfuse.com   # or your self-hosted instance
python compile_recommendation_prompt.py

Get a free key pair at cloud.langfuse.com (generous free tier) or self-host (LANGFUSE_BASE_URL → http://localhost:3000 per Langfuse's own docs). Omit these three vars entirely to skip tracing. _init_langfuse_tracing() checks for MAROLA_LANGFUSE_PUBLIC_KEY first and no-ops silently if it's unset, and degrades to a warning (never a crash) if the credentials are set but unreachable/invalid, so a broken Langfuse setup never blocks the actual compile step.

Uses the MAROLA_LANGFUSE_* prefix (marola's env var convention; see marola-app's AppConfig.scala and .env.example) rather than Langfuse's own bare LANGFUSE_* names directly, so _init_langfuse_tracing() bridges one to the other internally.

Optional: logging compile runs to MLflow

Additive to the Langfuse tracing above: both can run at the same time (MIP-0010 §11 OQ4). Where Langfuse traces every individual LLM call made while compiling, this logs one MLflow run per dspy.teleprompt.Teleprompter.compile() call (two per script invocation: "summarize" and "review") to the marola/prompt-compile experiment: params (model, optimizer, trainset_size), the compiled program's own metric score (a fresh dspy.Evaluate() pass over its trainset, using the same metric it was compiled against), and the artifact JSON it wrote (recommendation_prompt.json or review_prompt.json). A local server is one just mlflow-up away in a marola-app checkout.

export MAROLA_MLFLOW_TRACKING_URI=http://127.0.0.1:5000
export MAROLA_MLFLOW_EXPERIMENT=marola/prompt-compile      # optional, this is the default
python compile_recommendation_prompt.py

Omit MAROLA_MLFLOW_TRACKING_URI entirely to skip this: no mlflow import, no extra LLM calls for the evaluation pass, no network, same degrade-silently shape as the Langfuse hook (a stopped mlflow server prints a warning and continues rather than failing the compile step).

mlflow.dspy.autolog() (MIP-0010 §11 OQ3) exists at the pinned versions but isn't used. Confirmed live against a real mlflow==3.16.0 + dspy==3.3.1 install: autolog(log_compiles=True) patches Teleprompter.compile to open its own run and log the optimizer's hyperparameters plus a best_model.json/trainset.json artifact pair, but it never computes an aggregate metric score, and its artifact names don't match the actual recommendation_prompt.json/review_prompt.json files this step needs on record. Narrower than what this script needs on both counts, so it hand-logs instead; see _log_compile_run_to_mlflow()'s docstring in compile_recommendation_prompt.py for the full reasoning.

Run python compile_recommendation_prompt.py --self-test to check the params/metrics dict this logging builds, offline: no LLM call, no MLflow server, no mlflow import (only the "unconfigured" path is exercised; see the function's own docstring).

Status

Both compile steps have actually been run against a real LLM, end to end, more than once, a local Ollama model, not a paid API, so this didn't cost anything: ollama_chat/dolphin-mixtral:8x7b (26GB, MAROLA_DSPY_API_BASE=http://localhost:11434) and, separately while writing RUN-LOCALLY, the much smaller llama3.2:1b (1.3GB). Both produced real compiled artifacts with genuine LLM-bootstrapped demos (each demo carries "augmented": true; inspect the JSON directly to see this). Along the way this also surfaced and fixed a real, unrelated environment issue: tokenizers' Rust extension needs libstdc++.so.6, which a Nix-based Python environment doesn't put on the default linker path:

export LD_LIBRARY_PATH="$(ls -d /nix/store/*-gcc-*-lib 2>/dev/null | sort -V | tail -1)/lib:$LD_LIBRARY_PATH"

(picks the newest gcc-*-lib derivation on your store if more than one is present; verified this resolves correctly on the machine this was written on, which had both a 15.2.0 and 15.3.0 copy) needed once per shell before running python compile_recommendation_prompt.py, or DSPy fails with ImportError: libstdc++.so.6: cannot open shared object file.

The Scala-side loader (marola.llm.CompiledPrompt, marola.llm.Reviewer) exists and was run live against both compiled artifacts (just run -- --summarize in a marola-app checkout); see the system architecture §5a's own Status notes for the full detail, including one case where the reviewer correctly caught a deliberately-planted flaw in a draft summary.

langfuse==4.15.1 + openinference-instrumentation-dspy are confirmed importable and the graceful-degradation path (credentials set, endpoint unreachable) was exercised directly, but the actual Langfuse happy path (a trace landing in a real project) remains unverified: no Langfuse account was set up in this environment.