Fine-tuning is justified only after prompt and RAG baselines leave a repeatable gap, and every training run remains versioned, held out, and comparable.

A stable behavior gap has to earn the training experiment. Prompt and RAG baselines get the same held-out tasks first. If the gap remains material, each versioned run is weighed for lift, critical failure, general regression, consistency, cost, and latency.

The release decision weighs target lift against regression, consistency, cost, and latency, and the experiment trail behind it can be rerun by someone else.

Illustration of LLM Fine-Tuning & Model Customization: a team tuning a model against a target behavior and baseline

Some of the 500+ brands we've worked with

See all references
  • Acıbadem Sağlık Grubu
  • Atasun Optik
  • Tazedirekt
  • Sporjinal
  • QNB Finansfaktoring
  • Vispera

The reason for training, the intervention itself, and the result stay connected from the first baseline comparison through the release recommendation.

  1. Prove there is a training problem

    We test the target behavior with prompt and RAG baselines, record the gap that remains, and agree the held-out tasks, critical behaviors, regression limits, budget, and decision owner.

    AI assist
    Prompt and RAG baseline results feed a model-drafted gap comparison for the product owner to challenge.
    Human gate
    Is a tuning experiment more appropriate than the simpler alternatives? Your product owner confirms the remaining gap justifies a tuning experiment.
  2. Keep every run identifiable

    Data and configuration references are fixed for each experiment. We track every run and artifact, including assumptions and changes introduced after the previous result.

    AI assist
    The model applies the agreed data-version and configuration tags to each run record, followed by a specialist's metadata check.
    Human gate
    Can every result be tied to a specific data version, configuration, and model artifact? A specialist verifies the run metadata before results are compared.
  3. Read the lift beside the loss

    Target-task improvement is compared with critical-behavior failures, general-capability regression, consistency, inference cost, and latency, split by the cases that matter.

    AI assist
    The model prepares a slice-by-slice table of target lift, regressions, and cost for specialists to verify.
    Human gate
    Do the gains survive the regression, cost, and latency checks by important slice? Your product owner decides whether the trade-off justifies the tuned model.
  4. Put the release decision on record

    The handover includes the chosen artifact, experiment trail, held-out comparison, limitations, release conditions, and the event that should trigger another review.

    AI assist
    A model-drafted release summary captures the accepted run and its documented limits for the product owner to review.
    Human gate
    Does your owner accept the evidence and the limits of the selected run? Your product owner accepts the tuned artifact and its release conditions.

Another specialist should be able to reproduce the run, compare it with the alternatives, and challenge the recommendation without taking the best examples on trust.

  • Decision record

    Prompt, RAG, and tuning baseline comparison

    The target behavior, prompt and RAG baseline, remaining gap, tuning rationale, acceptance checks, exclusions, and owner.

  • Dataset

    Versioned training and evaluation manifest list

    Model, data, configuration, environment, run, evaluation-set, and tuned-artifact references for each experiment.

  • Dashboard

    Reproducible run ledger and selected-model reference

    The run history, changes, failures, observations, selected artifact, and links to held-out results.

  • Model card

    Held-out comparison report and limitations summary

    Baseline and tuned results by task and critical slice, plus regressions, consistency, cost, latency, intended use, and known limits.

This engagement fits a known, repeatable behavior gap that simpler prompt and retrieval changes have already failed to close against the same baseline.

A good fit when

  • Prompt and RAG baselines have run on the target tasks, but the same output behavior still falls short across held-out cases.
  • Curated training data is ready, yet the separate evaluation set must still expose any critical behavior that quietly regresses.
  • The experiment budget and latency needs are known, but nobody has compared the expected lift with inference cost and refresh demands.
  • Prompt, RAG, and tuning options have separate results, but no baseline record shows why training is the justified intervention.
  • Another specialist cannot reproduce the training runs because their data versions and configuration changes are not attached to each artifact.
  • Held-out tasks show target lift, but critical slices, general regression, consistency, cost, and latency are not read beside that gain.
  • A tuned artifact can be selected, yet its limitations, release conditions, owner, and retraining trigger sit in four separate places.

Better handled as other work when

  • You want fine-tuning before testing a prompt or retrieval change on the same cases. The simpler baseline must establish the remaining gap first.
  • Evaluation examples entered training or model selection, yet their result is being counted as lift. That leakage needs a clean comparison.
  • You need production deployment or continuous retraining. Those duties require a separately scoped engagement.

If one of these is closer to your situation, start here instead: See model customization services

This is the part of Zeo that writes and ships code. Our senior engineers build agents, chatbots, and RAG pipelines, along with the automation and data work around them, and they keep operating those systems once they're live. We've worked with more than 500 brands since 2011.

  • Modal

    provisions on-demand GPU compute for each training run without a dedicated cluster

  • Weights & Biases

    tracks each versioned run against the prompt and RAG baselines it must beat

  • Braintrust

    scores each trained version's critical failure cases, not just its average lift

  • Hugging Face

    hosts the base model and training pipeline the fine-tuning run executes on

  • Unsloth

    runs each versioned training experiment at a fraction of the usual compute cost

  • Predibase

    serves multiple fine-tuned versions side by side for direct comparison

Have the baseline, curated dataset, held-out cases, and experiment budget ready, and know who will make the release call.
Discuss fine-tuning

An agreed behavior target, prompt and RAG baseline, curated permitted data, a separate evaluation set, critical behaviors, candidate-model constraints, experiment budget, inference requirements, and regression limits. Data and evaluation versions are frozen before comparison.