← Writing

Most Teams Reach for Fine-Tuning Too Early

In this article 4 sections

When a general-purpose LLM doesn’t quite do what you need, whether that’s the raw model or an agent built on top of it, the reflex, especially for people who’ve built classic ML before, is to fine-tune it: train it on your data, make it yours. That instinct comes from a world where the model was the system and retraining was the main lever you had. With today’s LLMs it usually isn’t where you should start. There are three ways to close the gap between a general model and your specific problem: give it better grounding context, optimize the harness around it, or fine-tune the weights. I would identify which gap I actually have before committing to training.

The distinction becomes clearer across three needs: domain knowledge, reliable workflow behavior and deployment constraints. A pharmaceutical recipe project put the grounding choice in front of me. My later agent workflow tested the harness argument. Pushing that kind of workload onto an edge device is the counterexample I would test next, because it may change which intervention deserves priority.

When the missing piece is domain knowledge

I saw the grounding question most clearly on a pharmaceutical project: a recipe-design assistant inside a pharma MES. A recipe engineer takes an existing manufacturing recipe and authors a new one, specifying asset integration, step execution and the required checks. The result has to use the right terminology and follow company and process guidelines. An answer the engineer cannot trace is hard to accept.

The prevailing view was that a general model could not be trusted here, so it needed segment-specific training. What I kept coming back to was the information it needed at authoring time: the recipe library, definitions and governing guidelines. I wanted to test how much of the gap those sources could close before committing to changes in the weights.

That choice also creates different maintenance work. Retrieval needs current, usable sources. Fine-tuning needs suitable training data, evaluation and a decision about when to retrain. I do not have a matched cost comparison for this engagement, so I would not put a universal price or cost ordering on those paths.

The project landed on context management with retrieval-augmented generation (RAG), supplying relevant source material when the model answered. Looking back, I would also have put more effort into the harness: applying the guidelines and assembling a valid recipe step by step. The observed choice was grounding; the additional harness work is what I would test now.

When the model can do the work but skips the process

A second MES-style job is an agent that reads tables, files and notes, extracts an insight or flags a possible anomaly, then gives a person the evidence to decide. Before training a model for that job, I would check how it uses the capability it already has. Does it read the whole table, follow the tool result and hand back a claim the person can inspect?

The harness is the prompts, tool descriptions and surrounding code that guide and constrain those steps. Improving it may close a workflow gap without retraining. That still needs an evaluation on the actual task; it is not a claim that one general agent replaces every specialized anomaly detector.

I later tested that argument against a narrower failure: an agent that kept skipping steps in a repeatable workflow. The model already had the capability to do each step. Moving sequence, verification and state into a deterministic harness enforced the declared step order. Putting that harness into recurring work exposed the next boundary: deterministic order alone did not make interruption and restart safe.

LangChain’s July 2026 Nemotron 3 Ultra study provides a separate, reported example. Keeping the model fixed, the team raised its typical Deep Agents score from about 0.80 to 0.84. Its best run reached 0.86, close to Opus 4.8’s reported best of 0.87. The roughly tenfold cost comparison applied to those best runs; the reported advantage varied with precision and caching.

One correction moved a keep-reading instruction from the tool description into the returned file content. The model had been answering from the first page. That is relevant to the file-reading failure I care about, but it does not establish how the same change would perform in an MES deployment. The NVIDIA walkthrough describes the implementation.

When the deployment target changes the choice

Now put that workload on an inline factory cell with a fixed GPU budget and a hard response deadline. A large model with extensive retrieved context may not fit. I would first measure whether a smaller existing model, a tighter prompt and shorter tool loops meet the requirement. Fine-tuning or distillation becomes a candidate when that simpler path leaves a demonstrated gap.

This is the counterexample to a rigid grounding-first order. The deployment constraint may make model size the first problem to solve. I have not run that edge comparison here, and the cost depends on the chosen model, retrieval, workload volume and required quality.

Choose against the gap, then measure the cost

The three approaches can be combined. I would compare the work each adds and the failure it is meant to fix, then measure both implementation effort and the complete cost of an accepted result.

Scroll sideways for more columns.

ApproachFirst questionWork it addsWhat I would measure
Grounding contextIs the required information missing or out of date?Retrieval, source maintenance and usable contextCorrectness, traceability, retrieval latency and prompt cost
Harness optimizationCan the model do the task but follow the process unreliably?Tool contracts, workflow checks and evaluationCompleted requirements, retries, tool calls and review effort
Fine-tuning / distillationDoes a tested capability or deployment gap remain?Training data, fitting, evaluation and model maintenanceHeld-out quality, runtime resources and the cost of future changes

My next comparison is an open model on one defined task: establish a baseline, test grounding and harness changes, then train only if a gap remains. The result I want is the quality and cost of an accepted output, including corrections and review, rather than a cheaper model call that leaves more work for the person using it.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All writing