Training And Post-Training
Training methods are capability-shaping paths that update model weights. They differ from runtime methods such as CoT, ReAct, and Plan-and-Solve: runtime methods organize one run, while training changes what the model tends to do before extra scaffolding is added.
The cleaner storyline is not “training methods” and “reinforcement learning” as two peer topics. It is:
- Pretraining gives the model language, code, factual associations, and latent skills.
- Post-training turns the base model into something better at following instructions, using tools, matching preferences, and learning from feedback.
Post-training is not the same thing as RL. Supervised fine-tuning, instruction tuning, preference optimization, RLHF, DPO, process rewards, verifiable rewards, and agentic RL all happen after pretraining. They differ in their learning signal: demonstrations, preference comparisons, or environment rewards.
A harness places the model inside a runnable, observable, governable system; training and post-training decide which behaviors the model already brings into that system.
Pretraining: Base Capability
Pretraining usually trains a model with next-token prediction over large corpora: given previous tokens, predict the next one. The objective is simple, but it forces the model to compress language structure, code patterns, factual associations, style, common sense, and some reasoning ability. The resulting base model is a general-purpose generator.
The key variable is not just parameter count. Classic scaling laws track model size, training tokens, and compute. Early intuition leaned toward making models larger; Chinchilla-style results emphasized that, under a fixed compute budget, model size and training tokens should scale together. A smaller model trained on many more tokens can outperform a larger undertrained model. In practice, pretraining quality comes from the joint choice of model scale, data scale, data quality, and compute allocation.
Data determines the shape of the base model. Different mixtures of web pages, books, code, papers, Q&A, forums, math text, and synthetic data produce different defaults. Deduplication, filtering, quality scoring, and contamination control affect factual memory, code ability, long-text coherence, and benchmark trustworthiness. Tokenization, context length, architecture details, and training stability also matter, but the main point here is simple: pretraining is not just “feed the model more text.” It is a tradeoff between data, compute, and model capacity.
Still, a base model is not automatically a reliable instruction-following agent. It is good at continuing distributions; it has not necessarily learned to act within task constraints, output protocols, and safety boundaries. GPT-3 showed that scale can produce few-shot behavior: many new tasks can be performed from examples in the prompt without weight updates. But that is not the same as stable product behavior. For harness design, pretraining defines the substrate: what the model can infer, express, and generalize. The harness can expose and constrain latent capability, but it cannot easily supply a capability the model fundamentally lacks.
Post-Training: Demonstrations
Supervised fine-tuning (SFT) / instruction tuning turns the model toward a target behavior distribution. SFT and instruction tuning should not be treated as two fixed sequential stages: instruction tuning is often the instruction-response form of SFT. It teaches the model to understand user intent, respect system constraints, produce expected formats, and move from text continuation toward task completion.
For agents, the shape of the demonstration data matters. Single-turn examples mainly teach the model how to answer; trajectory data teaches it how to work:
- how to decompose a task;
- when to reason and when to call a tool;
- how to update a plan after an observation;
- how to retry or hand back control after failure;
- how to deliver a verifiable result.
This is the first clear meeting point between training and harness design: high-quality agent training examples often come from harness traces, not static Q&A data.
Post-Training: Tool Use
Tool use can be exposed only at runtime through schemas and constraints, or it can be trained into the model so it learns when to call tools, which tool to choose, how to fill arguments, and how to use the result.
- Toolformer uses a small number of API demonstrations to build self-supervised data, training the model to decide when to call an API, what arguments to pass, and how to incorporate the returned result.
- Gorilla fine-tunes on large API-call data and combines the model with retrieval so it can adapt to test-time documentation changes, with a focus on reducing hallucinated API names, arguments, and usage.
- ToolLLM / ToolBench connect real APIs, tool-use instructions, call-path annotation, and evaluation to train and evaluate more complex single-tool and multi-tool behavior.
Tool ability is not a single property of the model. It is jointly determined by training, tool design, and runtime infrastructure:
- training data determines whether the model has seen patterns for selecting tools, filling arguments, reading observations, and continuing reasoning;
- tool schemas determine whether the model can understand boundaries, parameter meanings, and failure modes;
- runtime selection determines whether the tool menu is too large, dynamically disclosed, or permissioned;
- verification and feedback determine whether wrong calls are detected and can become learning signals.
Training a model to call tools is therefore not enough. If tool descriptions are vague, permissions are too broad, observations are not fed back, or failures are not verified, tool behavior can still break down inside a production harness.
Post-Training: Preferences
SFT learns from demonstrated behavior: given an input, imitate a high-quality output or trajectory. Preference optimization learns tradeoffs between behaviors: which candidate answer better matches human preference, helpfulness, honesty, and safety.
RLHF became a standard pipeline in the InstructGPT line:
- use human demonstrations for SFT to produce an initial policy;
- collect human rankings or comparisons over model outputs and train a reward model;
- optimize the policy against that reward model, often with PPO and a KL penalty to keep the model close to the SFT policy.
The central lesson from InstructGPT is that alignment training can make a much smaller model preferred to a much larger base model. But the paper also notes that aligned models still make simple mistakes. For agents, that distinction matters: RLHF optimizes answer behavior under human preference, not real environment completion. A response can be polite, compliant, and helpful-looking while still failing to submit the form, fix the bug, pass the tests, or satisfy permission requirements.
DPO rewrites the RLHF preference objective as a classification loss over preference pairs. It removes the need for a separate reward model and online PPO loop during fine-tuning. It is lighter and more stable in engineering practice, but it is still preference learning: if the data only compares which answer looks better, the model may become better at sounding right rather than completing the task.
Post-Training: Process And Verifiable Rewards
For multi-step reasoning and agent tasks, final outcomes are often too coarse. Reward signals can appear at several levels:
- outcome supervision rewards only the final answer or task result;
- process supervision gives feedback on intermediate steps;
- verifiable rewards (RLVR) use executable correctness checks when they exist: tests, compilers, answer checkers, formal verifiers, or environment success states.
Let’s Verify Step by Step shows that process supervision can outperform outcome-only supervision on mathematical reasoning; the paper released PRM800K, a dataset of 800,000 step-level human feedback labels. DeepSeek-R1 pushed verifiable rewards into the center of the conversation: on verifiable tasks such as mathematics, coding competitions, and STEM problems, RL can elicit long reasoning, self-verification, and strategy shifts.
Extending this idea from single answers to multi-step agent trajectories gives agentic RL: train tool use, recovery, and long-horizon decision-making from task success, environment feedback, and verifiers.
The key word is environment. Single-answer RLVR needs an answer checker. Agentic RL needs task environments that are executable, resettable, observable, and scoreable:
- Execution provides sandboxes, dependencies, reset semantics, and side-effect boundaries;
- Lifecycle provides multi-step control flow and recovery;
- Observability captures traces, cost, retries, and runtime signals;
- Verification turns trajectories into rewards, scores, and attribution;
- Governance limits the action space so reward-seeking does not become authority escalation.
Agentic RL is therefore not only a model-training concern. It pulls ETCLOVG infrastructure into the training loop.
Data Quality Shapes Defaults
Training does not just add skills; it shapes defaults: when to be concise, when to reason explicitly, when to ask for clarification, when to refuse, when to call tools, whether to guess after tool failure, and whether to acknowledge uncertainty.
This is especially sensitive for agents. Low-quality data can teach bad defaults:
- brittle formatting: small schema or prompt changes produce unparsable output;
- tool overuse: the model calls tools even when it can answer internally, increasing cost and risk;
- tool underuse: the model answers from memory when external facts are required;
- overconfidence: the model wraps failed verification as success;
- trace shortcuts: the model omits intermediate observations and handoff state, making later takeover difficult;
- reward gaming: the model learns to exploit the evaluator or tests rather than satisfy the real goal.
That is why production harnesses should not fine-tune only on final good answers. The more valuable data includes traces, evaluation outcomes, and failure attribution: not just what succeeded, but why.
Relation To Harnesses
Training and post-training set the model’s default capability boundary, which changes what the harness needs to supply. As models become better at long-context maintenance, tool use, self-checking, and recovery, some explicit planning prompts, repeated verifiers, or reset scaffolds can become unnecessary overhead.
But training does not replace the harness. Even when a model is trained to use tools and rewards well, the harness remains responsible for:
- execution boundaries and sandboxes;
- tool permissions and identity;
- observation logging and replay;
- visibility into cost, failures, and retries;
- audit and governance of final behavior.
The better framing is a moving boundary: training internalizes some behaviors into the model, while the harness places those behaviors inside a controllable, measurable, accountable runtime system.
Sources
- Training language models to follow instructions with human feedback
- Scaling Laws for Neural Language Models
- Language Models are Few-Shot Learners
- Training Compute-Optimal Large Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Gorilla: Large Language Model Connected with Massive APIs
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Let’s Verify Step by Step
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning