Skip to main content
Engineering ••11 min read•

Agentic Reinforcement Learning in the Harness You Ship

Agentic reinforcement learning belongs in the harness you ship, not a training clone. Learn the five failure modes and when to skip RL entirely.

The agent harness, with its routing, retries, and tool schemas around the model, defines the behavior that reinforcement learning actually optimizes.

Most agent teams that reach for RL quietly train the wrong agent. They rebuild a lightweight stand-in inside the training framework, tune until evals move, then ship the original harness around the new weights and wonder why production lags the benchmark. The deployed policy is the checkpoint composed with everything wrapped around it: routing, retries, context trimming, tool schemas. Agentic reinforcement learning has mostly trained the weights and hoped the wrapper came along for the ride. Agent Lightning v1.0, a roughly 3,500-line open-source framework from Microsoft Research Asia, is built on the opposite bet: run training against the same harness you deploy. The bet is correct, and it is expensive. What follows is the engineering ledger that release notes skip, five failure modes with named mitigations, and a rubric for when in-harness RL beats supervised fine-tuning (SFT), prompting, or a bought checkpoint.

Stay in the loop.

Get the latest posts and exclusive content delivered to your inbox.

Join 10 readers. No spam. Unsubscribe in one click, anytime.

What Agent Lightning v1.0 Changes for Agentic Reinforcement Learning

Skip the release-note tour; two facts matter. First, the multi-library bridge is original Agent Lightning, not a v1.0 invention: a deliberately small integration surface connecting existing RL training libraries such as verl, TRL, SLIME, and AReaL to existing agent frameworks such as LangGraph, Microsoft Agent Framework, and Autogen. You point your harness's model endpoint at an OpenAI-compatible proxy, and the trainer observes model calls while the harness keeps owning its interaction loop. Second, v1.0 hardens that design into Harnessed Agentic RL: the customized trainer now rides on verl, the proxy-based integration surface carries forward, and whichever harness runs in deployment is the one that participates in training. The pitch is simple to state and hard to do: train agents in the same harness you deploy.

The reported evidence that it is now practical rather than a research curiosity: an end-to-end coding pipeline on Qwen3.5-9B with mini-SWE-agent lifted SWE-bench Verified Pass@1 from 41.8% to 56.4% using about 6,000 training samples. Respectable, but the figure that matters even more than 56.4 is 3,500, the approximate line count of the entire control plane. A control plane you can read in an afternoon changes what you should demand from any agent RL stack: an integration surface small enough to audit. The Agent Lightning repository ships v1.0; the Agent Lightning paper documents the original bridge design. Everything below is what you owe back once you adopt the approach.

The Policy You Train Is the Harness-Weight Pair

Classical RL for LLM agents assumes the training framework owns the loop: the model emits an action, the framework appends the observation, and one rollout is one continuous token trajectory. Real harnesses outgrew that assumption years ago. When your harness trims context near token limits, retries a failed tool call, routes between models, or spawns a subagent, it is executing policy. Retry policy is policy. Context trimming is policy. Tool schema ordering is policy. The trajectory distribution you sample from belongs to the harness-weight pair, not to the weights.

So a reimplemented training-time agent is a different policy even with identical weights, and RL optimizes whatever distribution it actually samples. If the training clone's retry logic, context windows, or tool wrappers differ even modestly from production, the gradients push toward a different optimum, and measured gains can fail to transfer at deploy time. This is ordinary distribution shift, hidden inside your own training pipeline. A recent survey of agentic reinforcement learning treats environments and tool use as first-class design problems for exactly this reason. And if you already accept that scaffolding decides deployment quality, the training implication is immediate: the scaffold is part of what you should be training.

How Harness-in-the-Loop RL Actually Runs

One concrete setup, end to end. The agent: a support-triage service built as a LangGraph graph that classifies a ticket, looks up the order, drafts a refund, and calls the ticketing API. The trainer: verl running Group Relative Policy Optimization (GRPO). Agent Lightning sits between them with three components.

  • The API gateway is an OpenAI-compatible proxy. The harness runs unchanged, as a local process or a Kubernetes job, believing it is calling a normal model endpoint. The gateway tags each model call with the rollout it belongs to and logs the inputs, completions, and per-token log probabilities the trainer consumes.
  • The rollout controller starts and manages agent execution, keeping agent CPU, memory, and dependencies off the trainer's GPUs.
  • The customized trainer, built on verl, creates rollouts, waits, collects samples, and assembles training batches. Because a harness splits one rollout into a variable number of samples (retokenization shifts token boundaries, subagents and summarization fragment the trajectory), v1.0 computes advantages and loss normalization at rollout level rather than sample level; the announcement reports this kept policy entropy more stable during training.

What the RL backend owns is deliberately narrow: policy updates, KL control, weight sync. What it never owns: tool execution, retries, or tool state. That boundary is the entire value proposition, and it only works because modern harnesses treat the model endpoint as a pluggable component, which is exactly the seam the proxy occupies. The Microsoft Agent Framework docs describe that same runtime layer from the harness side: tools, state, and orchestration live outside the model call.

Five Things That Break When Real Tools Enter the Training Loop

Reward hacking emerges when an agent manipulates tool state, such as closing a ticket without diagnosis, to inflate its measured reward.

Training against the real harness imports production's failure modes directly into the training loop. These five failure modes, each with a mitigation, are the price of fidelity. Budget them before you commit GPU hours.

Rollout latency moves into the tool path

With real tools in the loop, episode wall-clock time is often dominated by tool latency rather than model forward passes. Plausible arithmetic for the triage agent: 12 model calls at 2 seconds each is 24 seconds of GPU-relevant work; 14 tool calls (orders, ticketing, knowledge base) averaging 5 seconds, plus a retry cycle or two, is 70 to 80 seconds. The model is roughly a quarter of the episode, and during the other three quarters the trainer's GPUs sit idle unless rollouts overlap. Agent Lightning's collocated async mode, which pauses the gateway and drains in-flight requests around each update, is reported at about 2x end-to-end speedup over synchronous RL. That helps; it does not repeal the arithmetic. Mitigation: measure your per-tool latency distribution first. If p50 tool time already exceeds total model time, fix the environment (cache, replay, co-locate) before spending on GPUs.

Seeding does not buy reproducibility

Pinning the sampler seed pins the model's tokens, not the loop. A 503 from the ticketing API in run A but not run B sends two identically seeded runs down different branches; a cache that warms mid-run shifts observation timing; retry and checkpoint layers mean the same logical step can execute a different number of times, semantics the LangGraph checkpointing docs cover. Environment-driven nondeterminism has been documented in deep RL for years: identical seeds, divergent results, whenever the environment has a moving part. Agent RL reproducibility inherits all of it, with more moving parts than most. Mitigation: treat every eval as a distribution, not a number. Run N rollouts, report a band, alert on band shifts instead of point deltas, and record every tool request and response. You will need the tapes when a run diverges.

Reward hacking migrates into tool state

Computing rewards from real tool state (ticket actually closed, refund actually issued) beats grading transcripts, but it changes the attack surface: the agent can now act on the very state your reward reads. Worked example: your reward reads ticket.status == resolved. The policy finds an API route that sets status directly without doing the diagnosis, measured reward climbs, actual work stops. This is reward hacking through tool outputs, a pattern Lilian Weng's overview catalogs in detail, and attacks on the reward computation itself are not hypothetical: Anthropic's reward-tampering investigation documents agents subverting their own reward functions. Mitigation: separate the reward channel from the agent's action surface. Read state through an account the agent cannot write, verify outcomes out of band, and log every write the agent makes to reward-relevant objects.

Throughput ceilings of a small framework

Back-of-envelope: 64 concurrent Kubernetes agent jobs at about 90 seconds per episode is roughly 2,500 episodes per hour before scheduling overhead. A recipe wanting 8,000 rollouts per update waits about three hours per step on that fleet, and a deliberately small control plane was never meant to be a planet-scale orchestrator. The 3,500-line choice is a real feature, auditability and fast integration, that trades away built-in multi-node rollout throughput; teams with heavy async rollout fleets may need to wrap their own orchestration around it. Mitigation: start with collocated async on a single GPU pool; if you become throughput-bound, shard training or schedule rollouts with your own orchestrator instead of growing the framework.

Eval fidelity flips from asset to liability

The point of in-harness training is measuring what you ship. Evaluate on a harness-free, closed-book benchmark afterward and you reintroduce the exact mismatch you paid to remove, so benchmark motion will mislead you about deployment motion. Mitigation: hold out tasks, not just prompts. Evaluate in the same harness, same tools (sandboxed), same retry configuration, and pin a frozen harness snapshot per training run so numbers stay comparable across experiments.

Harness-in-the-Loop RL Versus SFT, Prompting, and a Bought Checkpoint

Choosing among agentic reinforcement learning, SFT, prompting, and a bought checkpoint reduces to one question: where does the deficit live? Orchestration points to in-harness training, reasoning depth to a bought checkpoint, instruction following to prompts. Find the row that matches your diagnosis.

OptionWins whenDominant failure modeCost position
Harness-in-the-loop RLYour scaffold is thick (multi-tool, routing, retries) and task success is checkable from tool stateReward hacking through tool surfaces; idle trainer GPUs during tool waitsHighest: GPU fleet, environment engineering, evals
SFT on tracesYou have or can buy strong trajectories and want stable gains quicklyTrace imitation typically memorizes surface form: trained on one harness's transcripts, a mere tool-call reorder or new retry path can leave the imitated sequence unreachable, and the ceiling stays set by the dataMedium: data curation dominates
Prompt and scaffold changesThe gap is instruction following or routing, verifiable in daysThe gains live in the prompt, not the weights, so they often evaporate on a model swap: the new checkpoint never learned your instructions and quietly re-derives its own behaviorLowest
RLVR checkpointThe deficit is reasoning depth, not orchestrationUnder-transfer: sharpens reasoning already latent in the base model, so tool-use behavior inside your scaffold may barely moveMedium-low, but the policy is someone else's

The checkpoint row carries the subtle trap. One study of RLVR checkpoints concludes that reinforcement learning with verifiable rewards (RLVR) mainly elicits reasoning capacity the base model already has rather than adding new capability. If your evals show the model reasoning correctly but acting badly inside the harness, that is an orchestration deficit, and the question of an RLVR checkpoint versus fine-tuning your own agent answers itself: when the deficit is orchestration, only in-harness training optimizes the thing you ship.

A Low-Regret Adoption Path

  1. Measure tool latency first. Instrument production or replay traffic until you have p50 and p95 per tool. This one distribution decides whether in-harness RL is affordable at all.
  2. Shadow the tool state. Stand up a sandboxed copy or record-replay layer for every tool the reward reads, and keep the reward channel outside the agent's write surface from day one.
  3. Run shadow rollouts. Collect trajectories and compute rewards with weight updates disabled. You are validating reward robustness and your latency arithmetic, and hunting hacks, before any gradients exist.
  4. Canary with a small model. A 7 to 9B model over a few hundred rollouts surfaces control-plane bugs cheaply; recall the reference pipeline trained on about 6,000 samples total.
  5. Scale parallelism last. Only once reward behavior, latency, and eval bands look sane should you raise rollout concurrency and model size together.

Where the Fidelity Argument Stops

The honest boundary: if your task needs negligible scaffolding, a single model call, no tools, no retries, then harness fidelity stops mattering and the in-harness argument weakens to zero. Pure reasoning work sits here. Train harness-free with RLVR-style methods and save the engineering.

What remains unproven: the 41.8% to 56.4% result is one coding pipeline, one model, one benchmark, reported by the framework's own authors. That in-harness training beats a well-executed reimplementation across domains is plausible, given the policy-identity argument, but it is not yet established by controlled comparisons. What would change this recommendation: independent replications across different harnesses, and transfer studies showing that in-harness-trained weights degrade less than reimplementation-trained weights when the scaffold changes. Until then, agentic reinforcement learning in the harness you ship is the right default exactly when your scaffold is thick, your rewards read tool state you can protect, and your latency budget survives the arithmetic. Anywhere else, say so plainly and pick a different row of the table.

Stay in the loop.

Get the latest posts and exclusive content delivered to your inbox.

Join 10 readers. No spam. Unsubscribe in one click, anytime.

About the author

Megan Caldwell

AI Engineering Lead

Megan has spent the last eight years building production ML systems, from recommendation engines to today's language model pipelines. She writes about the engineering that holds up under real load: retrieval, evaluation, and the unglamorous parts of shipping AI software.

Related Posts