Agentic Reinforcement Learning in the Harness You Ship
Agentic reinforcement learning belongs in the harness you ship, not a training clone. Learn the five failure modes and when to skip RL entirely.
Category
Technical deep dives into how AI systems are built and run: models, retrieval, agents, evaluation, fine-tuning, and infrastructure.
24 posts
Agentic reinforcement learning belongs in the harness you ship, not a training clone. Learn the five failure modes and when to skip RL entirely.
Git for AI coding agents: how agent fleets multiply clone, ref, and CI load, and the shallow, partial, and sparse checkout patterns that cap it.
MCP agent verification grades each claim's source, not just truth. The decision grid, provenance binding, and eval additions catch wrong-source answers.
Pass@k crossovers show RLVR sharpens rather than adds capability. Learn the decision rule and statistical test that find your model's crossover point.
Async GRPO training with LoRA swaps NCCL for a bucket and proxy, turning RL fine-tuning from reserved-cluster spend into elastic spot-GPU economics.
AI agent monitoring fails when dashboards track requests, not resolutions. This taxonomy maps silent failure modes to the signals that catch them.
Training a diffusion model from scratch can cost hundreds, not millions. The full math behind a 210M DiT built in 3.5 days on one RTX PRO 6000.
KV cache math explains why every decoded token reads the whole context, and what eviction and quantization can cut from million-token agent bills today.
Public benchmarks say 89%, your warehouse says otherwise. Build a text-to-SQL evaluation with schema-specific oracles that catches silent wrong answers.
MoE serving cost for a 6-of-125B model is not 6B per token. All 125B stay in VRAM, so run the builder math on residency, routing, and break-even.
Speculative decoding turns idle CPU cores into 4x faster LLM generation. Learn why it works, when gains collapse, and when CPU beats GPU or API.
Gradient accumulation can make identical batches train at different speeds. Learn why micro-batch shape drives T4 vs L4 wall-clock time and throughput.