ZIXUAN WANG*, 2026
<aside> 📌
This is a personal technical writeup. Work done during an internship on MiroMind RL team. Technical writeup only for personal learning
</aside>
<aside> 🧊
TL;DR In every modern RL framework, one engine samples your rollouts (vLLM / SGLang) and a different engine learns from them (FSDP / Megatron). They share the same weights $\theta$ but not the same arithmetic, so they are two different policies, and training silently goes off-policy.
Here is the entire problem in one inequality. Two engines share parameters $\theta$, but differ in kernels, precision, and MoE discrete routing, so they realize two distinct distributions:
$$ \underbrace{\textcolor{red}{\pi_{\mathrm{infer}}}(\cdot\,;\theta)}{\text{sampler rolls out (vLLM / SGLang)}} \;\neq\; \underbrace{\textcolor{blue}{\pi{\mathrm{train}}}(\cdot\,;\theta)}_{\text{learner scores \& differentiates (FSDP / Megatron)}} $$
even though both are "the model with weights $\theta$." A numerically tiny per-token disagreement is, formally, an off-policy bug. The roadmap:
<aside> 🎨

Figure 1. Rollout Routing Replay (R³) at a glance. The $\textcolor{red}{\mathsf{Record}}$ (inference) → $\textcolor{green}{\mathsf{Replay}}$ (bridge) → $\textcolor{blue}{\mathsf{Train}}$ flow: the inference engine records its per-token top-$K$ expert mask during rollout, that exact mask is replayed into the training forward pass, so the learner scores precisely the experts that fired. (Mechanism detailed in Part 3.)
Vanilla policy gradient for a response $a$ with reward $R(a)$:
$$ \theta \leftarrow \theta + \mu\, \underbrace{\mathbb{E}{a\sim \textcolor{blue}{\pi{\mathrm{train}}}(\theta)}\!\Big[\,R(a)\,\nabla_\theta \log \textcolor{blue}{\pi_{\mathrm{train}}}(a;\theta)\,\Big]}_{\text{on-policy: sample and score with the SAME distribution}} $$
This estimator is unbiased only because the sampling distribution and the scored distribution coincide.