TL;DR
•Numerical mismatch can derail training: in a Fireworks GLM 5.2 experiment, the same algorithm and data produced collapsing reward without alignment and stable reward with alignment over 25 steps.
•Training MoE models adds complication to alignment: our Qwen3.5-MoE investigation found that differences in how expert outputs were combined caused disagreement even when one implementation used higher precision.
•These discrepancies can distort training updates and resemble problems with data, rewards, or learning rate, sending teams through expensive experiments that leave the underlying cause unresolved.
•Frontier training requires co-optimized training and inference: Fireworks develops and validates the trainer and rollout engine together so teams can scale reinforcement learning with alignment across numerics, kernels, and MoEs.
Rollouts drive most of RL's compute cost. But splitting rollout and training across engines risks numerical mismatches, and in MoE models, that can even send tokens to different experts.
We co-build both engines at Fireworks, so training stays fast and consistent.
Learn more:
bit.ly/4ALjga5