QF3
Fast Flow RL with Filtered Q-Gradients
TL;DR: QF3 stabilizes off-policy RL for flow policies by clipping critic‑driven velocity updates around the replay data, and we validate it on humanoid locomotion and manipulation fine-tuning.
Improving Flow Policies with RL
Flow policies are widely used across robotics today, but how can we improve them with reinforcement learning? The most sample-efficient and direct approach would be to follow the critic’s gradient \nabla_a Q through the flow’s one-step prediction, but this can become unstable when the policy chases the critic’s overestimates.
A flow policy turns noise into actions.
It follows its velocity, one small step at a time, from noise to an action.
Given a dataset, it learns by flow matching:
take a data point and noise it partway back,
then train the velocity to point back at it.
But what if there is no dataset?
Training flow policies with off-policy RL.
Well, it can learn from a replay buffer of its own actions.
A critic Q is trained to predict each action’s reward, so it can rate any action, even untried ones.
Steer the velocity with critic gradients.
To improve, the policy should flow toward higher-reward actions.
QF3 takes a replay action and noises it partway back…
…and predicts where its flow lands, in one step.
\nabla_a Q there steers the velocity toward higher reward.
This can work well. But when is this safe?
Where can the critic gradient be trusted?
If a prediction lands far away, where the critic has no data…
If it rates that action too high, the policy might chase a fake peak!
QF3 masks the gradient where the critic can’t be trusted.
We prevent the policy from using this bad gradient,
…by clipping the velocity to stay near the replay action.
Inside the clip window, we can follow \nabla_a Q more safely.
Round by round, it can find new behaviors.
Better actions enter the buffer, and the critic refits…
…and the policy’s flow shifts toward them.
To avoid chasing overestimates, QF3 follows the pathwise Q-gradient only when the policy’s velocity v_\theta is sufficiently close to the replay action’s conditional velocity u_{\mathrm{buf}}:
QF3’s actor update, for a replay action a_{\mathrm{buf}}, noise \varepsilon and flow time \tau:
For a replay action a_{\mathrm{buf}}, noise \varepsilon and flow time \tau:
We achieve this by clipping the velocity difference \textcolor{#c98a3b}{\delta_\theta}, which masks the critic’s gradient once the policy strays too far from the replay data. We can also interpret this clip in terms of likelihood, as restricting \nabla_a Q to near-on-policy samples, since a small flow-matching residual means (via the ELBO) a likely action.
Why use the critic’s gradient?
A Q-value only scores the action it is given, so finding a direction of improvement means sampling many actions and ranking or reweighting them by Q. \nabla_a Q gives this direction directly, in every action dimension, from a single action.
How does the clip work?
One actor update in one action dimension: flow time \tau runs left to right, and the strips on the right are the policy \pi(a) and the critic Q(a).
1Sample a replay action a_{\mathrm{buf}}, noise \varepsilon, and flow time \tau.x_\tau = (1-\tau)\,\varepsilon + \tau\, a_{\mathrm{buf}}
2The flow-matching target is the conditional velocity.u_{\mathrm{buf}} = a_{\mathrm{buf}} - \varepsilon
3The policy’s one-step prediction lands away from a_{\mathrm{buf}}.\hat x_1 = x_\tau + (1-\tau)\, v_\theta
4QF3 instead clips v_\theta to u_{\mathrm{buf}} \pm \alpha per dimension, keeping the prediction near a_{\mathrm{buf}}.\hat x_{1,\mathrm{clip}} = x_\tau + (1-\tau)\big(u_{\mathrm{buf}} + \operatorname{clip}(v_\theta - u_{\mathrm{buf}}, -\alpha, \alpha)\big)
5The actor maximizes Q at the clipped prediction, alongside flow matching.\mathcal{L}_{\text{actor}} = -Q(s, \hat x_{1,\mathrm{clip}}) + \lambda \lVert v_\theta - u_{\mathrm{buf}} \rVert^2
6\nabla_a Q pulls \hat x_{1,\mathrm{clip}} away from the data, but past the clip that pull no longer reaches v_\theta.\partial Q / \partial v_\theta = 0
How does QF3 compare to prior work?
FlowRL samples an action a_\theta from the flow and backpropagates Q through the full sampling chain. It also flow-matches onto replay actions a, weighted by how much a outscores a_\theta under Q^{\beta^*}, a critic for the best behavior in the buffer.
vs. QF3. FlowRL learns an additional critic and squashes actions with a tanh for stability. QF3 needs only the standard critic and clips the velocity instead.
EXPO keeps its expressive base policy, and trains a small Gaussian edit policy \pi_{\text{edit}}(\Delta \mid s, a) with SAC to shift actions by at most \beta toward higher Q. It acts with the highest-Q of N base samples and their edits.
vs. QF3. Q never updates the base policy; improvement lives in a helper network and best-of-N sampling. QF3 trains the flow policy itself.
FPO/FPO++ scores each executed action a by its GAE advantage \hat A. At each sample (\varepsilon_i, \tau_i), the flow-matching loss takes the place of the log-likelihood in PPO. \hat A > 0 pulls v_\theta toward u_i, and \hat A < 0 pushes it away.
vs. QF3. Every update needs fresh rollouts, since FPO/FPO++ is on-policy and has no action critic. QF3 reuses a replay buffer and follows \nabla_a Q.
DSRL freezes the pretrained flow and trains a separate noise policy \pi^W with SAC to pick the flow’s input noise w.
vs. QF3. Steering only the noise keeps actions within what the pretrained flow can already produce; QF3 trains the flow policy itself.
OGPO trains a critic Q on the flow policy’s actions. To improve the policy, it treats each denoising chain as a short trajectory and runs PPO on it. A chain’s advantage is its final action’s Q minus the mean over a group of chains from the same state.
vs. QF3. OGPO only uses Q’s values to rank chains. QF3 uses \nabla_a Q directly.
FlowDPG steps the one-step estimate \hat x_1 along \nabla_a Q to an improved action a^\star. It then regresses v_\theta onto an interpolation between the velocity u^\star = a^\star - \varepsilon toward a^\star and the replay velocity u.
vs. QF3. FlowDPG turns \nabla_a Q into a regression target by shifting the one-step estimate, while QF3 maximizes Q by backpropagating through the one-step estimate, then clips the velocity to stay near the data.
Results
In simulation, we use QF3 to finetune the pretrained bimanual manipulation policies (ABC-VLA), training small LoRA adapters on their flow action head while the VLM stays frozen. We find QF3 speeds up both tasks, either by speeding up the motion or by settling on a more efficient behavior.
We also show QF3 on robomimic tasks (full finetuning)!
We train humanoid policies from scratch by dropping QF3's actor update into a high-throughput off-policy recipe, and deploy them on hardware. Because QF3 changes only the actor, the rest of FastTD3, including massively parallel simulation, large-batch updates, and a distributional critic, can carry over unchanged.
We find that the same update can also be used for finetuning a flow-based text-to-image model (Stable Diffusion 3.5 Medium)! We apply QF3's actor update by treating each prompt as a one-step episode and using a pretrained differentiable reward model directly as the critic.
Citation
@article{kim2026qf3,
title={{QF3}: Fast Flow {RL} with Filtered {Q}-Gradients},
author={Kim, Chung Min and Yi, Brent and McAllister, David and Choi, Hongsuk and Singh, Himanshu Gaurav and Cao, Jinkun and Goldberg, Ken and Abbeel, Pieter and Sferrazza, Carmelo and Kanazawa, Angjoo},
journal={arXiv preprint arXiv:2610.08789},
year={2026}
}