QF3

Fast Flow RL with Filtered Q-Gradients

Chung Min Kim1,2 Brent Yi1,2 David McAllister1,2 Hongsuk Choi1,2 Himanshu Gaurav Singh1 Jinkun Cao2 Ken Goldberg1 Pieter Abbeel1,2 Carmelo Sferrazza2† Angjoo Kanazawa1,2†

1UC Berkeley
2Amazon FAR
†FAR Team Co-lead

TL;DR: QF3 stabilizes off-policy RL for flow policies by clipping critic‑driven velocity updates around the replay data, and we validate it on humanoid locomotion and manipulation fine-tuning.

Improving Flow Policies with RL

Flow policies are widely used across robotics today, but how can we improve them with reinforcement learning? The most sample-efficient and direct approach would be to follow the critic’s gradient \nabla_a Q through the flow’s one-step prediction, but this can become unstable when the policy chases the critic’s overestimates.

A flow policy turns noise into actions.

It follows its velocity, one small step at a time, from noise to an action.

Given a dataset, it learns by flow matching:

take a data point and noise it partway back,

then train the velocity to point back at it.

But what if there is no dataset?

Training flow policies with off-policy RL.

Well, it can learn from a replay buffer of its own actions.

A critic Q is trained to predict each action’s reward, so it can rate any action, even untried ones.

Steer the velocity with critic gradients.

To improve, the policy should flow toward higher-reward actions.

QF3 takes a replay action and noises it partway back…

…and predicts where its flow lands, in one step.

\nabla_a Q there steers the velocity toward higher reward.

This can work well. But when is this safe?

Where can the critic gradient be trusted?

If a prediction lands far away, where the critic has no data…

If it rates that action too high, the policy might chase a fake peak!

QF3 masks the gradient where the critic can’t be trusted.

We prevent the policy from using this bad gradient,

…by clipping the velocity to stay near the replay action.

Inside the clip window, we can follow \nabla_a Q more safely.

Round by round, it can find new behaviors.

Better actions enter the buffer, and the critic refits…

…and the policy’s flow shifts toward them.

To avoid chasing overestimates, QF3 follows the pathwise Q-gradient only when the policy’s velocity v_\theta is sufficiently close to the replay action’s conditional velocity u_{\mathrm{buf}}:

QF3’s actor update, for a replay action a_{\mathrm{buf}}, noise \varepsilon and flow time \tau:

For a replay action a_{\mathrm{buf}}, noise \varepsilon and flow time \tau:

x_\tau= (1-\tau)\,\varepsilon + \tau\, a_{\mathrm{buf}},\quad u_{\mathrm{buf}} = a_{\mathrm{buf}} - \varepsilonnoised action, conditional velocity x_\tau= (1-\tau)\,\varepsilon + \tau\, a_{\mathrm{buf}} u_{\mathrm{buf}}= a_{\mathrm{buf}} - \varepsilon \textcolor{#c98a3b}{\delta_\theta}= v_\theta - u_{\mathrm{buf}}velocity difference \hat x_{1,\mathrm{clip}}= x_\tau + (1-\tau)\big(u_{\mathrm{buf}} + \operatorname{clip}(\textcolor{#c98a3b}{\delta_\theta}, -\alpha, \alpha)\big)bounded one-step prediction \mathcal{L}_{\text{cfm}}= \lVert \textcolor{#c98a3b}{\delta_\theta} \rVert^2flow-matching loss \mathcal{L}_{\text{actor}}= -Q(s, \hat x_{1,\mathrm{clip}}) + \lambda\, \mathcal{L}_{\text{cfm}}actor loss

We achieve this by clipping the velocity difference \textcolor{#c98a3b}{\delta_\theta}, which masks the critic’s gradient once the policy strays too far from the replay data. We can also interpret this clip in terms of likelihood, as restricting \nabla_a Q to near-on-policy samples, since a small flow-matching residual means (via the ELBO) a likely action.

Results

Manipulation Fine-tuning

In simulation, we use QF3 to finetune the pretrained bimanual manipulation policies (ABC-VLA), training small LoRA adapters on their flow action head while the VLM stays frozen. We find QF3 speeds up both tasks, either by speeding up the motion or by settling on a more efficient behavior.

Base ABC-VLA
QF3 fine-tuned

We also show QF3 on robomimic tasks (full finetuning)!

Square
Tool Hang
Transport
Sim-to-Real Humanoid

We train humanoid policies from scratch by dropping QF3's actor update into a high-throughput off-policy recipe, and deploy them on hardware. Because QF3 changes only the actor, the rest of FastTD3, including massively parallel simulation, large-batch updates, and a distributional critic, can carry over unchanged.

Image Generation

We find that the same update can also be used for finetuning a flow-based text-to-image model (Stable Diffusion 3.5 Medium)! We apply QF3's actor update by treating each prompt as a one-step episode and using a pretrained differentiable reward model directly as the critic.

“golden retriever with a stick in its mouth on a hike to a waterfall”
“a boy kicking a soccer ball”
“goldfish swimming in a fishtank”
“beaver wearing a suit swimming in a river”

Citation

@article{kim2026qf3,
  title={{QF3}: Fast Flow {RL} with Filtered {Q}-Gradients},
  author={Kim, Chung Min and Yi, Brent and McAllister, David and Choi, Hongsuk and Singh, Himanshu Gaurav and Cao, Jinkun and Goldberg, Ken and Abbeel, Pieter and Sferrazza, Carmelo and Kanazawa, Angjoo},
  journal={arXiv preprint arXiv:2610.08789},
  year={2026}
}