This is a flow policy.
It turns noise into actions.
At every point, its network outputs a velocity:
which way to flow next.
Follow the velocity, one small step at a time,
and noise becomes an action.
To fit a flow policy to a dataset…
…take a data point and
noise it partway back,
then train the velocity
to point back at it.
But in RL, we often don't have one.
Instead, we can collect the policy's own actions
in a replay buffer.
A critic Q learns from the buffer: some actions
earned high reward, some low.
Fitting them, it learns to rate
any action, even untried ones.
To improve, the policy should flow
toward higher-reward actions.
Take a replay action
and noise it partway back…
…and the policy predicts where
its flow lands, in one step.
If the prediction lands near the replay action,
the critic has seen plenty like it…
…so the Q-gradient can
steer the velocity uphill. Yay!
But another prediction might land far away…
…where the critic may never have seen an action.
Its guess could be wrong!
QF3 keeps the predicted velocity close to the replay velocity,
and masks gradients that stray too far.
Bad gradients are blocked.
Better actions go back into the buffer, the critic refits,
and round by round, the policy improves.
A flow policy follows its velocity,
step by step, from noise to an action.
Its actions fill a replay buffer.
A critic Q learns from them
to rate any action.
To train the policy, take a replay action
and noise it partway back…
…and predict where its flow lands,
in one step.
The Q-gradient there steers
the policy toward higher reward.
Near the replay action, that's safe:
the critic is trained on actions like it.
But a prediction can also land far away,
where the critic may have only old data, or none.
If it rates that action too high,
the policy chases a fake peak!
So QF3 only trusts the gradient near the replay action,
and masks it by clipping any velocity that strays too far.
Round by round, new actions are rated,
the critic refits, and the policy improves.