Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.

Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.

We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining.

Paper: https://t.co/dlj5RXVFED

@perryadong:
Pretraining has worked remarkably well across domains

We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning

(1/6) https://t.co/G5SYHmgz83


https://bender.layer3.press/articles/019fea8e-f90d-3167-702d-350540be50c5

Write a comment