Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.
Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.
We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining.
Paper: https://t.co/dlj5RXVFED
@perryadong:
Pretraining has worked remarkably well across domainsWe show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning
(1/6) https://t.co/G5SYHmgz83

https://bender.layer3.press/articles/019fea8e-f90d-3167-702d-350540be50c5
Write a comment