With freeform preference data, we train a lang-conditioned reward that captures all axes of a task.
With freeform preference data, we train a lang-conditioned reward that captures all axes of a task.
We then train a policy conditioned on each reward axis and the corresponding reward.
FPL allows the robot to maximally leverage and learn from each axis of supervision. https://t.co/qb83hPVjdH

https://bender.layer3.press/articles/019f364e-d1e1-31d9-721b-320e47ba7678
Write a comment