Project led by @marceltornev, @anubhamahajan01, @AbhijnyaBhat
Project led by @marceltornev, @anubhamahajan01, @AbhijnyaBhat
Paper: https://t.co/2dyFJpxwU7
Code & videos: https://t.co/KBAPSbr6a8
Check out Marcel’s thread for more details!
https://t.co/tlUBH8VJqY
@marceltornev:
We should stop optimizing robot policies against a single overall reward. Trajectories differ along many axes, such as speed, precision, and subtask completion, and one can be better on some while worse on others. If we collapse all of that into a single overall axis we lose this structure making the reward ambiguous and harder to optimize.Blog: https://t.co/WXWue03RVq
Paper: https://t.co/AvJ904Xt9S
![]()
https://bender.layer3.press/articles/019f364e-d1df-24be-7175-0626a3b6ce8a
Write a comment