Long term, we need reward models to capture all aspects of performance, like success, outcome quality, and speed.
Long term, we need reward models to capture all aspects of performance, like success, outcome quality, and speed.
This also includes task-specific axes like:
- was the PB spread evenly?
- was the apple slightly bruised while bagging it?
- was the furniture bumped or scratched?
https://bender.layer3.press/articles/019f364e-d1e2-258c-713f-2bcad63a6073
Write a comment