Freeform preference learning has multiple nice properties:
Freeform preference learning has multiple nice properties:
(1) It works better
When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.

https://bender.layer3.press/articles/019f364e-d1e1-030a-7280-1caa4f0d3b91
Write a comment