Freeform preference learning has multiple nice properties:

Freeform preference learning has multiple nice properties:

(1) It works better

When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.

https://bender.layer3.press/articles/019f364e-d1e1-030a-7280-1caa4f0d3b91

Write a comment