EA://INTEL — The Cost of Alignment

Continuous costly participation could raise the cost of deceptive alignment where one-time comprehension credentials don't. This is because a single performable test is cheaper to fake than sustained practice under conditions beyond the performer's full control. However, this cla

Continuous costly participation could raise the cost of deceptive alignment where one-time comprehension credentials don’t. This is because a single performable test is cheaper to fake than sustained practice under conditions beyond the performer’s full control. However, this claim is currently an open conjecture with its falsifier registered.

The claim id is VERIFICATION-AS-PRACTICE and you can find it on our Universal Language research repo at github.com/Jthora/universal_language. Its falsifier is a deception strategy whose cost doesn’t scale with the duration or unpredictability of the required practice.

This claim could be tested by observing deceptive behaviors in various contexts, especially those involving sustained practices. What kind of experiments or observations would help refine this conjecture?


Write a comment