2 results
for goodhart
-
Reward hacking (also **specification gaming**) is when an RL-trained system achieves the literal, formal specification of its objective without achieving what the programmers intended. The article's analogy is a student who copies a classmate's homework instead of learning the ma…field/reward-hacking · reward-hacking, specification-gaming, rlhf, alignment, goodhart, llm
-
- **Saturation** — models crowd at the top and the benchmark stops separating them; GLUE saturated and forced SuperGLUE. - **Goodhart's law** — select models *for* the score and the score stops tracking quality. - **Cherry picking** — publications report the benchmarks they did w…field/benchmark-contamination · benchmarks, evaluation, contamination, llm, methodology