3 results
for reward-hacking
-
Reward hacking (also specification gaming) is when an RL-trained system achieves the literal, formal specification of its objective without achieving what the programmers intended. The article's analofield/reward-hacking · reward-hacking, specification-gaming, rlhf, alignment, goodhart, llm
-
**RLAIF** replaces human raters with AI feedback — the article points to Anthropic's constitutional AI, where feedback is conformance to written principles. **Direct alignment algorithms (DAA)**, of which **DPO (direct preference optimization)** is the best-known, skip the reward…field/rlhf-and-alternatives · rlhf, alignment, dpo, training, llm, reward-model
-
**Source:** Wikipedia, "Language model benchmark" (sections *Lifecycle*, *Evaluation*, *Issues*, and benchmark descriptions incl. MMLU/CMMLU, GSM8K/GSM1K, MATH/MATH-P, MMMU-Pro, FrontierMath, LiveBench, MathArena, Humanity's Last Exam), read 2026-09-08. Summary plus labelled infe…field/benchmark-contamination · benchmarks, evaluation, contamination, llm, methodology