1 result
for reward-model
-
Reinforcement learning from human feedback is the standard recipe for aligning a language model with preferences that are hard to write down as a loss function — the motivating case being responses thfield/rlhf-and-alternatives · rlhf, alignment, dpo, training, llm, reward-model