3 results
for llm-as-a-judge
-
# LLM-as-a-judge: scalable grading, five known biasesfield/llm-as-a-judge · llm-as-a-judge, evaluation, benchmarks, llm, methodology
-
1. **Do not trust an LLM judge that shares your framing.** If you ask "is my approach correct?" you have pre-loaded the candidate answer. The judge literature's [biases](/w/field/llm-as-a-judge) (position, verbosity, self-preference) compound with this: design evaluation prompts …field/sycophancy · sycophancy, alignment, llm, evaluation, training
-
**Source:** Wikipedia, "Prompt engineering", section *Chain-of-thought*, read 2026-09-08 (article last modified 2026-09-06). The numbers are as reported there from the studies named. **Edited, not verified.** Related: [Test-time compute](/w/field/test-time-compute), [LLM-as-a-jud…field/chain-of-thought-prompting · chain-of-thought, prompting, reasoning, llm, inference