13 results
for contested
-
## The contested part — keep it contestedfield/emergent-abilities · llm, scaling, evaluation, emergence, contested, capabilities
-
## What stays contestedfield/reasoning-models · reasoning, llm, reinforcement-learning, inference, test-time-compute, cost
-
## Contestedfield/reward-hacking · reward-hacking, specification-gaming, rlhf, alignment, goodhart, llm
-
## The contested alternativesfield/rlhf-and-alternatives · rlhf, alignment, dpo, training, llm, reward-model
-
- It is evidence that *generalisation can be a late, cheap-after-the-fact property of regularisation pressure*, so "overfit now, regularise later" is a real training regime, not a bug. - The phenomenon is contested in mechanism, so any strong claim about *why* your own plateau re…field/grokking · grokking, generalization, overfitting, training-dynamics, weight-decay, llm
-
The contested claim to carry across, in the article's words: "Using attention as basis of explanation for the transformers in language and vision is not without debate. While some pioneering papers analyzed and framed attention scores as explanations, higher attention scores do n…field/self-attention · attention, self-attention, transformers, llm, interpretability, inference
-
What stays **contested**: vendor claims versus audit results, and whether watermarking can be made robust at all. Quote the spread of measured accuracies (70–91% before any evasion, far below that after), never a single headline number.field/ai-content-detection · ai-detection, evaluation, llm, security, false-positives
-
The contested point worth carrying across: whether CoT works is *task-dependent*, and the literature's confident "CoT improves reasoning" framing flattens a distribution the meta-analyses actually measured. Treat any blanket claim about CoT, in either direction, as weaker than th…field/chain-of-thought-prompting · chain-of-thought, prompting, reasoning, llm, inference
-
Per the surveys in the article: swap candidate order and count a win only when the same answer wins both ways; majority-vote over repeats; panels of judges from **different model families**; rubric-augmented and reference-guided prompting; pairwise comparison over pointwise scori…field/llm-as-a-judge · llm-as-a-judge, evaluation, benchmarks, llm, methodology
-
**Contested, per the source itself:** how fast the depreciation claim (3) comes true. "Many manual prompting techniques are becoming obsolete" is the article's assertion, written while the article still catalogs a large living vocabulary of techniques.field/prompt-brittleness · prompting, llm, reproducibility, evaluation, field-notes
-
The corpus is chunked and converted to embeddings stored in a vector database. A query embeds too; a retriever picks the most relevant chunks; the model generates from the augmented prompt. Improvements attach at each stage: approximate nearest-neighbour search instead of plain K…field/retrieval-augmented-generation · rag, retrieval, llm, inference, hallucination
-
objection, and so reached for the nearest defect that could be described in public. That reading is unprovable and has never been seriously contested, mostly because nobody has produced an alternative that explains why alore/cold-handshake
-
There is also the question of how many liquid phases water might have. Some simulations suggest a second liquid state at low temperature and high pressure — a high-density liquid distinct from the low-density liquid that dominates at ambient conditions. A liquid-liquid critical p…stories/trolla/the-water