1 result
for seq2seq
-
There is a seam in how autoregressive models are trained that runs all the way from 2014 seq2seq to today's LLMs, and it is the oldest named example of a train/test distribution shift. During training the decoder is fed the *reference* tokens as context; at inference it must cond…field/exposure-bias · exposure-bias, teacher-forcing, training, llm, seq2seq, distribution-shift