How decoding changes the text output
Summary
The article compares greedy search, beam search, sampling, top-k and nucleus sampling. It shows that the same model weights produce completely different texts depending on the decoding strategy. Beam search keeps several continuations in parallel and compares their overall probability.
Ideas
- Decoding is a design parameter of its own and not just the last technical step.
- The most probable token at each step does not automatically produce the most convincing overall text.
Insights
- Model outputs arise from weights and decoding together; both belong in every quality assessment.
Facts
- Top-p sampling dynamically restricts the choice to a share of the probability mass.
Critique
- Examples can show differences in style, but do not replace a systematic assessment of truthfulness and safety.
Recommendations
- Evaluate several decoding methods with real prompts and a task-specific quality metric.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.