bk99.de entertain the web since 1997

How decoding changes the text output

Summary

The article compares greedy search, beam search, sampling, top-k and nucleus sampling. It shows that the same model weights produce completely different texts depending on the decoding strategy. Beam search keeps several continuations in parallel and compares their overall probability.

Ideas

  • Decoding is a design parameter of its own and not just the last technical step.
  • The most probable token at each step does not automatically produce the most convincing overall text.

Insights

  • Model outputs arise from weights and decoding together; both belong in every quality assessment.

Facts

  • Top-p sampling dynamically restricts the choice to a share of the probability mass.

Critique

  • Examples can show differences in style, but do not replace a systematic assessment of truthfulness and safety.

Recommendations

  • Evaluate several decoding methods with real prompts and a task-specific quality metric.

References

Read the original article on Hugging Face

Search the Web Archive