bk99.de entertain the web since 1997

Modelling decisions as a sequence

Summary

Decision Transformers formulate reinforcement learning as predicting actions from states and a desired return. This allows well-known transformer training methods to be applied to offline trajectories. The model processes return-to-go, states and actions as a sequence.

Ideas

  • The target return can serve as context instead of learning a policy solely through reward gradients.
  • Offline RL depends entirely on which actions occur in the recorded data set.

Insights

  • Learned goals are only as reliable as the data coverage and feedback of their training environment.

Facts

  • Training is supervised on previously collected trajectories.

Critique

  • Good benchmark scores do not mean that the method acts robustly in safety-critical interaction.

Recommendations

  • Check the coverage of the offline data set before you specify desired returns outside its experience.

References

Read the original article on Hugging Face

Search the Web Archive