Modelling decisions as a sequence
Summary
Decision Transformers formulate reinforcement learning as predicting actions from states and a desired return. This allows well-known transformer training methods to be applied to offline trajectories. The model processes return-to-go, states and actions as a sequence.
Ideas
- The target return can serve as context instead of learning a policy solely through reward gradients.
- Offline RL depends entirely on which actions occur in the recorded data set.
Insights
- Learned goals are only as reliable as the data coverage and feedback of their training environment.
Facts
- Training is supervised on previously collected trajectories.
Critique
- Good benchmark scores do not mean that the method acts robustly in safety-critical interaction.
Recommendations
- Check the coverage of the offline data set before you specify desired returns outside its experience.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.