bk99.de entertain the web since 1997

Mixtral and mixture of experts

Summary

Mixtral activates only part of its experts per token and combines high model capacity with limited computation. The article explains the architecture, usage and open availability. Mixtral uses eight experts and activates two of them per token.

Ideas

  • Mixture of experts separates the total number of parameters from the active computing effort.
  • Routing becomes an additional model component with its own load and quality questions.

Insights

  • Open weights shift control to users and in return require more testing competence of their own.

Facts

  • The model weights were made openly available on the Hub.

Critique

  • The inactive experts save computing time, but their weights still have to be stored and moved.

Recommendations

  • Measure memory requirements, active computing costs and routing behaviour instead of only comparing the number of parameters.

References

Read the original article on Hugging Face

Search the Web Archive