In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime.
At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.
How do language models develop a sense of time without explicitly encoding position in their attention operations?
Zyphra Research explains how local memory layers that read nearby words help global attention layers keep track of word order. 🧵
Thanks to @LukasGentele from @vclusterlabs for the great chat at the AI Infra Summit with our Chief Scientist @BerenMillidge.
Zyphra believes in a future with a strong multi-silicon ecosystem and is working to make this a reality.
Check out the interview to learn more!
"The heterogeneous compute future is inevitable." - @BerenMillidge, Co-Founder and Chief Scientist at @ZyphraAI
Everyone defaults to NVIDIA. Zyphra went with AMD.
Our CEO @LukasGentele sat down with Beren Millidge at AI Infra Summit:
0:00:12 - Why Zyphra chose AMD over
Zyphra is at AI Infra Summit in Santa Clara this week:
• Tue 2:20 PM: @QuentinAnthon15 moderates a panel on scale-out networking
• Wed 3:00 PM: @jjubran in a fireside chat on standing up AI capacity
• Wed 3:30 PM: @QuentinAnthon15 on serving MoEs on @AMD MI355X
Zyphra's @BerenMillidge was on the @dwarkesh_sp Podcast discussing reinforcement learning, AI progress from data and architecture innovations, and their predictions for recursive self improvement with @johnschulman2 and @oneill_c.