Post
566
๐ The reasoning backbone quadruples from 8B to 32B , while the action expert remains at 2.3B!
๐ We took a closer look at the architectural evolution from nvidia/Alpamayo-1.5-10B to nvidia/Alpamayo2-Super .
Read the analysis here:
https://huggingface.co/blog/JonnaMat/alpamayo2-super
Our analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. ๐ง
๐ We took a closer look at the architectural evolution from nvidia/Alpamayo-1.5-10B to nvidia/Alpamayo2-Super .
Read the analysis here:
https://huggingface.co/blog/JonnaMat/alpamayo2-super
Our analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. ๐ง