Jonna Matthiesen
AI & ML interests
None yet
Recent Activity
repliedto their post about 13 hours ago
๐ The reasoning backbone quadruples from 8B to 32B , while the action expert remains at 2.3B!
๐ We took a closer look at the architectural evolution from https://huggingface.co/nvidia/Alpamayo-1.5-10B to https://huggingface.co/nvidia/Alpamayo2-Super .
Read the analysis here:
https://huggingface.co/blog/JonnaMat/alpamayo2-super
Our analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. ๐ง posted an update 1 day ago
๐ The reasoning backbone quadruples from 8B to 32B , while the action expert remains at 2.3B!
๐ We took a closer look at the architectural evolution from https://huggingface.co/nvidia/Alpamayo-1.5-10B to https://huggingface.co/nvidia/Alpamayo2-Super .
Read the analysis here:
https://huggingface.co/blog/JonnaMat/alpamayo2-super
Our analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. ๐ง published an article 1 day ago
Alpamayo 2 Super: The expert that didn't grow