A suit of multilingual MoE models with highly-sparse architectures
-
AIDC-AI/Marco-Nano-Base
Text Generation β’ 8B β’ Updated β’ 446 β’ 12 -
AIDC-AI/Marco-Mini-Base
Text Generation β’ 17B β’ Updated β’ 483 β’ 4 -
AIDC-AI/Marco-Mini-Global-Base
Text Generation β’ 17B β’ Updated β’ 406 β’ 5 -
AIDC-AI/Marco-Nano-Instruct
Text Generation β’ 8B β’ Updated β’ 1.7k β’ 28