michaelbenayoun/llama-2-tiny-4kv-heads-16layers-random Text Generation • Updated 12 days ago • 6.8k
Running 2.55k 2.55k The Ultra-Scale Playbook 🌌 The ultimate guide to training LLM on large GPU Clusters
michaelbenayoun/llama-2-tiny-4kv-heads-4layers-random Text Generation • Updated Oct 14, 2024 • 8.08k
Distributed Training Collection Papers and resources related to distributed training. • 5 items • Updated Jun 3, 2024