hustvl
/

mmMamba-linear

Image-Text-to-Text

feature-extraction

Model card Files Files and versions Community

HongyuanTao commited on Feb 18

Commit

9598d88

·

verified ·

1 Parent(s): 4b0576d

Update README.md

Files changed (1) hide show

README.md +2 -2

README.md CHANGED Viewed

@@ -9,11 +9,11 @@ We propose mmMamba, the first decoder-only multimodal state space model achieved
 Distilled from the decoder-only HoVLE-2.6B, our pure Mamba-2-based mmMamba-linear achieves performance competitive with existing linear and quadratic-complexity VLMs, including those with 2x larger parameter size like EVE-7B. The hybrid variant, mmMamba-hybrid, further enhances performance across all benchmarks, approaching the capabilities of the teacher model HoVLE. In long-context scenarios with 103K tokens, mmMamba-linear demonstrates remarkable efficiency gains with a 20.6× speedup and 75.8% GPU memory reduction compared to HoVLE, while mmMamba-hybrid achieves a 13.5× speedup and 60.2% memory savings.
 <div align="center">
-<img src="assets/teaser.png" />
 <b>Seeding strategy and three-stage distillation pipeline of mmMamba.</b>
-<img src="assets/pipeline.png" />
 </div>
 ## Quick Start Guide for mmMamba Inference

 Distilled from the decoder-only HoVLE-2.6B, our pure Mamba-2-based mmMamba-linear achieves performance competitive with existing linear and quadratic-complexity VLMs, including those with 2x larger parameter size like EVE-7B. The hybrid variant, mmMamba-hybrid, further enhances performance across all benchmarks, approaching the capabilities of the teacher model HoVLE. In long-context scenarios with 103K tokens, mmMamba-linear demonstrates remarkable efficiency gains with a 20.6× speedup and 75.8% GPU memory reduction compared to HoVLE, while mmMamba-hybrid achieves a 13.5× speedup and 60.2% memory savings.
 <div align="center">
+<img src="teaser.png" />
 <b>Seeding strategy and three-stage distillation pipeline of mmMamba.</b>
+<img src="pipeline.png" />
 </div>
 ## Quick Start Guide for mmMamba Inference