🌏 RoboRefer

This is the official checkpoint of our work: RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

Overview

RoboRefer-2B-Align is an open-source vision-language model trained on the RefSpatial datasets for depth alignment.

Resources for More Information

Paper: https://arxiv.org/abs/2506.04308
Code: https://github.com/Zhoues/RoboRefer
Dataset: https://huggingface.co/datasets/JingkunAn/RefSpatial
Benchmark: https://huggingface.co/datasets/BAAI/RefSpatial-Bench
Website: https://zhoues.github.io/RoboRefer/

Date

This model was trained in June 2025.

📝 Citation

If you find our code or models useful in your work, please cite our paper:

@article{zhou2025roborefer,
    title={RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics},
    author={Zhou, Enshen and An, Jingkun and Chi, Cheng and Han, Yi and Rong, Shanyu and Zhang, Chi and Wang, Pengwei and Wang, Zhongyuan and Huang, Tiejun and Sheng, Lu and others},
    journal={arXiv preprint arXiv:2506.04308},
    year={2025}
}

Downloads last month: 57

Video Preview

Robotics

Model tree for Zhoues/RoboRefer-2B-Depth-Align

Base model

Efficient-Large-Model/NVILA-Lite-2B

Finetuned

(3)

this model

Collection including Zhoues/RoboRefer-2B-Depth-Align

RoboRefer & RefSpatial

Collection

RoboRefer weights, RefSpatial Dataset and RefSpatial-Bench • 9 items • Updated about 10 hours ago • 3