Weiyun1025 commited on
Commit
5f61851
·
verified ·
1 Parent(s): e9d0cf2

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -103,7 +103,7 @@ In this work, we use the Best-of-N evaluation strategy and employ [VisualPRM-8B]
103
 
104
  ### Multimodal Reasoning and Mathematics
105
 
106
- ![image/png](https://huggingface.co/datasets/Weiyun1025/InternVL-Performance/resolve/main/internvl3/reasoning.png)
107
 
108
  ### OCR, Chart, and Document Understanding
109
 
@@ -159,7 +159,7 @@ The evaluation results in the Figure below shows that the model with native mult
159
 
160
  As shown in the table below, models fine-tuned with MPO demonstrate superior reasoning performance across seven multimodal reasoning benchmarks compared to their counterparts without MPO. Specifically, InternVL3-78B and InternVL3-38B outperform their counterparts by 4.1 and 4.5 points, respectively. Notably, the training data used for MPO is a subset of that used for SFT, indicating that the performance improvements primarily stem from the training algorithm rather than the training data.
161
 
162
- ![image/png](https://huggingface.co/datasets/Weiyun1025/InternVL-Performance/resolve/main/internvl3/ablation-mpo.png)
163
 
164
  ### Variable Visual Position Encoding
165
 
 
103
 
104
  ### Multimodal Reasoning and Mathematics
105
 
106
+ ![image/png](https://huggingface.co/OpenGVLab/VisualPRM-8B-v1_1/resolve/main/visualprm-performance.png)
107
 
108
  ### OCR, Chart, and Document Understanding
109
 
 
159
 
160
  As shown in the table below, models fine-tuned with MPO demonstrate superior reasoning performance across seven multimodal reasoning benchmarks compared to their counterparts without MPO. Specifically, InternVL3-78B and InternVL3-38B outperform their counterparts by 4.1 and 4.5 points, respectively. Notably, the training data used for MPO is a subset of that used for SFT, indicating that the performance improvements primarily stem from the training algorithm rather than the training data.
161
 
162
+ ![image/png](https://huggingface.co/datasets/OpenGVLab/MMPR-v1.2-prompts/resolve/main/ablation-mpo.png)
163
 
164
  ### Variable Visual Position Encoding
165