Delete checkpoint-50
Browse filesThis view is limited to 50 files because it contains too many changes.
See raw diff
- checkpoint-50/README.md +0 -202
- checkpoint-50/adapter_config.json +0 -39
- checkpoint-50/adapter_model.safetensors +0 -3
- checkpoint-50/added_tokens.json +0 -28
- checkpoint-50/chat_template.jinja +0 -89
- checkpoint-50/global_step44/bf16_zero_pp_rank_0_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_1_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_2_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_3_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_4_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_5_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_6_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/bf16_zero_pp_rank_7_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_0_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_1_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_2_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_3_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_4_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_5_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_6_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step44/zero_pp_rank_7_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_0_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_1_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_2_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_3_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_4_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_5_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_6_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/bf16_zero_pp_rank_7_mp_rank_00_optim_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_0_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_1_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_2_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_3_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_4_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_5_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_6_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/global_step50/zero_pp_rank_7_mp_rank_00_model_states.pt +0 -3
- checkpoint-50/latest +0 -1
- checkpoint-50/merges.txt +0 -0
- checkpoint-50/rng_state_0.pth +0 -3
- checkpoint-50/rng_state_1.pth +0 -3
- checkpoint-50/rng_state_2.pth +0 -3
- checkpoint-50/rng_state_3.pth +0 -3
- checkpoint-50/rng_state_4.pth +0 -3
- checkpoint-50/rng_state_5.pth +0 -3
- checkpoint-50/rng_state_6.pth +0 -3
- checkpoint-50/rng_state_7.pth +0 -3
- checkpoint-50/scheduler.pt +0 -3
- checkpoint-50/special_tokens_map.json +0 -31
- checkpoint-50/tokenizer.json +0 -3
checkpoint-50/README.md
DELETED
@@ -1,202 +0,0 @@
|
|
1 |
-
---
|
2 |
-
base_model: Qwen/Qwen3-8B
|
3 |
-
library_name: peft
|
4 |
-
---
|
5 |
-
|
6 |
-
# Model Card for Model ID
|
7 |
-
|
8 |
-
<!-- Provide a quick summary of what the model is/does. -->
|
9 |
-
|
10 |
-
|
11 |
-
|
12 |
-
## Model Details
|
13 |
-
|
14 |
-
### Model Description
|
15 |
-
|
16 |
-
<!-- Provide a longer summary of what this model is. -->
|
17 |
-
|
18 |
-
|
19 |
-
|
20 |
-
- **Developed by:** [More Information Needed]
|
21 |
-
- **Funded by [optional]:** [More Information Needed]
|
22 |
-
- **Shared by [optional]:** [More Information Needed]
|
23 |
-
- **Model type:** [More Information Needed]
|
24 |
-
- **Language(s) (NLP):** [More Information Needed]
|
25 |
-
- **License:** [More Information Needed]
|
26 |
-
- **Finetuned from model [optional]:** [More Information Needed]
|
27 |
-
|
28 |
-
### Model Sources [optional]
|
29 |
-
|
30 |
-
<!-- Provide the basic links for the model. -->
|
31 |
-
|
32 |
-
- **Repository:** [More Information Needed]
|
33 |
-
- **Paper [optional]:** [More Information Needed]
|
34 |
-
- **Demo [optional]:** [More Information Needed]
|
35 |
-
|
36 |
-
## Uses
|
37 |
-
|
38 |
-
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
39 |
-
|
40 |
-
### Direct Use
|
41 |
-
|
42 |
-
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
43 |
-
|
44 |
-
[More Information Needed]
|
45 |
-
|
46 |
-
### Downstream Use [optional]
|
47 |
-
|
48 |
-
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
49 |
-
|
50 |
-
[More Information Needed]
|
51 |
-
|
52 |
-
### Out-of-Scope Use
|
53 |
-
|
54 |
-
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
55 |
-
|
56 |
-
[More Information Needed]
|
57 |
-
|
58 |
-
## Bias, Risks, and Limitations
|
59 |
-
|
60 |
-
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
61 |
-
|
62 |
-
[More Information Needed]
|
63 |
-
|
64 |
-
### Recommendations
|
65 |
-
|
66 |
-
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
67 |
-
|
68 |
-
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
69 |
-
|
70 |
-
## How to Get Started with the Model
|
71 |
-
|
72 |
-
Use the code below to get started with the model.
|
73 |
-
|
74 |
-
[More Information Needed]
|
75 |
-
|
76 |
-
## Training Details
|
77 |
-
|
78 |
-
### Training Data
|
79 |
-
|
80 |
-
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
81 |
-
|
82 |
-
[More Information Needed]
|
83 |
-
|
84 |
-
### Training Procedure
|
85 |
-
|
86 |
-
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
87 |
-
|
88 |
-
#### Preprocessing [optional]
|
89 |
-
|
90 |
-
[More Information Needed]
|
91 |
-
|
92 |
-
|
93 |
-
#### Training Hyperparameters
|
94 |
-
|
95 |
-
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
96 |
-
|
97 |
-
#### Speeds, Sizes, Times [optional]
|
98 |
-
|
99 |
-
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
100 |
-
|
101 |
-
[More Information Needed]
|
102 |
-
|
103 |
-
## Evaluation
|
104 |
-
|
105 |
-
<!-- This section describes the evaluation protocols and provides the results. -->
|
106 |
-
|
107 |
-
### Testing Data, Factors & Metrics
|
108 |
-
|
109 |
-
#### Testing Data
|
110 |
-
|
111 |
-
<!-- This should link to a Dataset Card if possible. -->
|
112 |
-
|
113 |
-
[More Information Needed]
|
114 |
-
|
115 |
-
#### Factors
|
116 |
-
|
117 |
-
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
118 |
-
|
119 |
-
[More Information Needed]
|
120 |
-
|
121 |
-
#### Metrics
|
122 |
-
|
123 |
-
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
124 |
-
|
125 |
-
[More Information Needed]
|
126 |
-
|
127 |
-
### Results
|
128 |
-
|
129 |
-
[More Information Needed]
|
130 |
-
|
131 |
-
#### Summary
|
132 |
-
|
133 |
-
|
134 |
-
|
135 |
-
## Model Examination [optional]
|
136 |
-
|
137 |
-
<!-- Relevant interpretability work for the model goes here -->
|
138 |
-
|
139 |
-
[More Information Needed]
|
140 |
-
|
141 |
-
## Environmental Impact
|
142 |
-
|
143 |
-
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
144 |
-
|
145 |
-
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
146 |
-
|
147 |
-
- **Hardware Type:** [More Information Needed]
|
148 |
-
- **Hours used:** [More Information Needed]
|
149 |
-
- **Cloud Provider:** [More Information Needed]
|
150 |
-
- **Compute Region:** [More Information Needed]
|
151 |
-
- **Carbon Emitted:** [More Information Needed]
|
152 |
-
|
153 |
-
## Technical Specifications [optional]
|
154 |
-
|
155 |
-
### Model Architecture and Objective
|
156 |
-
|
157 |
-
[More Information Needed]
|
158 |
-
|
159 |
-
### Compute Infrastructure
|
160 |
-
|
161 |
-
[More Information Needed]
|
162 |
-
|
163 |
-
#### Hardware
|
164 |
-
|
165 |
-
[More Information Needed]
|
166 |
-
|
167 |
-
#### Software
|
168 |
-
|
169 |
-
[More Information Needed]
|
170 |
-
|
171 |
-
## Citation [optional]
|
172 |
-
|
173 |
-
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
174 |
-
|
175 |
-
**BibTeX:**
|
176 |
-
|
177 |
-
[More Information Needed]
|
178 |
-
|
179 |
-
**APA:**
|
180 |
-
|
181 |
-
[More Information Needed]
|
182 |
-
|
183 |
-
## Glossary [optional]
|
184 |
-
|
185 |
-
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
186 |
-
|
187 |
-
[More Information Needed]
|
188 |
-
|
189 |
-
## More Information [optional]
|
190 |
-
|
191 |
-
[More Information Needed]
|
192 |
-
|
193 |
-
## Model Card Authors [optional]
|
194 |
-
|
195 |
-
[More Information Needed]
|
196 |
-
|
197 |
-
## Model Card Contact
|
198 |
-
|
199 |
-
[More Information Needed]
|
200 |
-
### Framework versions
|
201 |
-
|
202 |
-
- PEFT 0.15.2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
checkpoint-50/adapter_config.json
DELETED
@@ -1,39 +0,0 @@
|
|
1 |
-
{
|
2 |
-
"alpha_pattern": {},
|
3 |
-
"auto_mapping": null,
|
4 |
-
"base_model_name_or_path": "Qwen/Qwen3-8B",
|
5 |
-
"bias": "none",
|
6 |
-
"corda_config": null,
|
7 |
-
"eva_config": null,
|
8 |
-
"exclude_modules": null,
|
9 |
-
"fan_in_fan_out": false,
|
10 |
-
"inference_mode": true,
|
11 |
-
"init_lora_weights": true,
|
12 |
-
"layer_replication": null,
|
13 |
-
"layers_pattern": null,
|
14 |
-
"layers_to_transform": null,
|
15 |
-
"loftq_config": {},
|
16 |
-
"lora_alpha": 16,
|
17 |
-
"lora_bias": false,
|
18 |
-
"lora_dropout": 0.05,
|
19 |
-
"megatron_config": null,
|
20 |
-
"megatron_core": "megatron.core",
|
21 |
-
"modules_to_save": null,
|
22 |
-
"peft_type": "LORA",
|
23 |
-
"r": 8,
|
24 |
-
"rank_pattern": {},
|
25 |
-
"revision": null,
|
26 |
-
"target_modules": [
|
27 |
-
"v_proj",
|
28 |
-
"up_proj",
|
29 |
-
"down_proj",
|
30 |
-
"o_proj",
|
31 |
-
"gate_proj",
|
32 |
-
"k_proj",
|
33 |
-
"q_proj"
|
34 |
-
],
|
35 |
-
"task_type": "CAUSAL_LM",
|
36 |
-
"trainable_token_indices": null,
|
37 |
-
"use_dora": false,
|
38 |
-
"use_rslora": false
|
39 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
checkpoint-50/adapter_model.safetensors
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:79c0f0c6711730284cfc8c2a01daafd1ce46570f5a69dcd429c85d0adc13ac16
|
3 |
-
size 43713984
|
|
|
|
|
|
|
|
checkpoint-50/added_tokens.json
DELETED
@@ -1,28 +0,0 @@
|
|
1 |
-
{
|
2 |
-
"</think>": 151668,
|
3 |
-
"</tool_call>": 151658,
|
4 |
-
"</tool_response>": 151666,
|
5 |
-
"<think>": 151667,
|
6 |
-
"<tool_call>": 151657,
|
7 |
-
"<tool_response>": 151665,
|
8 |
-
"<|box_end|>": 151649,
|
9 |
-
"<|box_start|>": 151648,
|
10 |
-
"<|endoftext|>": 151643,
|
11 |
-
"<|file_sep|>": 151664,
|
12 |
-
"<|fim_middle|>": 151660,
|
13 |
-
"<|fim_pad|>": 151662,
|
14 |
-
"<|fim_prefix|>": 151659,
|
15 |
-
"<|fim_suffix|>": 151661,
|
16 |
-
"<|im_end|>": 151645,
|
17 |
-
"<|im_start|>": 151644,
|
18 |
-
"<|image_pad|>": 151655,
|
19 |
-
"<|object_ref_end|>": 151647,
|
20 |
-
"<|object_ref_start|>": 151646,
|
21 |
-
"<|quad_end|>": 151651,
|
22 |
-
"<|quad_start|>": 151650,
|
23 |
-
"<|repo_name|>": 151663,
|
24 |
-
"<|video_pad|>": 151656,
|
25 |
-
"<|vision_end|>": 151653,
|
26 |
-
"<|vision_pad|>": 151654,
|
27 |
-
"<|vision_start|>": 151652
|
28 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
checkpoint-50/chat_template.jinja
DELETED
@@ -1,89 +0,0 @@
|
|
1 |
-
{%- if tools %}
|
2 |
-
{{- '<|im_start|>system\n' }}
|
3 |
-
{%- if messages[0].role == 'system' %}
|
4 |
-
{{- messages[0].content + '\n\n' }}
|
5 |
-
{%- endif %}
|
6 |
-
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
7 |
-
{%- for tool in tools %}
|
8 |
-
{{- "\n" }}
|
9 |
-
{{- tool | tojson }}
|
10 |
-
{%- endfor %}
|
11 |
-
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
12 |
-
{%- else %}
|
13 |
-
{%- if messages[0].role == 'system' %}
|
14 |
-
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
15 |
-
{%- endif %}
|
16 |
-
{%- endif %}
|
17 |
-
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
18 |
-
{%- for message in messages[::-1] %}
|
19 |
-
{%- set index = (messages|length - 1) - loop.index0 %}
|
20 |
-
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
21 |
-
{%- set ns.multi_step_tool = false %}
|
22 |
-
{%- set ns.last_query_index = index %}
|
23 |
-
{%- endif %}
|
24 |
-
{%- endfor %}
|
25 |
-
{%- for message in messages %}
|
26 |
-
{%- if message.content is string %}
|
27 |
-
{%- set content = message.content %}
|
28 |
-
{%- else %}
|
29 |
-
{%- set content = '' %}
|
30 |
-
{%- endif %}
|
31 |
-
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
32 |
-
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
33 |
-
{%- elif message.role == "assistant" %}
|
34 |
-
{%- set reasoning_content = '' %}
|
35 |
-
{%- if message.reasoning_content is string %}
|
36 |
-
{%- set reasoning_content = message.reasoning_content %}
|
37 |
-
{%- else %}
|
38 |
-
{%- if '</think>' in content %}
|
39 |
-
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
40 |
-
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
41 |
-
{%- endif %}
|
42 |
-
{%- endif %}
|
43 |
-
{%- if loop.index0 > ns.last_query_index %}
|
44 |
-
{%- if loop.last or (not loop.last and reasoning_content) %}
|
45 |
-
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
46 |
-
{%- else %}
|
47 |
-
{{- '<|im_start|>' + message.role + '\n' + content }}
|
48 |
-
{%- endif %}
|
49 |
-
{%- else %}
|
50 |
-
{{- '<|im_start|>' + message.role + '\n' + content }}
|
51 |
-
{%- endif %}
|
52 |
-
{%- if message.tool_calls %}
|
53 |
-
{%- for tool_call in message.tool_calls %}
|
54 |
-
{%- if (loop.first and content) or (not loop.first) %}
|
55 |
-
{{- '\n' }}
|
56 |
-
{%- endif %}
|
57 |
-
{%- if tool_call.function %}
|
58 |
-
{%- set tool_call = tool_call.function %}
|
59 |
-
{%- endif %}
|
60 |
-
{{- '<tool_call>\n{"name": "' }}
|
61 |
-
{{- tool_call.name }}
|
62 |
-
{{- '", "arguments": ' }}
|
63 |
-
{%- if tool_call.arguments is string %}
|
64 |
-
{{- tool_call.arguments }}
|
65 |
-
{%- else %}
|
66 |
-
{{- tool_call.arguments | tojson }}
|
67 |
-
{%- endif %}
|
68 |
-
{{- '}\n</tool_call>' }}
|
69 |
-
{%- endfor %}
|
70 |
-
{%- endif %}
|
71 |
-
{{- '<|im_end|>\n' }}
|
72 |
-
{%- elif message.role == "tool" %}
|
73 |
-
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
74 |
-
{{- '<|im_start|>user' }}
|
75 |
-
{%- endif %}
|
76 |
-
{{- '\n<tool_response>\n' }}
|
77 |
-
{{- content }}
|
78 |
-
{{- '\n</tool_response>' }}
|
79 |
-
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
80 |
-
{{- '<|im_end|>\n' }}
|
81 |
-
{%- endif %}
|
82 |
-
{%- endif %}
|
83 |
-
{%- endfor %}
|
84 |
-
{%- if add_generation_prompt %}
|
85 |
-
{{- '<|im_start|>assistant\n' }}
|
86 |
-
{%- if enable_thinking is defined and enable_thinking is false %}
|
87 |
-
{{- '<think>\n\n</think>\n\n' }}
|
88 |
-
{%- endif %}
|
89 |
-
{%- endif %}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_0_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:a9de36fe4d5bcea143cac9d6c10df60648c387213936d0d79d7fda4d251abc0f
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_1_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:5b0732c53b076adc1100b2a70493aa0c75c1937fbd62dc4ef3c7bc268a1bafd8
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_2_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:5d6a809ba92d4bf901d77ad7d334d988765ba17cdd51cb25b43d59608de4b46d
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_3_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:2a67bff855608b804ad3623e6afe3e413a56e27791a4d443c7a80ea3b0ad0081
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_4_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:352207d4cffd1ec3eac2844b287eb27ff47b0dad349a7c50cb42c8218377915c
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_5_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:4aa163426b44d8986649c2173b1b2407bacc789896ad51dab813fa217a74e4d1
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_6_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:edc251cb7e94e454255d6a8fcc3b93e4ccf96a139b3586e37ed384c6381b26ce
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/bf16_zero_pp_rank_7_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:522cc928155eca7288e4c77eaef9cfe2c051d4dec3180bb8652a8441e18bff5d
|
3 |
-
size 32739589
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_0_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:372befffe03989dbdf4d378ff65b1822c6960e896481e6905bbafe72aba4a148
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_1_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:d55e5f34caa9fffb786489ae7ea7c404bac2759f6ef858ae426bf9f6585d2780
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_2_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:0a5b606de7b5beda42dbd0ca4172e6a7c578eb9c839fc9e31dd28b017c18b09d
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_3_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:fe36de82e6915098c4b20165d6458703640987a0b94b80103e6499047378cdeb
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_4_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:8973e5b60d1f4a8baddb88dd82158c90e377ea341397d8d4c9796f096526b1c2
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_5_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:b5a565b461f68a48a83fcf8b43b62165cb2db9404b7edfecf4ace2310109c5fe
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_6_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:2b1381fdc9460c3f5cbcf6b7fa3e8dbdf572c95ae6749de86b0e782a64448b33
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step44/zero_pp_rank_7_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:13d42064f99b710ebd2d6222050d94593a0d7808fae98085e9263aa20b2b9008
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_0_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:4f4d72a1737209c0b8e72fe13570baaba4a1c3055bb2b05d089e5333fd91d55a
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_1_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:ed2d08ce3db0e511a88c576a99a0d1da12d357c07481a05fd4b451aa8f845340
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_2_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:2be85d7801d57683658e759384e437063da074113cfff86785620a5961c033b3
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_3_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:c695ed726ba293243b46d082ead16d286c1bdcc596e1438f642beb293220ea24
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_4_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:4b78cf28d8c9fb9cefca4520938bb3b2ca5718e95dbce4eda2e4a2c027fd98c9
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_5_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:d04adb6bbdfce6d7a8d9636f41a95839f464274bf555274d95e92b1111544bc2
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_6_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:b6494495fbdfb2ca804879a2511f137d20d9359f90437f193cfc71131c4b3d43
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/bf16_zero_pp_rank_7_mp_rank_00_optim_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:d3eb73c09639ebf76e6d2d4f15073e17264aaff026471f30596106657dedbdbc
|
3 |
-
size 261886213
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_0_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:7c01f8c0fa03f1d7befad2eb63031c91b9f5209cdd72a8c3b84187a007f19288
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_1_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:1407d5600bbb62c8fdee7903421f5f95101723dcf31f6eefc34ade6ad1ca949d
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_2_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:f6c4141b6fea5086e602ee16ca835a4abaaba189dd20ec863001b4363373acae
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_3_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:58b4d6e57ec646377961cfcdc6267fb433d24779950ec15dd43e53efcdf891dc
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_4_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:f43ca1dd28cc70a6711980ca193ac641efbe919aefe298114c8093293b6739db
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_5_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:258cbb29f6c827e929d627d012ced023ca4497f586dd581b394c69022daa0b3c
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_6_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:58c72ad0dee2e25386289b1a498d0efa5ebfa20fb19140f728a6fd1ea2ab30c4
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/global_step50/zero_pp_rank_7_mp_rank_00_model_states.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:b530209e28d101f43f9456112e8f36500558a2c6e653f90f2445772c4c95dc27
|
3 |
-
size 504529
|
|
|
|
|
|
|
|
checkpoint-50/latest
DELETED
@@ -1 +0,0 @@
|
|
1 |
-
global_step44
|
|
|
|
checkpoint-50/merges.txt
DELETED
The diff for this file is too large to render.
See raw diff
|
|
checkpoint-50/rng_state_0.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:8f083214d6e54ffa40ba7c89a0b8ceb60d9a743f4a02eb5fce0f4a9f9b61934b
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_1.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:85b629acd710a60eaf73732dc4205f6b24be2e0fc11b426ee424516aed5f72cd
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_2.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:dfba5cff1fab0d5a4a59e9d26d73b86d8680c089ba3427d3f700fba6cf47f8bd
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_3.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:cd1708a2da57906a954edb626e700e5d0e430574a92e540144aa5fa0dd292254
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_4.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:32378442f326246f7297bea3283fe469a5fd67f2eb6cc5f15960a96d660727a2
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_5.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:3374ddd16a0f504dcf5427e9f025684db5821c8077bfa4cf630e53749193de5b
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_6.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:374a336911b22c295a8e9c974e056d18bc0b56d3e10e25d5ee0476ad78c607da
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/rng_state_7.pth
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:17f3bd6ab1ca7dc6ca2a9ebbdb9b694ea511330e6ebcdc3004c9b216d4ab91ce
|
3 |
-
size 16389
|
|
|
|
|
|
|
|
checkpoint-50/scheduler.pt
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:03c7bda8d6f35b03629d916e81eb3e59c5cf1cef30e8cc48bbdd17025b136e17
|
3 |
-
size 1401
|
|
|
|
|
|
|
|
checkpoint-50/special_tokens_map.json
DELETED
@@ -1,31 +0,0 @@
|
|
1 |
-
{
|
2 |
-
"additional_special_tokens": [
|
3 |
-
"<|im_start|>",
|
4 |
-
"<|im_end|>",
|
5 |
-
"<|object_ref_start|>",
|
6 |
-
"<|object_ref_end|>",
|
7 |
-
"<|box_start|>",
|
8 |
-
"<|box_end|>",
|
9 |
-
"<|quad_start|>",
|
10 |
-
"<|quad_end|>",
|
11 |
-
"<|vision_start|>",
|
12 |
-
"<|vision_end|>",
|
13 |
-
"<|vision_pad|>",
|
14 |
-
"<|image_pad|>",
|
15 |
-
"<|video_pad|>"
|
16 |
-
],
|
17 |
-
"eos_token": {
|
18 |
-
"content": "<|im_end|>",
|
19 |
-
"lstrip": false,
|
20 |
-
"normalized": false,
|
21 |
-
"rstrip": false,
|
22 |
-
"single_word": false
|
23 |
-
},
|
24 |
-
"pad_token": {
|
25 |
-
"content": "<|endoftext|>",
|
26 |
-
"lstrip": false,
|
27 |
-
"normalized": false,
|
28 |
-
"rstrip": false,
|
29 |
-
"single_word": false
|
30 |
-
}
|
31 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
checkpoint-50/tokenizer.json
DELETED
@@ -1,3 +0,0 @@
|
|
1 |
-
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
|
3 |
-
size 11422654
|
|
|
|
|
|
|
|