Upload folder using huggingface_hub

Browse files

Files changed (13) hide show

.gitattributes +1 -0
LICENSE.DeepSeek +21 -0
README.md +152 -0
config.json +122 -0
configuration_deepseek.py +199 -0
intelligence_score_vs_output_tokens.png +3 -0
model.safetensors.index.json +0 -0
nextn_layer_parameters.safetensors +3 -0
tensor_types.json +1 -0
tokenizer.json +0 -0
tokenizer_config.json +35 -0
tool_chat_template.jinja +100 -0
tool_parser_vllm.py +583 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+intelligence_score_vs_output_tokens.png filter=lfs diff=lfs merge=lfs -text

LICENSE.DeepSeek ADDED Viewed

	@@ -0,0 +1,21 @@

+MIT License
+Copyright (c) 2023 DeepSeek
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to deal
+in the Software without restriction, including without limitation the rights
+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+The above copyright notice and this permission notice shall be included in all
+copies or substantial portions of the Software.
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+SOFTWARE.

README.md ADDED Viewed

	@@ -0,0 +1,152 @@

+---
+license: mit
+library_name: transformers
+base_model:
+- deepseek-ai/DeepSeek-R1-0528
+- deepseek-ai/DeepSeek-R1
+- deepseek-ai/DeepSeek-V3-0324
+pipeline_tag: text-generation
+---
+# DeepSeek-TNG-R1T2-Chimera
+<div align="center">
+<img src="https://354918363417-runtime-assets.s3.eu-central-1.amazonaws.com/company_logo_light.svg"
+     alt="TNG Logo"
+     width="400"
+     style="display: inline-block; vertical-align: middle;"/>
+</div>
+<br>
+<div align="center">
+  <a href="https://huggingface.co/tngtech/DeepSeek-TNG-R1T2-Chimera/blob/main/LICENSE.DeepSeek" style="margin: 2px;">
+    <img alt="License" src="https://img.shields.io/badge/License-MIT-f5de53?&color=f5de53" style="display: inline-block; vertical-align: middle;"/>
+  </a>
+</div>
+<br>
+<div align="center">
+    <img alt="Intelligence Score" src="intelligence_score_vs_output_tokens.png" style="display: inline-block; vertical-align: middle;" width="750"/>
+    <figcaption><a href="https://x.com/tngtech/status/1940531045432283412">Release Announcement on X</a></figcaption>
+</div>
+## Assembly of Experts Chimera model constructed with the DeepSeek [R1-0528](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528), [R1](https://huggingface.co/deepseek-ai/DeepSeek-R1) and [V3-0324](https://huggingface.co/deepseek-ai/DeepSeek-V3-0324) parent models
+We present our new **DeepSeek-TNG R1T2 Chimera** 671B model, the first successor to our original [*DeepSeek R1T Chimera*](https://huggingface.co/tngtech/DeepSeek-R1T-Chimera) that was released on April 26th. Unlike the original Chimera, which was based on the *two parent models* V3-0324 and R1, the new Chimera is a **Tri-Mind** *with three parents*, namely additionally R1-0528. It is constructed using the Assembly of Experts-method with relatively fine-granular direct brain edits. This more refined assembly allowed, among other improvements, the fixing of the &lt;think&gt; token consistency issue, which was a weakness of R1T and is now solved for R1T2.
+**Sweet spot**
+R1T2 operates at a new sweet spot in intelligence vs. output token length. It appears to be...
+- about **20% faster than** the regular **R1**, and more than **twice as fast as R1-0528**
+- significantly **more intelligent than** the regular **R1** in benchmarks such as **GPQA**, **AIME-24** and **Aider Polyglot**
+- much **more intelligent** and also **think-token consistent** compared to the first **R1T Chimera** 0426
+- and generally well-behaved and a **nice persona** to talk to, even without any system prompt.
+**Recommendations for your model decision**
+*R1T2* compared...
+- *vs R1:* We hope that R1T2 is a very desirable, almost universally **better drop-in replacement for R1**
+- *vs R1-0528:* R1T2 is a much **cheaper alternative to the full R1-0528**, if the full 0528-level intelligence is not required
+- *vs R1T:* R1T2 is usually **recommended over R1T**, unless the specific personality of R1T was optimal, the think-token issue not important, or R1T's higher speed crucial
+- *vs V3-0324:* V3 is so much faster that if you can live with the **lower intelligence, take V3**, however, if you **need reasoning, R1T2** is the go-to model
+**Limitations**
+- **R1-0528** is thinking much longer, but also is achieving **better hard benchmark results** than R1T2
+- As measured by SpeechMap.ai (courtesy of xlr8harder), **R1T2** is significantly **more reserved** than R1T, but not as much as R1-0528
+- When switching from R1T to R1T2 development, we changed from AIME24 and MT-Bench to AIME24, AIME25 and GPQA-Diamond for the intelligence score. With the new benchmark set, there is a larger score difference between R1 and the original R1T Chimera than published earlier.
+- Function calling is supported in general, but both vLLM and SGLang currently require some specific adaptions, see the section below.
+**Evaluation results**
+Evaluation was performed using the evalchemy framework (pass@1 averaged over 10/5 runs for AIME/GPQAD, at a temperature of 0.6).
+We report measured benchmark results for our R1T2, R1T models and published benchmark results for V3-0324, R1, R1-0528.
+|                                    | R1T2 |  R1T | V3-0324 |   R1 | R1-0528 | Comment | Special source |
+|:-----------------------------------|-----:|-----:|--------:|-----:|--------:|:--------|:--------|
+| AIME-24                            | 82.3 | 74.7 |    59.4 | 79.8 |    91.4 |         |         |
+| AIME-25                            | 70.0 | 58.3 |    49.6 | 70.0 |    87.5 |         | V3-0324 AIME-25 measured by us |
+| GPQA-Diamond                       | 77.9 | 72.0 |    68.4 | 71.5 |    81.0 |         |         |
+| Aider Polyglot                     | 64.4 | 48.4 |    44.9 | 52.0 |    71.6 | R1T2 beats two of its parents, V3-0324 and R1, and was measured to be about 2.2 times more token efficient, i.e. faster, than its third parent, R1-0528 | R1T2 source: Aider discord, t=0.75 |
+| MMLU-Pro Computer Science          | 83.7-85.6 | 82.9-84.6 | 81.5-82.4 | 85.1-85.3 | 84.6-86.1 |         |         |
+| EQ-Bench Longform Creative Writing | 76.4 |  ./. |    78.1 | 74.6 |    78.9 | EQ Bench version before August 8th, 2025 | see [EQ Bench](https://eqbench.com/creative_writing_longform.html)  |
+| Vectara Hallucination Rate         |  5.5 |  ./. |     8.0 | 14.3 |     7.7 | lower hallucination rates are better, R1T2 is better than all its three parents | see [Hallucination Leaderboard](https://github.com/vectara/hallucination-leaderboard) |
+## Technological background
+For details on the AoE construction process, you can read our [Paper on arXiV](https://arxiv.org/abs/2506.14794).
+**Runtime parameter settings**
+- Most of our evaluation was done with a maximum context size of 60,000 tokens.
+  With a context size of 130,000 tokens, the model proved very helpful in interpreting very long debug logs. Long-context testing was less extensive, though.
+- We're running the model using vLLM on 8xH200 and MI325X nodes, additionally we've tested the model using SGLang, which is also used by [chutes.ai](https://chutes.ai/app/chute/4fa0c7f5-82f7-59d1-8996-661bb778893d).
+- For SGLang, we recommend using versions >= v0.4.8 in combination with argument `--reasoning-parser qwen3` to properly handle rare cases when the model skips the `<think>` reasoning step.
+### Function calling
+R1T2 does support function calling using an updated chat template (since 01 Aug 2025). However, neither vLLM nor SGLang provide an R1T2-compatible tool call parser natively but require some adaptions.
+_vLLM:_
+For function calling with vLLM, a new tool parser is required. While we opened [a PR to vLLM](https://github.com/vllm-project/vllm/pull/22074) to include an R1T2-compatible tool parser off-the-shelf, we also ship the tool parser file `tool_parser_vllm.py` within this repository.
+With this file, tool calling can be enabled via
+```
+--tool-parser-plugin <ABSOLUTE_MODEL_SNAPSHOT_PATH>/tool_parser_vllm.py  \
+--tool-call-parser tng_r1t2
+```
+Here, put in the path to the snapshot folder such as `~/.cache/huggingface/hub/models--tngtech--DeepSeek-TNG-R1T2-Chimera/snapshots/SNAPSHOT/tool_parser_vllm.py`
+_SGLang:_
+Tool call support for R1T2 requires a recent SGLang version >= v0.4.10 (alternatively, you need to patch [this bugfix for the reasoning parser](https://github.com/sgl-project/sglang/pull/8606) for older versions of SGLang).
+An R1T2-compatible tool call parser will be added with [this PR to SGLang](https://github.com/sgl-project/sglang/pull/8672).
+Unfortunately, and unlike vLLM, there is no simple plugin system for tool call parsers in SGLang.
+Until our PR is merged an relased with a new SGLang version, you can still install it manually by patching your SGLang source code as outlined in the PR:
+The new tool call parser must be added and registered (so in total one file must be added, a second one edited, see [details here](https://github.com/sgl-project/sglang/pull/8672/files)).
+Once the SGLang installation has been updated correctly, tool calling with R1T2 can be activated by starting SGLang with
+```
+--tool-call-parser tng_r1t2
+```
+## Model Details
+- **Architecture**: DeepSeek-MoE transformer-based language model
+- **Combination Method**: Assembly of Experts from the three DeepSeek parent models R1-0528, R1 and V3-0324
+- **Release Date**: 2025-07-02
+- **Design Team**: Robert Dahlke, Henrik Klagges, Benjamin Merkel, Fabian Klemm and David Reiss, Munich, Germany
+- **Extra Thanks**: Big thanks to DeepSeek for their great models and open-source generosity, and to the other researchers that have published on model merging methodologies.
+## Use, Out-of-scope Use, Other Limitations, Risks, Recommendations et al.
+Regarding the R1T/R1T2-Chimeras, we ask you to follow the careful guidelines that Microsoft has created for their "MAI-DS-R1" DeepSeek-based model.
+These professional guidelines are available [here on Hugging Face](https://huggingface.co/microsoft/MAI-DS-R1).
+## EU AI Act
+Due to the strict new guidelines of the EU AI Act that take effect on August 2nd 2025, we recommend that each R1T/R1T2 user in the EU either familiarizes themselves with these requirements and assess their compliance, or ceases using the model in the EU after August 1st, 2025.
+## Contact, especially for your user feedback
+Please give us your feedback, especially if you find deficiencies in the model:
+- Email: [email protected]
+- X.com: @tngtech
+## Citation
+```
+@misc{tng_technology_consulting_gmbh_2025_07_02,
+	author       = { TNG Technology Consulting GmbH },
+	title        = { DeepSeek-TNG-R1T2-Chimera },
+	year         = 2025,
+    month        = { July },
+	url          = { https://huggingface.co/tngtech/DeepSeek-TNG-R1T2-Chimera },
+	doi          = { 10.57967/hf/5950 },
+	publisher    = { Hugging Face }
+}
+```

config.json ADDED Viewed

	@@ -0,0 +1,122 @@

+{
+  "_name_or_path": "/data/hf/hub/models--tngtech--DeepSeek-TNG-R1T2-Chimera/snapshots/c0f05b5836356f60dee8ac364394ceb28215e7ba",
+  "add_cross_attention": false,
+  "architectures": [
+    "DeepseekV3ForCausalLMNextN"
+  ],
+  "attention_bias": false,
+  "attention_dropout": 0.0,
+  "auto_map": {
+    "AutoConfig": "configuration_deepseek.DeepseekV3Config",
+    "AutoModel": "modeling_deepseek.DeepseekV3Model",
+    "AutoModelForCausalLM": "modeling_deepseek.DeepseekV3ForCausalLM"
+  },
+  "bad_words_ids": null,
+  "begin_suppress_tokens": null,
+  "bos_token_id": 0,
+  "chunk_size_feed_forward": 0,
+  "cross_attention_hidden_size": null,
+  "decoder_start_token_id": null,
+  "diversity_penalty": 0.0,
+  "do_sample": false,
+  "dtype": "bfloat16",
+  "early_stopping": false,
+  "encoder_no_repeat_ngram_size": 0,
+  "eos_token_id": 1,
+  "ep_size": 1,
+  "exponential_decay_length_penalty": null,
+  "finetuning_task": null,
+  "first_k_dense_replace": 3,
+  "forced_bos_token_id": null,
+  "forced_eos_token_id": null,
+  "hidden_act": "silu",
+  "hidden_size": 7168,
+  "id2label": {
+    "0": "LABEL_0",
+    "1": "LABEL_1"
+  },
+  "initializer_range": 0.02,
+  "intermediate_size": 18432,
+  "is_decoder": false,
+  "is_encoder_decoder": false,
+  "kv_lora_rank": 512,
+  "label2id": {
+    "LABEL_0": 0,
+    "LABEL_1": 1
+  },
+  "length_penalty": 1.0,
+  "max_length": 20,
+  "max_position_embeddings": 163840,
+  "min_length": 0,
+  "model_type": "deepseek_v3",
+  "moe_intermediate_size": 2048,
+  "moe_layer_freq": 1,
+  "n_group": 8,
+  "n_routed_experts": 256,
+  "n_shared_experts": 1,
+  "no_repeat_ngram_size": 0,
+  "norm_topk_prob": true,
+  "num_attention_heads": 128,
+  "num_beam_groups": 1,
+  "num_beams": 1,
+  "num_experts_per_tok": 8,
+  "num_hidden_layers": 1,
+  "num_key_value_heads": 128,
+  "num_nextn_predict_layers": 1,
+  "num_return_sequences": 1,
+  "output_attentions": false,
+  "output_hidden_states": false,
+  "output_scores": false,
+  "pad_token_id": null,
+  "prefix": null,
+  "problem_type": null,
+  "pruned_heads": {},
+  "q_lora_rank": 1536,
+  "qk_nope_head_dim": 128,
+  "qk_rope_head_dim": 64,
+  "quantization_config": {
+    "activation_scheme": "dynamic",
+    "fmt": "e4m3",
+    "quant_method": "fp8",
+    "weight_block_size": [
+      128,
+      128
+    ]
+  },
+  "remove_invalid_values": false,
+  "repetition_penalty": 1.0,
+  "return_dict": true,
+  "return_dict_in_generate": false,
+  "rms_norm_eps": 1e-06,
+  "rope_scaling": {
+    "beta_fast": 32,
+    "beta_slow": 1,
+    "factor": 40,
+    "mscale": 1.0,
+    "mscale_all_dim": 1.0,
+    "original_max_position_embeddings": 4096,
+    "type": "yarn"
+  },
+  "rope_theta": 10000,
+  "routed_scaling_factor": 2.5,
+  "scoring_func": "sigmoid",
+  "sep_token_id": null,
+  "suppress_tokens": null,
+  "task_specific_params": null,
+  "temperature": 1.0,
+  "tf_legacy_loss": false,
+  "tie_encoder_decoder": false,
+  "tie_word_embeddings": false,
+  "tokenizer_class": null,
+  "top_k": 50,
+  "top_p": 1.0,
+  "topk_group": 4,
+  "topk_method": "noaux_tc",
+  "torchscript": false,
+  "transformers_version": "4.56.1",
+  "typical_p": 1.0,
+  "use_bfloat16": false,
+  "use_cache": true,
+  "v_head_dim": 128,
+  "vocab_size": 129280
+}

configuration_deepseek.py ADDED Viewed

	@@ -0,0 +1,199 @@

+from transformers.configuration_utils import PretrainedConfig
+from transformers.utils import logging
+logger = logging.get_logger(__name__)
+DEEPSEEK_PRETRAINED_CONFIG_ARCHIVE_MAP = {}
+class DeepseekV3Config(PretrainedConfig):
+    r"""
+    This is the configuration class to store the configuration of a [`DeepseekV3Model`]. It is used to instantiate an DeepSeek
+    model according to the specified arguments, defining the model architecture. Instantiating a configuration with the
+    defaults will yield a similar configuration to that of the DeepSeek-V3.
+    Configuration objects inherit from [`PretrainedConfig`] and can be used to control the model outputs. Read the
+    documentation from [`PretrainedConfig`] for more information.
+    Args:
+        vocab_size (`int`, *optional*, defaults to 129280):
+            Vocabulary size of the Deep model. Defines the number of different tokens that can be represented by the
+            `inputs_ids` passed when calling [`DeepseekV3Model`]
+        hidden_size (`int`, *optional*, defaults to 4096):
+            Dimension of the hidden representations.
+        intermediate_size (`int`, *optional*, defaults to 11008):
+            Dimension of the MLP representations.
+        moe_intermediate_size (`int`, *optional*, defaults to 1407):
+            Dimension of the MoE representations.
+        num_hidden_layers (`int`, *optional*, defaults to 32):
+            Number of hidden layers in the Transformer decoder.
+        num_nextn_predict_layers (`int`, *optional*, defaults to 1):
+            Number of nextn predict layers in the DeepSeekV3 Model.
+        num_attention_heads (`int`, *optional*, defaults to 32):
+            Number of attention heads for each attention layer in the Transformer decoder.
+        n_shared_experts (`int`, *optional*, defaults to None):
+            Number of shared experts, None means dense model.
+        n_routed_experts (`int`, *optional*, defaults to None):
+            Number of routed experts, None means dense model.
+        routed_scaling_factor (`float`, *optional*, defaults to 1.0):
+            Scaling factor or routed experts.
+        topk_method (`str`, *optional*, defaults to `gready`):
+            Topk method used in routed gate.
+        n_group (`int`, *optional*, defaults to None):
+            Number of groups for routed experts.
+        topk_group (`int`, *optional*, defaults to None):
+            Number of selected groups for each token(for each token, ensuring the selected experts is only within `topk_group` groups).
+        num_experts_per_tok (`int`, *optional*, defaults to None):
+            Number of selected experts, None means dense model.
+        moe_layer_freq (`int`, *optional*, defaults to 1):
+            The frequency of the MoE layer: one expert layer for every `moe_layer_freq - 1` dense layers.
+        first_k_dense_replace (`int`, *optional*, defaults to 0):
+            Number of dense layers in shallow layers(embed->dense->dense->...->dense->moe->moe...->lm_head).
+                                                            \--k dense layers--/
+        norm_topk_prob (`bool`, *optional*, defaults to False):
+            Whether to normalize the weights of the routed experts.
+        scoring_func (`str`, *optional*, defaults to 'softmax'):
+            Method of computing expert weights.
+        aux_loss_alpha (`float`, *optional*, defaults to 0.001):
+            Auxiliary loss weight coefficient.
+        seq_aux = (`bool`, *optional*, defaults to True):
+            Whether to compute the auxiliary loss for each individual sample.
+        num_key_value_heads (`int`, *optional*):
+            This is the number of key_value heads that should be used to implement Grouped Query Attention. If
+            `num_key_value_heads=num_attention_heads`, the model will use Multi Head Attention (MHA), if
+            `num_key_value_heads=1 the model will use Multi Query Attention (MQA) otherwise GQA is used. When
+            converting a multi-head checkpoint to a GQA checkpoint, each group key and value head should be constructed
+            by meanpooling all the original heads within that group. For more details checkout [this
+            paper](https://arxiv.org/pdf/2305.13245.pdf). If it is not specified, will default to
+            `num_attention_heads`.
+        hidden_act (`str` or `function`, *optional*, defaults to `"silu"`):
+            The non-linear activation function (function or string) in the decoder.
+        max_position_embeddings (`int`, *optional*, defaults to 2048):
+            The maximum sequence length that this model might ever be used with.
+        initializer_range (`float`, *optional*, defaults to 0.02):
+            The standard deviation of the truncated_normal_initializer for initializing all weight matrices.
+        rms_norm_eps (`float`, *optional*, defaults to 1e-06):
+            The epsilon used by the rms normalization layers.
+        use_cache (`bool`, *optional*, defaults to `True`):
+            Whether or not the model should return the last key/values attentions (not used by all models). Only
+            relevant if `config.is_decoder=True`.
+        pad_token_id (`int`, *optional*):
+            Padding token id.
+        bos_token_id (`int`, *optional*, defaults to 1):
+            Beginning of stream token id.
+        eos_token_id (`int`, *optional*, defaults to 2):
+            End of stream token id.
+        tie_word_embeddings (`bool`, *optional*, defaults to `False`):
+            Whether to tie weight embeddings
+        rope_theta (`float`, *optional*, defaults to 10000.0):
+            The base period of the RoPE embeddings.
+        rope_scaling (`Dict`, *optional*):
+            Dictionary containing the scaling configuration for the RoPE embeddings. Currently supports two scaling
+            strategies: linear and dynamic. Their scaling factor must be a float greater than 1. The expected format is
+            `{"type": strategy name, "factor": scaling factor}`. When using this flag, don't update
+            `max_position_embeddings` to the expected new maximum.
+        attention_bias (`bool`, defaults to `False`, *optional*, defaults to `False`):
+            Whether to use a bias in the query, key, value and output projection layers during self-attention.
+        attention_dropout (`float`, *optional*, defaults to 0.0):
+            The dropout ratio for the attention probabilities.
+    ```python
+    >>> from transformers import DeepseekV3Model, DeepseekV3Config
+    >>> # Initializing a Deepseek-V3 style configuration
+    >>> configuration = DeepseekV3Config()
+    >>> # Accessing the model configuration
+    >>> configuration = model.config
+    ```"""
+    model_type = "deepseek_v3"
+    keys_to_ignore_at_inference = ["past_key_values"]
+    def __init__(
+        self,
+        vocab_size=129280,
+        hidden_size=7168,
+        intermediate_size=18432,
+        moe_intermediate_size = 2048,
+        num_hidden_layers=61,
+        num_nextn_predict_layers=1,
+        num_attention_heads=128,
+        num_key_value_heads=128,
+        n_shared_experts = 1,
+        n_routed_experts = 256,
+        ep_size = 1,
+        routed_scaling_factor = 2.5,
+        kv_lora_rank = 512,
+        q_lora_rank = 1536,
+        qk_rope_head_dim = 64,
+        v_head_dim = 128,
+        qk_nope_head_dim = 128,
+        topk_method = 'noaux_tc',
+        n_group = 8,
+        topk_group = 4,
+        num_experts_per_tok = 8,
+        moe_layer_freq = 1,
+        first_k_dense_replace = 3,
+        norm_topk_prob = True,
+        scoring_func = 'sigmoid',
+        hidden_act="silu",
+        max_position_embeddings=4096,
+        initializer_range=0.02,
+        rms_norm_eps=1e-6,
+        use_cache=True,
+        pad_token_id=None,
+        bos_token_id=0,
+        eos_token_id=1,
+        tie_word_embeddings=False,
+        rope_theta=10000.0,
+        rope_scaling=None,
+        attention_bias=False,
+        attention_dropout=0.0,
+        **kwargs,
+    ):
+        self.vocab_size = vocab_size
+        self.max_position_embeddings = max_position_embeddings
+        self.hidden_size = hidden_size
+        self.intermediate_size = intermediate_size
+        self.moe_intermediate_size = moe_intermediate_size
+        self.num_hidden_layers = num_hidden_layers
+        self.num_nextn_predict_layers = num_nextn_predict_layers
+        self.num_attention_heads = num_attention_heads
+        self.n_shared_experts = n_shared_experts
+        self.n_routed_experts = n_routed_experts
+        self.ep_size = ep_size
+        self.routed_scaling_factor = routed_scaling_factor
+        self.kv_lora_rank = kv_lora_rank
+        self.q_lora_rank = q_lora_rank
+        self.qk_rope_head_dim = qk_rope_head_dim
+        self.v_head_dim = v_head_dim
+        self.qk_nope_head_dim = qk_nope_head_dim
+        self.topk_method = topk_method
+        self.n_group = n_group
+        self.topk_group = topk_group
+        self.num_experts_per_tok = num_experts_per_tok
+        self.moe_layer_freq = moe_layer_freq
+        self.first_k_dense_replace = first_k_dense_replace
+        self.norm_topk_prob = norm_topk_prob
+        self.scoring_func = scoring_func
+        # for backward compatibility
+        if num_key_value_heads is None:
+            num_key_value_heads = num_attention_heads
+        self.num_key_value_heads = num_key_value_heads
+        self.hidden_act = hidden_act
+        self.initializer_range = initializer_range
+        self.rms_norm_eps = rms_norm_eps
+        self.use_cache = use_cache
+        self.rope_theta = rope_theta
+        self.rope_scaling = rope_scaling
+        self.attention_bias = attention_bias
+        self.attention_dropout = attention_dropout
+        super().__init__(
+            pad_token_id=pad_token_id,
+            bos_token_id=bos_token_id,
+            eos_token_id=eos_token_id,
+            tie_word_embeddings=tie_word_embeddings,
+            **kwargs,
+        )

intelligence_score_vs_output_tokens.png ADDED Viewed

Git LFS Details

SHA256: ace1e8df27abccaf153f01b719117cbc024839c02cab6e2a300aa401ba196af7
Pointer size: 131 Bytes
Size of remote file: 197 kB

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff

nextn_layer_parameters.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:22b0f714c7fd5bfc0b7c10761c73f3dff0e2d7c6a383b784a7a9cad4c51e2012
+size 11717707368

tensor_types.json ADDED Viewed

	@@ -0,0 +1 @@


1	+ {}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,35 @@

+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "bos_token": {
+    "__type": "AddedToken",
+    "content": "<｜begin▁of▁sentence｜>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "clean_up_tokenization_spaces": false,
+  "eos_token": {
+    "__type": "AddedToken",
+    "content": "<｜end▁of▁sentence｜>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "legacy": true,
+  "model_max_length": 131072,
+  "pad_token": {
+    "__type": "AddedToken",
+    "content": "<｜end▁of▁sentence｜>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "sp_model_kwargs": {},
+  "unk_token": null,
+  "tokenizer_class": "LlamaTokenizerFast",
+  "chat_template": "{% if not add_generation_prompt is defined %}{% set add_generation_prompt = false %}{% endif %}{% set ns = namespace(is_first=false, is_tool=false, is_output_first=true, system_prompt='', is_first_sp=true, is_last_user=false) %}{%- for message in messages %}{%- if message['role'] == 'system' %}{%- if ns.is_first_sp %}{% set ns.system_prompt = ns.system_prompt + message['content'] %}{% set ns.is_first_sp = false %}{%- else %}{% set ns.system_prompt = ns.system_prompt + '\\n\\n' + message['content'] %}{%- endif %}{%- endif %}{%- endfor -%}{#- Adapted from https://github.com/sgl-project/sglang/blob/main/examples/chat_template/tool_chat_template_deepseekr1.jinja #}{% if tools is defined and tools is not none %}{% set tool_ns = namespace(text='You are a helpful assistant with tool calling capabilities. When a tool call is needed, you MUST use the following format to issue the call:\\n<tool_call>\\n{\"name\": FUNCTION_NAME, \"arguments\": {\"param1\": \"value1\", \"param2\": \"value2\"}}\\n</tool_call>\\n\\nMake sure the JSON is valid.\\n\\n## Tools\\n\\n### Function\\n\\nYou have the following functions available:\\n\\n') %}{% for tool in tools %}{% set tool_ns.text = tool_ns.text + '\\n- ' + tool['function']['name'] + '\\n```json\\n' + (tool['function'] | tojson) + '\\n```\\n' %}{% endfor %}{% set ns.system_prompt = ns.system_prompt + '\\n\\n' + tool_ns.text %}{% endif %}{{- bos_token }}{{- ns.system_prompt }}{%- for message in messages %}{% set content = message['content'] %}{%- if message['role'] == 'user' %}{%- set ns.is_tool = false -%}{%- set ns.is_first = false -%}{%- set ns.is_last_user = true -%}{{'<｜User｜>' + content + '<｜Assistant｜>'}}{%- endif %}{%- if message['role'] == 'assistant' %}{% if '</think>' in content %}{% set content = content.split('</think>')[-1] %}{% endif %}{% endif %}{%- if message['role'] == 'assistant' and message['tool_calls'] is defined and message['tool_calls'] is not none %}{%- set ns.is_last_user = false -%}{%- if ns.is_tool %}{{- ''}}{%- endif %}{%- set ns.is_first = false %}{%- set ns.is_tool = false -%}{%- set ns.is_output_first = true %}{%- for tool in message['tool_calls'] %}{%- if tool['function']['arguments'] is string %}{%- set arguments = tool['function']['arguments'] %}{%- else %}{%- set arguments = tool['function']['arguments'] | tojson %}{%- endif %}{%- if not ns.is_first %}{%- if content is none %}{{- '<tool_call>\\n{\"tool_call_id\": ' + tool['id']|tojson + ', \"name\": ' + tool['function']['name']|tojson + ', \"arguments\": ' + arguments + '}\\n</tool_call>'}}{%- else %}{{- content + '\\n\\n<tool_call>\\n{\"tool_call_id\": ' + tool['id']|tojson + ', \"name\": ' + tool['function']['name']|tojson + ', \"arguments\": ' + arguments + '}\\n</tool_call>'}}{%- endif %}{%- set ns.is_first = true -%}{%- else %}{{- '\\n\\n<tool_call>\\n{\"tool_call_id\": ' + tool['id']|tojson + ', \"name\": ' + tool['function']['name']|tojson + ', \"arguments\": ' + arguments + '}\\n</tool_call>'}}{%- endif %}{%- endfor %}{{- '<｜end▁of▁sentence｜>'}}{%- endif %}{%- if message['role'] == 'assistant' and (message['tool_calls'] is not defined or message['tool_calls'] is none)%}{%- set ns.is_last_user = false -%}{%- if ns.is_tool %}{{- '\\n' + content + '<｜end▁of▁sentence｜>'}}{%- set ns.is_tool = false -%}{%- else %}{{- content + '<｜end▁of▁sentence｜>'}}{%- endif %}{%- endif %}{%- if message['role'] == 'tool' %}{%- set ns.is_last_user = false -%}{%- set ns.is_tool = true -%}{%- set tool_call_id_param = '' %}{%- if message['tool_call_id'] %}{%- set tool_call_id_param = '\"tool_call_id\": ' + message['tool_call_id']|tojson + ', ' %}{%- endif %}{%- if ns.is_output_first %}{{- '\\n\\n{' + tool_call_id_param + '\"content\": ' + content|tojson + '}'}}{%- set ns.is_output_first = false %}{%- else %}{{- '\\n{' + tool_call_id_param + '\"content\": ' + content|tojson + '}'}}{%- endif %}{%- endif %}{%- endfor -%}{% if ns.is_tool %}{{- '\\n\\n'}}{%- endif %}{% if add_generation_prompt and not ns.is_last_user %}{{- '<｜Assistant｜>'}}{%- endif %}"
+}

tool_chat_template.jinja ADDED Viewed

	@@ -0,0 +1,100 @@

+{% if not add_generation_prompt is defined %}
+    {% set add_generation_prompt = false %}
+{% endif %}
+{% set ns = namespace(is_first=false, is_tool=false, is_output_first=true, system_prompt='', is_first_sp=true, is_last_user=false) %}
+{%- for message in messages %}
+    {%- if message['role'] == 'system' %}
+        {%- if ns.is_first_sp %}
+            {% set ns.system_prompt = ns.system_prompt + message['content'] %}
+            {% set ns.is_first_sp = false %}
+        {%- else %}
+            {% set ns.system_prompt = ns.system_prompt + '\n\n' + message['content'] %}
+        {%- endif %}
+    {%- endif %}
+{%- endfor -%}
+{#- Adapted from https://github.com/sgl-project/sglang/blob/main/examples/chat_template/tool_chat_template_deepseekr1.jinja #}
+{% if tools is defined and tools is not none %}
+    {% set tool_ns = namespace(text='You are a helpful assistant with tool calling capabilities. '
+        'When a tool call is needed, you MUST use the following format to issue the call:\n'
+        '<tool_call>\n{"name": FUNCTION_NAME, "arguments": {"param1": "value1", "param2": "value2"}}\n</tool_call>\n\n'
+        'Make sure the JSON is valid.\n\n'
+        '## Tools\n\n### Function\n\nYou have the following functions available:\n\n') %}
+    {% for tool in tools %}
+        {% set tool_ns.text = tool_ns.text + '\n- ' + tool['function']['name'] + '\n```json\n' + (tool['function'] | tojson) + '\n```\n' %}
+    {% endfor %}
+    {% set ns.system_prompt = ns.system_prompt + '\n\n' + tool_ns.text %}
+{% endif %}
+{{- bos_token }}
+{{- ns.system_prompt }}
+{%- for message in messages %}
+    {% set content = message['content'] %}
+    {%- if message['role'] == 'user' %}
+        {%- set ns.is_tool = false -%}
+        {%- set ns.is_first = false -%}
+        {%- set ns.is_last_user = true -%}
+        {{'<｜User｜>' + content + '<｜Assistant｜>'}}
+    {%- endif %}
+    {%- if message['role'] == 'assistant' %}
+        {% if '</think>' in content %}
+            {% set content = content.split('</think>')[-1] %}
+        {% endif %}
+    {% endif %}
+    {%- if message['role'] == 'assistant' and message['tool_calls'] is defined and message['tool_calls'] is not none %}
+        {%- set ns.is_last_user = false -%}
+        {%- if ns.is_tool %}
+            {{- ''}}
+        {%- endif %}
+        {%- set ns.is_first = false %}
+        {%- set ns.is_tool = false -%}
+        {%- set ns.is_output_first = true %}
+        {%- for tool in message['tool_calls'] %}
+            {%- if tool['function']['arguments'] is string %}
+                {%- set arguments = tool['function']['arguments'] %}
+            {%- else %}
+                {%- set arguments = tool['function']['arguments'] | tojson %}
+            {%- endif %}
+            {%- if not ns.is_first %}
+                {%- if content is none %}
+                    {{- '<tool_call>\n{"tool_call_id": ' + tool['id']|tojson + ', "name": ' + tool['function']['name']|tojson + ', "arguments": ' + arguments + '}\n</tool_call>'}}
+                {%- else %}
+                    {{- content + '\n\n<tool_call>\n{"tool_call_id": ' + tool['id']|tojson + ', "name": ' + tool['function']['name']|tojson + ', "arguments": ' + arguments + '}\n</tool_call>'}}
+                {%- endif %}
+                {%- set ns.is_first = true -%}
+            {%- else %}
+                {{- '\n\n<tool_call>\n{"tool_call_id": ' + tool['id']|tojson + ', "name": ' + tool['function']['name']|tojson + ', "arguments": ' + arguments + '}\n</tool_call>'}}
+            {%- endif %}
+        {%- endfor %}
+        {{- '<｜end▁of▁sentence｜>'}}
+    {%- endif %}
+    {%- if message['role'] == 'assistant' and (message['tool_calls'] is not defined or message['tool_calls'] is none)%}
+        {%- set ns.is_last_user = false -%}
+        {%- if ns.is_tool %}
+            {{- '\n' + content + '<｜end▁of▁sentence｜>'}}
+            {%- set ns.is_tool = false -%}
+        {%- else %}
+            {{- content + '<｜end▁of▁sentence｜>'}}
+        {%- endif %}
+    {%- endif %}
+    {%- if message['role'] == 'tool' %}
+        {%- set ns.is_last_user = false -%}
+        {%- set ns.is_tool = true -%}
+        {%- set tool_call_id_param = '' %}
+        {%- if message['tool_call_id'] %}
+            {%- set tool_call_id_param = '"tool_call_id": ' + message['tool_call_id']|tojson + ', ' %}
+        {%- endif %}
+        {%- if ns.is_output_first %}
+            {{- '\n\n{' + tool_call_id_param + '"content": ' + content|tojson + '}'}}
+            {%- set ns.is_output_first = false %}
+        {%- else %}
+            {{- '\n{' + tool_call_id_param + '"content": ' + content|tojson + '}'}}
+        {%- endif %}
+    {%- endif %}
+{%- endfor -%}
+{% if ns.is_tool %}
+    {{- '\n\n'}}
+{%- endif %}
+{% if add_generation_prompt and not ns.is_last_user %}
+    {{- '<｜Assistant｜>'}}
+{%- endif %}

tool_parser_vllm.py ADDED Viewed

	@@ -0,0 +1,583 @@

+# SPDX-License-Identifier: Apache-2.0
+# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
+# ruff: noqa
+import json
+from collections.abc import Sequence
+from enum import Enum
+from typing import Any, Union
+import partial_json_parser
+import regex as re
+from partial_json_parser.core.options import Allow
+from vllm.entrypoints.openai.protocol import (ChatCompletionRequest,
+                                              DeltaFunctionCall, DeltaMessage,
+                                              DeltaToolCall,
+                                              ExtractedToolCallInformation,
+                                              FunctionCall, ToolCall)
+from vllm.entrypoints.openai.tool_parsers.abstract_tool_parser import (
+    ToolParser, ToolParserManager)
+from vllm.logger import init_logger
+from vllm.transformers_utils.tokenizer import AnyTokenizer
+from vllm.utils import random_uuid
+logger = init_logger(__name__)
+class ParsedStructure(Enum):
+    CONTENT = 1
+    REASONING_CONTENT = 2
+    TOOL_CALL = 3
+    TOOL_CALL_DELIMITER = 4
+    TOOL_CALL_START_TAG = 5
+    TOOL_CALL_END_TAG = 6
+@ToolParserManager.register_module("tng_r1t2")
+class TngR1T2ToolParser(ToolParser):
+    """Tool Parser for models like tngtech/DeepSeek-TNG-R1T2-Chimera,
+    It is compatible with hermes tool call templates
+    but does not require <tool_call> and </tool_call>
+    to be single tokens in the vocabulary;
+    instead only the string representation of the model output
+    is parsed, making this tool call parser robust and versatile."""
+    def __init__(self, tokenizer: AnyTokenizer):
+        super().__init__(tokenizer)
+        # For backward compatibility with serving code
+        self.prev_tool_call_arr: list[dict] = []
+        self.streamed_args_for_tool: list[str] = []
+        self.think_tag_pattern = r"(<think>[\s\S]*?</think>)"
+        self.think_start_tag = "<think>"
+        self.think_end_tag = "</think>"
+        self.tool_call_tag_pattern = r"<tool_call>([\s\S]*?)</tool_call>"
+        self.tool_call_start_tag = "<tool_call>"
+        self.tool_call_end_tag = "</tool_call>"
+        # Define streaming state type to be initialized later
+        self.streaming_state: dict[str, Any] = {
+            "streamed_tool_calls": [],
+            "buffer": "",
+            "parsed_structure": ParsedStructure.CONTENT,
+        }
+    def extract_tool_call_from_nonthink_output(
+            self, raw_text: str) -> tuple[str, list[dict]]:
+        parts = re.split(self.tool_call_tag_pattern, raw_text)
+        content = ""
+        tool_calls: list[dict] = []
+        for i, part in enumerate(parts):
+            is_potential_tool_call = i % 2 == 1
+            if is_potential_tool_call:
+                try:
+                    more_tool_calls = json.loads(part)
+                    if isinstance(more_tool_calls, list):
+                        tool_calls.extend(more_tool_calls)
+                    else:
+                        tool_calls.extend([more_tool_calls])
+                except json.JSONDecodeError:
+                    logger.warning("Invalid tool call json "
+                                   "-> parse as text content")
+                    content += part
+                    continue
+            else:
+                content += part
+        return content, tool_calls
+    def extract_tool_calls(
+        self,
+        model_output: str,
+        request: ChatCompletionRequest,
+    ) -> ExtractedToolCallInformation:
+        """
+        Extract tool calls from a complete model output.
+        """
+        # split at think traces -> those will not be parsed for tool calls
+        think_parts = re.split(self.think_tag_pattern, model_output)
+        content = ""
+        tool_calls = []
+        for i, part in enumerate(think_parts):
+            parse_output = i % 2 == 0
+            if parse_output:
+                more_content, more_tool_calls = (
+                    self.extract_tool_call_from_nonthink_output(part))
+                content += more_content
+                tool_calls += more_tool_calls
+            else:
+                content += part
+        if not tool_calls:
+            return ExtractedToolCallInformation(
+                tools_called=False,
+                tool_calls=[],
+                content=content,
+            )
+        tool_call_objs: list[ToolCall] = []
+        for idx, call in enumerate(tool_calls):
+            if (not isinstance(call, dict) or "name" not in call
+                    or "arguments" not in call):
+                logger.warning("Invalid tool call format, ignore.")
+                continue
+            tool_call = ToolCall(
+                id=f"call_{idx}_{random_uuid()}",
+                type="function",
+                function=FunctionCall(
+                    name=call["name"],
+                    arguments=(json.dumps(call["arguments"]) if isinstance(
+                        call["arguments"], dict) else call["arguments"]),
+                ),
+            )
+            tool_call_objs.append(tool_call)
+        return ExtractedToolCallInformation(
+            tools_called=len(tool_call_objs) > 0,
+            tool_calls=tool_call_objs,
+            content=content,
+        )
+    def _parse_think_trace(self, raw_text: str) -> tuple[str, bool, str]:
+        """
+        Returns: (unambiguous_text_content, found_think_end, rest_string)
+        """
+        # Either a complete think_end_tag can be somewhere in raw_text
+        think_end_pos = raw_text.find(self.think_end_tag)
+        if think_end_pos >= 0:
+            # in contrast to tool_call_start_tags, </think> remains part of content
+            think_end_tag_end_pos = think_end_pos + len(self.think_end_tag)
+            return (raw_text[:think_end_tag_end_pos], True,
+                    raw_text[think_end_tag_end_pos:])
+        # or the end of raw_text can be continued to a complete think_end_tag
+        think_end_pos = (
+            len(raw_text) -
+            self._ends_with_partial_token(raw_text, self.think_end_tag))
+        return raw_text[:think_end_pos], False, raw_text[think_end_pos:]
+    def _parse_unambiguous_text_content(
+            self, raw_text: str) -> tuple[str, Union[str, None], str]:
+        """
+        Returns: (unambiguous_text_content, interrupting_tag, rest_string)
+        """
+        # Either a complete tool_call_start_tag or think_start can be somewhere in raw_text
+        search_tags = [self.think_start_tag, self.tool_call_start_tag]
+        tag_positions = [(tag, pos) for tag in search_tags
+                         if (pos := raw_text.find(tag)) >= 0]
+        tag_positions.sort(key=lambda tag_and_pos: tag_and_pos[1])
+        if len(tag_positions) > 0:
+            first_tag, tag_pos = tag_positions[0]
+            return raw_text[:tag_pos], first_tag, raw_text[tag_pos:]
+        # or the end of raw_text can be continued to a complete tag
+        tag_positions = [
+            (tag, len(raw_text) - self._ends_with_partial_token(raw_text, tag))
+            for tag in search_tags
+        ]
+        tag_positions.sort(key=lambda tag_and_pos: tag_and_pos[1])
+        first_tag, tag_pos = tag_positions[0]
+        if tag_pos < len(raw_text):
+            return raw_text[:tag_pos], None, raw_text[tag_pos:]
+        return raw_text, None, ""
+    def _parse_tool_call_start_tag(self, raw_text: str) -> tuple[bool, str]:
+        """
+        Removes tool_call_start_tag from the beginning of raw_text,
+        and an optional "[", and leading whitespace.
+        Returns: (found_complete_tool_call_start_tag, rest_string)
+        """
+        if not raw_text.startswith(self.tool_call_start_tag):
+            return False, raw_text
+        rest = raw_text[len(self.tool_call_start_tag):].lstrip()
+        if rest.startswith("["):
+            rest = rest[1:].lstrip()
+        return True, rest
+    def _parse_tool_call_end_tag(
+            self, raw_text: str) -> tuple[Union[bool, None], str]:
+        """
+        Removes tool_call_end_tag from the beginning of raw_text,
+        and an optional "]" before it, and leading whitespace.
+        Returns: tuple
+            found_complete_tool_call_end_tag (or None if not decidable yet)
+            rest_string
+        """
+        # remove optional whitespace and closing ] bracket from json list notation
+        rest = raw_text.lstrip()
+        if rest.startswith("]"):
+            rest = rest[1:].lstrip()
+        if rest.startswith(self.tool_call_end_tag):
+            # found a complete tool call end tag
+            return True, rest[len(self.tool_call_end_tag):]
+        if (len(rest) >= len(self.tool_call_end_tag)
+                or rest != self.tool_call_end_tag[:len(rest)]):
+            # evidence that rest_string does not start with a tool call end tag
+            return False, raw_text
+        # incomplete tool call end tag, can not be decided yet
+        return None, raw_text
+    def _extract_arguments_from_partial_tool_call(
+            self, raw_text: str) -> Union[str, None]:
+        """
+        Extracts the raw text of the "arguments" field of a complete
+        or partial tool call.
+        Args:
+            raw_text: tool call raw text,
+                      e.g `{"name": "my_tool", "arguments": {"firstarg": "some`
+        Returns:
+            raw text of the "arguments" field, which is not valid JSON
+            unless the tool call is complete,
+            e.g. `{"firstarg": "some` for the example raw_text above
+        """
+        # assumptions:
+        # - "arguments" is always an object
+        # - there is no other field of type object in the function call
+        # - `raw_text` contains first "name", then "arguments" (otherwise,
+        #   we'd have to find the end of "arguments" before returning its
+        #   raw text value)
+        # typically, at position 0, but there might be leading whitespace
+        tool_call_start_pos = raw_text.find("{")
+        assert raw_text[:tool_call_start_pos].strip() == ""
+        arguments_start_pos = raw_text.find("{", tool_call_start_pos + 1)
+        if arguments_start_pos < 0:
+            return None
+        arguments_raw_text = raw_text[arguments_start_pos:]
+        return arguments_raw_text
+    def _parse_complete_tool_call(
+            self, raw_text: str) -> tuple[Union[dict, None], str]:
+        """
+        Returns: tuple
+            parsed tool call if complete, None otherwise
+            rest_string that needs to be parsed again or may contain
+                        a partial tool call
+        """
+        # raw_text must start without whitespace for correct parsing
+        obj, end_pos = self.extract_complete_json_dict(raw_text)
+        if obj is None:
+            return None, raw_text
+        tool_call_raw_text = raw_text[:end_pos]
+        # `tool_call_raw_text` is something like:
+        #   '{"name": "tool-name", "arguments": {...xyz...} }'
+        # we want to extract `{...xyz...}`,
+        # but `extract_arguments_from_partial_tool_call` would return
+        # everything after the second '{', i.e. `{...xyz...} }`
+        arguments_raw_text = self._extract_arguments_from_partial_tool_call(
+            tool_call_raw_text.removesuffix("}").rstrip())
+        tool_call = {
+            "name": obj.get("name"),
+            "arguments": obj.get("arguments"),
+            "arguments_raw_text": arguments_raw_text,
+            "is_complete": True,
+        }
+        return tool_call, raw_text[end_pos:]
+    def _parse_partial_tool_call(self, raw_text: str) -> Union[dict, None]:
+        # raw_text must start without whitespace for correct parsing
+        obj = partial_json_parser.loads(raw_text, Allow.ALL)
+        arguments_raw_text = (
+            self._extract_arguments_from_partial_tool_call(raw_text))
+        tool_call = {
+            "name": obj.get("name"),
+            "arguments": obj.get("arguments"),
+            "arguments_raw_text": arguments_raw_text,
+            "is_complete": False,
+        }
+        return tool_call
+    def _parse_tool_call(
+            self,
+            raw_text: str) -> tuple[Union[bool, None], Union[dict, None], str]:
+        # remove optional whitespace and closing ] bracket
+        # from json list notation
+        rest = raw_text.lstrip()
+        if rest == "":
+            # no json has been received yet
+            # -> can't tell if this will be a valid tool call
+            return None, None, raw_text
+        if not rest.startswith("{"):
+            # can't be a tool call json
+            return False, None, raw_text
+        tool_call, rest = self._parse_complete_tool_call(rest)
+        if tool_call:
+            return True, tool_call, rest
+        try:
+            tool_call = self._parse_partial_tool_call(rest)
+            # need to re-parse partial tool call later again
+            # -> return None, not True
+            return None, tool_call, rest
+        except json.JSONDecodeError:
+            # invalid json -> neither complete nor partial tool call
+            return False, None, rest
+    def _parse_tool_call_delimiter(
+            self, raw_text: str) -> tuple[Union[bool, None], str]:
+        """
+        Returns: tuple
+            does raw_text start with tool call delimiter?
+                (None if undecidable/incomplete)
+            rest_string
+        """
+        rest = raw_text.lstrip()
+        if rest == "":
+            return None, raw_text
+        has_next_tool_call = rest.startswith(",")
+        if not has_next_tool_call:
+            return False, raw_text
+        rest = rest[1:].lstrip()
+        if rest == "":
+            return None, raw_text
+        has_next_tool_call = rest.startswith("{")
+        if not has_next_tool_call:
+            return False, raw_text
+        return True, rest
+    def _parse_all(
+        self, raw_text: str, start_mode: ParsedStructure
+    ) -> tuple[str, list[dict], str, ParsedStructure]:
+        if start_mode == ParsedStructure.REASONING_CONTENT:
+            content, found_closing_think, rest = self._parse_think_trace(
+                raw_text)
+            if found_closing_think:
+                more_content, tool_calls, rest, structure = self._parse_all(
+                    rest, start_mode=ParsedStructure.CONTENT)
+                return content + more_content, tool_calls, rest, structure
+            return content, [], rest, ParsedStructure.REASONING_CONTENT
+        elif start_mode == ParsedStructure.CONTENT:
+            content, interrupting_tag, rest = self._parse_unambiguous_text_content(
+                raw_text)
+            # rest might contain a tool call start tag or a think start tag
+            if interrupting_tag == self.tool_call_start_tag:
+                more_content, tool_calls, rest, structure = self._parse_all(
+                    rest, start_mode=ParsedStructure.TOOL_CALL_START_TAG)
+                return content + more_content, tool_calls, rest, structure
+            elif interrupting_tag == self.think_start_tag:
+                more_content, tool_calls, rest, structure = self._parse_all(
+                    rest, start_mode=ParsedStructure.REASONING_CONTENT)
+                return content + more_content, tool_calls, rest, structure
+            else:
+                return content, [], rest, ParsedStructure.CONTENT
+        elif start_mode == ParsedStructure.TOOL_CALL_START_TAG:
+            found_tool_call_start_tag, rest = self._parse_tool_call_start_tag(
+                raw_text)
+            if not found_tool_call_start_tag:
+                return "", [], raw_text, ParsedStructure.CONTENT
+            # we found a complete start tag, but we haven't seen the begin of a tool call json yet
+            content, tool_calls, rest, structure = self._parse_all(
+                rest, start_mode=ParsedStructure.TOOL_CALL)
+            if not content and not tool_calls:
+                # We haven't reached the opening "{" of the tool call yet.
+                # We might see a "[" before the "{", so let's process the start tag again next chunk.
+                return content, [], raw_text, ParsedStructure.CONTENT
+            return content, tool_calls, rest, structure
+        elif start_mode == ParsedStructure.TOOL_CALL:
+            found_tool_call, tool_call, rest = self._parse_tool_call(raw_text)
+            if found_tool_call is True:
+                tool_calls = [tool_call] if tool_call else []
+                content, more_tool_calls, rest, structure = self._parse_all(
+                    rest, start_mode=ParsedStructure.TOOL_CALL_DELIMITER)
+                return (content, tool_calls + more_tool_calls, rest, structure)
+            elif found_tool_call is None:
+                # partial tool call -> need to parse again with next chunk
+                tool_calls = ([tool_call] if tool_call is not None else [])
+                return "", tool_calls, rest, ParsedStructure.TOOL_CALL
+            else:
+                logger.warning(
+                    "Invalid tool call -> continue with parsing model output as text content"
+                )
+                return self._parse_all(raw_text,
+                                       start_mode=ParsedStructure.CONTENT)
+        elif start_mode == ParsedStructure.TOOL_CALL_DELIMITER:
+            found_tool_call_delimiter, rest = self._parse_tool_call_delimiter(
+                raw_text)
+            if found_tool_call_delimiter is True:
+                return self._parse_all(rest,
+                                       start_mode=ParsedStructure.TOOL_CALL)
+            elif found_tool_call_delimiter is None:
+                # could neither confirm nor deny that raw_text starts with a tool call delimiter
+                return "", [], rest, ParsedStructure.TOOL_CALL_DELIMITER
+            else:
+                return self._parse_all(
+                    raw_text, start_mode=ParsedStructure.TOOL_CALL_END_TAG)
+        elif start_mode == ParsedStructure.TOOL_CALL_END_TAG:
+            found_tool_call_end_tag, rest = self._parse_tool_call_end_tag(
+                raw_text)
+            if found_tool_call_end_tag is True:
+                return self._parse_all(rest,
+                                       start_mode=ParsedStructure.CONTENT)
+            elif found_tool_call_end_tag is None:
+                return "", [], rest, ParsedStructure.TOOL_CALL_END_TAG
+            else:
+                return self._parse_all(raw_text,
+                                       start_mode=ParsedStructure.CONTENT)
+        logger.warning(
+            f"Unknown tool call parser start_mode '{start_mode}'. Falling back to text content."
+        )
+        return self._parse_all(raw_text, start_mode=ParsedStructure.CONTENT)
+    def extract_tool_calls_streaming(
+        self,
+        previous_text: str,
+        current_text: str,
+        delta_text: str,
+        previous_token_ids: Sequence[int],
+        current_token_ids: Sequence[int],
+        delta_token_ids: Sequence[int],
+        request: ChatCompletionRequest,
+    ) -> Union[DeltaMessage, None]:
+        """
+        Extract tool calls for streaming mode.
+        """
+        raw_text = self.streaming_state["buffer"] + delta_text
+        structure = self.streaming_state["parsed_structure"]
+        content, tool_calls, rest, new_structure = self._parse_all(
+            raw_text, start_mode=structure)
+        self.streaming_state["buffer"] = rest
+        self.streaming_state["parsed_structure"] = new_structure
+        already_streamed_tool_calls = (
+            self.streaming_state["streamed_tool_calls"])
+        already_streamed_complete_tool_calls = [
+            tool_call for tool_call in already_streamed_tool_calls
+            if tool_call["is_complete"]
+        ]
+        all_tool_calls = (already_streamed_complete_tool_calls +
+                          (tool_calls or []))
+        to_be_streamed_tool_calls = self._calculate_delta_tool_calls(
+            all_tool_calls, already_streamed_tool_calls)
+        if not content and not to_be_streamed_tool_calls:
+            return None
+        self.update_state_vars(all_tool_calls)
+        return DeltaMessage(content=content if content else None,
+                            tool_calls=to_be_streamed_tool_calls)
+    def _calculate_delta_tool_calls(
+            self, current_tool_calls: Union[list[dict], None],
+            already_streamed_tool_calls: list[dict]) -> list[DeltaToolCall]:
+        if not current_tool_calls:
+            return []
+        new_deltas = []
+        for tool_call_idx, partial_tool_call in enumerate(current_tool_calls):
+            if (partial_tool_call.get("name") is None
+                    or partial_tool_call.get("arguments") is None):
+                # do not stream arguments for an unknown tool name;
+                # and unless arguments appear in the partial json,
+                # it might be that "name" has not been received completely
+                # (assuming a template like `{"name": "mytool", "arguments": ...}`)
+                continue
+            partial_tool_call["tool_call_idx"] = tool_call_idx
+            partial_tool_call["arguments_raw_text"] = (
+                partial_tool_call.get('arguments_raw_text') or "")
+            if len(already_streamed_tool_calls) > tool_call_idx:
+                # parts of this tool_call_idx have already been streamed
+                already_streamed_tool_call = already_streamed_tool_calls[
+                    tool_call_idx]
+                delta_tool_call = self._delta_for_partial_tool_call(
+                    partial_tool_call, already_streamed_tool_call)
+                if delta_tool_call is not None:
+                    new_deltas.append(delta_tool_call)
+                already_streamed_tool_calls[tool_call_idx] = (
+                    already_streamed_tool_call | partial_tool_call)
+            else:
+                # no parts of this tool_call_idx have been streamed yet
+                tool_call = self._delta_for_new_tool_call(partial_tool_call)
+                new_deltas.append(tool_call)
+                already_streamed_tool_calls.append(partial_tool_call)
+        return new_deltas
+    def _delta_for_new_tool_call(self, tool_call_dict: dict) -> DeltaToolCall:
+        """constructs DeltaToolCall for new tool call,
+        with tool_call_id, name, and all arguments seen so far.
+        Updates tool_call dictionary with tool_call_id.
+        """
+        tool_call_idx = tool_call_dict["tool_call_idx"]
+        tool_call_dict[
+            "tool_call_id"] = f"call_{tool_call_idx}_{random_uuid()}"
+        tool_call_dict["arguments_raw_text"] = tool_call_dict.get(
+            'arguments_raw_text') or ""
+        delta_tool_call = DeltaToolCall(
+            index=tool_call_idx,
+            type="function",
+            id=tool_call_dict["tool_call_id"],
+            function=DeltaFunctionCall(
+                name=tool_call_dict.get("name"),
+                arguments=tool_call_dict["arguments_raw_text"]))
+        return delta_tool_call
+    def _delta_for_partial_tool_call(
+            self, new_tool_call: dict,
+            already_streamed_tool_call: dict) -> Union[DeltaToolCall, None]:
+        """Calculate delta for a tool call of which some parts have already been streamed."""
+        assert new_tool_call["name"] == already_streamed_tool_call["name"]
+        assert already_streamed_tool_call.get("tool_call_id")
+        if already_streamed_tool_call.get("is_complete"):
+            return None
+        to_be_streamed_arguments = (
+            new_tool_call["arguments_raw_text"].removeprefix(
+                already_streamed_tool_call["arguments_raw_text"]))
+        if not to_be_streamed_arguments:
+            return None
+        delta_tool_call = DeltaToolCall(
+            index=new_tool_call["tool_call_idx"],
+            type="function",
+            function=DeltaFunctionCall(arguments=to_be_streamed_arguments))
+        return delta_tool_call
+    def update_state_vars(self, all_tools: list[dict]) -> None:
+        # `tool_parser.streamed_args_for_tool` and
+        # `tool_parser.prev_tool_call_arr` are checked in serving_chat.py
+        # relevant is {"arguments": {...}}
+        self.prev_tool_call_arr = all_tools
+        # json-serialized argument
+        self.streamed_args_for_tool = [
+            tool_call.get("arguments_raw_text", "") for tool_call in all_tools
+        ]
+    @classmethod
+    def _ends_with_partial_token(cls, buffer: str, tag: str) -> int:
+        """
+        Check if buffer ends with a partial tag.
+        Return the length of the partial tag.
+        """
+        for i in range(1, min(len(buffer) + 1, len(tag))):
+            if tag.startswith(buffer[-i:]):
+                return i
+        return 0
+    @classmethod
+    def extract_complete_json_dict(cls, json_str: str):
+        try:
+            decoder = json.JSONDecoder()
+            obj, end_pos = decoder.raw_decode(
+                json_str)  # ignore any text after the end of the json object
+            if isinstance(obj, dict):
+                return obj, end_pos
+            return None, 0
+        except json.JSONDecodeError:
+            return None, 0