DeusImperator
/

Dumpling-Qwen2.5-32B_exl2_4.5bpw_L

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Configuration Parsing Warning: In config.json: "quantization_config.bits" must be an integer

Dumpling-Qwen2.5-32B - EXL2 4.5bpw L

This is a 4.5bpw EXL2 quant of nbeerbower/Dumpling-Qwen2.5-32B

This quant was made using exllamav2-0.2.7 with default dataset and extended quantization sample length (8k instead of default 2k). It also uses -head_bits=8 and max accuracy quant for first and last layer (8bpw), all other layers of the model use normally chosen methods (method and name (4.5bpw_L) inspired by quants like Q4_K_L and Q6_K_L made by bartowski)

I tested it some some RPs (also ones over 12k context) and it seems to work. It fits nicely in 24GB VRAM on Windows with 16k fp16 context (should fit 2x that with q8 cache in exl2).

Prompt Templates

Seems to use ChatML

Original readme below

Dumpling-Qwen2.5-32B

nbeerbower/Rombos-EVAGutenberg-TIES-Qwen2.5-32B finetuned on:

Method

ORPO tuned with 8x A100 for 2 epochs.

Downloads last month: 3

Inference Providers NEW

Text Generation

This model is not currently available via any of the supported third-party Inference Providers, and the model is not deployed on the HF Inference API.

Model tree for DeusImperator/Dumpling-Qwen2.5-32B_exl2_4.5bpw_L

Base model

nbeerbower/Rombos-EVAGutenberg-TIES-Qwen2.5-32B

Finetuned

nbeerbower/Dumpling-Qwen2.5-32B

Quantized

(4)

this model

Datasets used to train DeusImperator/Dumpling-Qwen2.5-32B_exl2_4.5bpw_L