Instructions to use ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ") model = AutoModelForMultimodalLM.from_pretrained("ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ
- SGLang
How to use ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ
Qwen3.8-Flash-Next-P48NVFP4-MoESQ
A W4A4 + paired-4:8 sparse compressed checkpoint of
Qwen/Qwen3.8-Flash-Next, produced with MoESQ.
The routed MoE expert weights are NVFP4 with paired-4:8 structured sparsity and are stored
sparse: only the kept values plus a small mask are on disk. The target is NVIDIA Blackwell
(SM100 and SM120) sparse tensor cores.
- Base model: Qwen/Qwen3.8-Flash-Next (MoE, 512 routed experts with 10 active plus a shared expert, 48 layers: 36 linear-attention and 12 sparse-attention, hyper-connections, per-layer n-gram embedding)
- Compression: NVFP4 W4A4 plus paired-4:8 sparsity on the routed MoE experts
- Routed experts: 225.0 GiB in BF16 → 38.7 GiB here, 5.8× smaller. That is 2.75 bits per weight including the sparsity mask and scales (2 bits per weight for the kept values).
- GPU-resident weights: 239.9 GiB in BF16 → 53.5 GiB, 4.5× smaller (about 27 GiB per GPU at TP2). The per-layer n-gram embedding tables (95.4 GiB, BF16, unchanged) are lookup tables served from CPU memory, as in the base model.
- Checkpoint size: 149.0 GiB on disk, 64% of it the n-gram tables (see the breakdown below).
- Kernel: paired-4:8 sparse NVFP4 grouped GEMM (CUTLASS, SM100 and SM120) through vLLM's
paired48_nvfp4MoE backend
Links
- Paper: coming soon
- Code: IST-DASLab/MoESQ. It contains the compression code, the vLLM integration and the kernels, and becomes public together with the paper.
- Other MoESQ checkpoints: ISTA-DASLab/moesq collection
Usage
This checkpoint does not load in upstream vLLM. The MoESQ repository installs a patched
vLLM v0.30.0 (the paired48_nvfp4 MoE backend and sparse-storage loader) together with the
kernels. It needs an NVIDIA Blackwell SM100 (B200, GB200) or SM120 (RTX 5090, RTX PRO 6000) GPU and a
CUDA toolkit >= 12.8; SM103 (B300) and SM121 (DGX Spark) are not supported.
git clone --recurse-submodules https://github.com/IST-DASLab/MoESQ.git && cd MoESQ
bash integrations/vllm/install.sh && source .venv-vllm/bin/activate
vllm serve ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ \
--tensor-parallel-size 2 --enable-expert-parallel # 2x B200, ~27 GiB of weights per GPU
vLLM selects the backend automatically. The model is chat-only: use the chat endpoint (plain
few-shot completions end immediately, for the BF16 model too). Pass
--default-chat-template-kwargs '{"enable_thinking": false}' to serve without thinking.
Recommended sampling
Use the base model's thinking-mode setting: temperature=1.0, top_p=0.95, top_k=20.
For long single-turn reasoning (competitive programming, math), also set
presence_penalty=1.5. On very long reasoning traces this model is more likely than the BF16
model to fall into repetition loops ("maybe maybe maybe …"), which run until the token budget is
exhausted. The base model card allows a presence penalty of 0–2 to reduce endless repetition.
On LiveCodeBench v6 at 65,536 new tokens, the penalty lifts this model from 43.09 to 49.03 pass@1;
on one seed it moved the BF16 model by only +1.1. We saw no language mixing at 1.5.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ",
messages=[{"role": "user", "content": "..."}],
temperature=1.0,
top_p=0.95,
presence_penalty=1.5, # long single-turn reasoning only; keep 0 for agentic tool calling
max_tokens=131072,
extra_body={"top_k": 20},
)
For agentic tool calling, keep presence_penalty=0. vLLM applies the penalty to the tokens
already generated in the current response. After the model reasons about a command, the tool
call's own tokens get penalized: on SWE-bench Verified, tool calls without their required
argument nearly tripled, and resolution dropped (70.6% vs 74.7% on the same 170 instances).
Evaluation
OpenLLM Leaderboard v1: 6-task average
ARC-Challenge (25-shot), HellaSwag (10-shot), MMLU (5-shot), TruthfulQA-MC2 (0-shot) and
Winogrande (5-shot) are scored with lm-evaluation-harness on /v1/completions. GSM8K (5-shot,
greedy, strict match) runs through the chat template with the few-shot examples as turns and
thinking off, because the model is chat-only. Mean ± sd over few-shot seeds 1234, 0 and 1;
both rows served with the vLLM above (BF16 at TP4, this model at TP2 + EP, KV cache bf16).
| Model | ARC-C | GSM8K | HellaSwag | MMLU | TQA-MC2 | Winogrande | Avg | Recovery |
|---|---|---|---|---|---|---|---|---|
| Qwen3.8-Flash-Next (BF16) | 78.90 | 96.92 | 88.81 | 86.71 | 67.66 | 73.56 | 82.09 ± 0.28 | — |
| This model (MoESQ) | 74.74 | 94.34 | 83.54 | 84.56 | 61.84 | 75.16 | 79.03 ± 0.15 | 96.3 % |
Coding: LiveCodeBench v6 and SWE-bench Verified
Thinking on, temperature=1.0, top_p=0.95, top_k=20 for both models. This model adds
presence_penalty=1.5 on LiveCodeBench (see Recommended sampling);
the BF16 model runs at its card's default of 0. Both models are served with the vLLM above (BF16
at TP4, this model at TP2 + EP, KV cache bf16, context 131,072, or 147,456 for the 131,072-token
LiveCodeBench run).
| Model | LCB-v6, 65,536 new tokens | LCB-v6, 131,072 new tokens | SWE-bench Verified | Recovery (LCB 131k / SWE) |
|---|---|---|---|---|
| Qwen3.8-Flash-Next (BF16) | 55.26 ± 0.97 | 57.03 ± 1.26 | 79.8 | — |
| This model (MoESQ) | 49.03 ± 1.34 | 53.54 ± 1.45 | 67.0 | 93.9 % / 84.0 % |
- LiveCodeBench v6: 175 problems (
lcb:codegeneration_v6in lighteval), 0-shot, pass@1, mean ± sd over 10 seeds. An answer still inside its reasoning when the token budget runs out counts as wrong. At 131,072 tokens that is 19.8% of this model's answers and 1.4% of BF16's. Without the presence penalty, this model scores 43.09 ± 1.33 at 65,536 tokens. - SWE-bench Verified: 500 instances, one run per model. The agent is mini-swe-agent 2.4.6
with native tool calls (one bash tool, parsed by vLLM's
qwen3_coderparser), at most 250 steps, and no network inside the sandbox. Grading uses the SWE-bench harness logic under Apptainer instead of Docker. On the 484 instances whose reference patch passes in that setup, the scores are 81.8 (BF16) and 68.6 (this model). No presence penalty for either model.
Size breakdown
Measured from the safetensors headers of this checkpoint and of the BF16 base model.
| Component | BF16 base | This model | Where it lives when served |
|---|---|---|---|
| Routed experts (48 layers × 512, 120.8 B params) | 225.0 GiB | 38.7 GiB (paired-4:8 NVFP4, sparse storage) | GPU |
Attention, linear attention, hyper-connections, shared experts, embeddings, lm_head, norms, router, vision tower |
10.0 GiB | 10.0 GiB (BF16) | GPU |
| MTP module (speculative decoding) | 4.9 GiB | 4.9 GiB (BF16) | GPU |
| GPU-resident total | 239.9 GiB | 53.5 GiB | |
| Per-layer n-gram embedding tables (51.2 B entries) | 95.4 GiB | 95.4 GiB (BF16) | CPU |
| Checkpoint total | 335.3 GiB | 149.0 GiB |
Only the routed experts are compressed. The same expert weights stored as dense NVFP4 (pruned values kept as zeros) would take 59.8 GiB; the paired48 sparse storage drops the zeros and keeps a 4-bit mask per 4-byte chunk.
Compression details
| Field | Value |
|---|---|
| Weights | NVFP4 (E2M1), one FP8-E4M3 scale per 32 dense K elements, FP32 per-tensor global scale |
| Activations | NVFP4, dynamic per-32 group scales, FP32 per-tensor global scale stored per expert linear |
| Sparsity | Paired 4:8 along K: in every 8 consecutive elements (4 pairs), exactly 2 pairs are kept |
| Compressed layers | Routed MoE experts (gate_proj, up_proj, down_proj) in all 48 layers |
| Left in BF16 | lm_head, embeddings, attention (linear and sparse), norms, router (mlp.gate), shared expert, hyper-connections, per-layer n-gram embedding, MTP head |
| Format | compressed-tensors, nvfp4-pack-quantized, with a paired48_sparse storage marker |
Gate and up projections carry separate global scales; the patched vLLM folds their ratio into the block scales at load time (upstream vLLM would apply the gate scale to both).
Sparse storage layout (paired48_sparse, pair-bitmask v1)
Each routed-expert linear stores weight_sparse_packed (uint8 [out, K/4], the 2 kept bytes of
every 4), weight_sparse_mask (uint8 [out, K/16], 4 bits per 4-byte chunk with exactly 2 set),
weight_scale (float8_e4m3 [out, K/32]) and the float32 global scales. The encoding is
lossless, and every tensor in this upload was checked with a round trip.
Recipe (MoESQ, arm "gw2")
The full MoESQ config for this run is in moe_sq_config.yaml.
- Init: GPTQ, 4-bit, 512 calibration samples, percdamp 0.1, block size 128. Masks come from paired-4:8 pruning.
- Refinement: masks and weight values are learned jointly for 10 epochs on 8,192 mixed
calibration sequences of up to 4,096 tokens (masks LR 2e-4, weights LR 2.5e-5), with
activations fake-quantized to NVFP4. The objective is the expert reconstruction error weighted
by the router gate (
gate_weight_exponent = 2).
License
This checkpoint is a derivative of Qwen/Qwen3.8-Flash-Next and is distributed under the Qwen Community License 1.0.
Citation
@article{moesq,
title = {Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts},
author = {TODO},
journal = {arXiv preprint arXiv:TODO},
year = {TODO}
}
Contact
For questions, open a discussion here or contact kwanhee.lee@postech.ac.kr.
- Downloads last month
- 114
Model tree for ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ
Base model
Qwen/Qwen3.8-Flash-Next