DevQuasar

community

Verified

https://devquasar.com/

Activity Feed

AI & ML interests

Open-Source LLMs, Local AI Projects: https://pypi.org/project/llm-predictive-router/

Recent Activity

csabakecskemeti updated a model less than a minute ago

DevQuasar/THUDM.GLM-Z1-32B-0414-GGUF

csabakecskemeti updated a model about 1 hour ago

DevQuasar/THUDM.GLM-4-32B-0414-GGUF

csabakecskemeti published a model about 4 hours ago

DevQuasar/THUDM.GLM-Z1-32B-0414-GGUF

View all activity

DevQuasar's activity

csabakecskemeti

updated a model less than a minute ago

DevQuasar/THUDM.GLM-Z1-32B-0414-GGUF

Text Generation • Updated less than a minute ago

csabakecskemeti

updated a model about 1 hour ago

DevQuasar/THUDM.GLM-4-32B-0414-GGUF

Text Generation • Updated about 1 hour ago • 1

csabakecskemeti

published a model about 4 hours ago

DevQuasar/THUDM.GLM-Z1-32B-0414-GGUF

Text Generation • Updated less than a minute ago

csabakecskemeti

updated a model about 15 hours ago

DevQuasar/nvidia.Llama-3_1-Nemotron-Ultra-253B-v1-GGUF

Text Generation • Updated about 15 hours ago • 1.75k • 7

csabakecskemeti

published a model about 18 hours ago

DevQuasar/THUDM.GLM-4-32B-0414-GGUF

Text Generation • Updated about 1 hour ago • 1

csabakecskemeti

in DevQuasar/nvidia.Llama-3_1-Nemotron-Ultra-253B-v1-GGUF about 22 hours ago

Thank you!! (IQ quants)

#2 opened 4 days ago by

BoshiAI

csabakecskemeti

updated a model 2 days ago

DevQuasar/Salesforce.Llama-xLAM-2-8b-fc-r-GGUF

Text Generation • Updated 2 days ago • 13

csabakecskemeti

published a model 2 days ago

DevQuasar/Salesforce.Llama-xLAM-2-8b-fc-r-GGUF

Text Generation • Updated 2 days ago • 13

csabakecskemeti

updated a model 2 days ago

DevQuasar/Salesforce.Llama-xLAM-2-70b-fc-r-GGUF

Text Generation • Updated 2 days ago • 11

csabakecskemeti

published a model 3 days ago

DevQuasar/Salesforce.Llama-xLAM-2-70b-fc-r-GGUF

Text Generation • Updated 2 days ago • 11

csabakecskemeti

updated a model 3 days ago

DevQuasar/prithivMLmods.Bootes-Xuange-Qwen-14B-GGUF

Text Generation • Updated 3 days ago • 314

csabakecskemeti

posted an update 7 days ago

Post

1983

Local Llama4 Maverick Q2
https://youtu.be/4F8g_LThli0?si=MGba2SUTHt6xYw3T
Quants uploading now

Big thanks to @ngxson !

csabakecskemeti

posted an update 8 days ago

Post

1637

Why the 'how many r's in strawberry' prompt "breaks" llama4? :D

Quants DevQuasar/meta-llama.Llama-4-Scout-17B-16E-Instruct-GGUF

3 replies

csabakecskemeti

posted an update 23 days ago

Post

3356

I'm collecting llama-bench results for inference with a llama 3.1 8B q4 and q8 reference models on varoius GPUs. The results are average of 5 executions.
The system varies (different motherboard and CPU ... but that probably that has little effect on the inference performance).

https://devquasar.com/gpu-gguf-inference-comparison/
the exact models user are in the page

I'd welcome results from other GPUs is you have access do anything else you've need in the post. Hopefully this is useful information everyone.

csabakecskemeti

posted an update 25 days ago

Post

2381

Managed to get my hands on a 5090FE, it's beefy

| llama 8B Q8_0 | 7.95 GiB | 8.03 B | CUDA | 99 | pp512 | 12207.44 ± 481.67 |
| llama 8B Q8_0 | 7.95 GiB | 8.03 B | CUDA | 99 | tg128 | 143.18 ± 0.18 |

Comparison with others GPUs
http://devquasar.com/gpu-gguf-inference-comparison/

csabakecskemeti

posted an update 28 days ago

Post

1810

GTC new model announcement now from Nvidia
nvidia/Llama-3_3-Nemotron-Super-49B-v1

GGUFs:
DevQuasar/nvidia.Llama-3_3-Nemotron-Super-49B-v1-GGUF

Enjoy!

csabakecskemeti

posted an update about 1 month ago

Post

572

Cohere Command-a Q2 quant
DevQuasar/CohereForAI.c4ai-command-a-03-2025-GGUF

6.7t/s on a 3gpu setup (4080 + 2x3090)

(q3, q4 currently uploading)

csabakecskemeti

posted an update about 1 month ago

Post

827

Fine tuning on the edge. Pushing the MI100 to it's limits.
QWQ-32B 4bit QLORA fine tuning
VRAM usage 31.498G/31.984G :D

4 replies

csabakecskemeti

posted an update about 1 month ago

Post

1966

-UPDATED-
4bit inference is working! The blogpost is updated with code snippet and requirements.txt
https://devquasar.com/uncategorized/all-about-amd-and-rocm/
-UPDATED-
I've played around with an MI100 and ROCm and collected my experience in a blogpost:
https://devquasar.com/uncategorized/all-about-amd-and-rocm/
Unfortunately I've could not make inference or training work with model loaded in 8bit or use BnB, but did everything else and documented my findings.

4 replies

csabakecskemeti

posted an update about 2 months ago

Post

2789

Testing Training on AMD/ROCm the first time!

I've got my hands on an AMD Instinct MI100. It's about the same price used as a V100 but on paper has more TOPS (V100 14TOPS vs MI100 23TOPS) also the HBM has faster clock so the memory bandwidth is 1.2TB/s.
For quantized inference it's a beast (MI50 was also surprisingly fast)

For LORA training with this quick test I could not make the bnb config works so I'm running the FT on the fill size model.

Will share all the install, setup and setting I've learned in a blog post, together with the cooling shroud 3D design.

8 replies

AI & ML interests

Recent Activity

Team members 2

DevQuasar's activity

Thank you!! (IQ quants)