Thor Lin
coolthor
AI & ML interests
On-device & edge LLM inference on NVIDIA GB10 / DGX Spark.
Quantization (NVFP4 W4A4/W4A16, FP8), vLLM serving, speculative
decoding (EAGLE-3), and multimodal/omni models. Measure-first
benchmarking — I publish the numbers, including the ones that fail.
Recent Activity
updated a model 4 days ago
coolthor/Kandinsky-6.0-Pro-distill-5s-NVFP4 published a model 4 days ago
coolthor/Kandinsky-6.0-Pro-distill-5s-NVFP4 updated a model 13 days ago
coolthor/Qwen3.8-27B-Uncensored-NVFP4-FastLLMOrganizations
None yet