DEV Community

#quantization

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU

Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU

1
Comments
10 min read
KV Cache INT4 Quantization for 1M+ Token Context Windows

KV Cache INT4 Quantization for 1M+ Token Context Windows

Comments
3 min read
Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers

Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers

Comments
3 min read
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX

Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX

Comments
3 min read
Quantizing MedGemma to INT4 (GPTQ/W4A16): Everything That Broke Along the Way

Quantizing MedGemma to INT4 (GPTQ/W4A16): Everything That Broke Along the Way

Comments
6 min read
LLM Inference Optimization: From Quantization to Speculative Decoding

LLM Inference Optimization: From Quantization to Speculative Decoding

1
Comments 1
2 min read
Self-Hosting a Model Means Self-Hosting Its Evaluation Too

Self-Hosting a Model Means Self-Hosting Its Evaluation Too

1
Comments 2
5 min read
Gemma 4 QAT on a 1080 Ti: What 'Quantization-Aware' Actually Buys — and Fitting the 12B on 8 GB at 16k

Gemma 4 QAT on a 1080 Ti: What 'Quantization-Aware' Actually Buys — and Fitting the 12B on 8 GB at 16k

Comments
5 min read
Quantization formats compared: GGUF vs GPTQ vs AWQ vs NF4

Quantization formats compared: GGUF vs GPTQ vs AWQ vs NF4

Comments
7 min read
INT8 Q/DQ Calibration on Blackwell: 1.8 the TRT 10 + FP16 Baseline

INT8 Q/DQ Calibration on Blackwell: 1.8 the TRT 10 + FP16 Baseline

Comments
7 min read
LLM Quantization Levels Compared: Q4_K_M vs Q8_0 vs FP16 [2026]

LLM Quantization Levels Compared: Q4_K_M vs Q8_0 vs FP16 [2026]

Comments 1
15 min read
How to Pick a GGUF Quant Level for Your VRAM Budget

How to Pick a GGUF Quant Level for Your VRAM Budget

Comments
4 min read
Why your quantized LLM loses its MTP heads and how to keep them

Why your quantized LLM loses its MTP heads and how to keep them

1
Comments
5 min read
The Best Result This Week Was a Failed Prediction — Phase-3a Doesn't Transfer

The Best Result This Week Was a Failed Prediction — Phase-3a Doesn't Transfer

Comments
1 min read