What Does MLX 4-bit Cost? A Controlled Audit of Quantization for Code Generation on Apple Silicon

J. Melton · American Code Labs · Preprint, August 2026

Download PDF

Abstract

A controlled measurement of what MLX default 4-bit quantization costs five code models on HumanEval+ and MBPP+ against bf16 baselines converted from the same weights. Pooled across five models, 220 discordant pairs favour bf16 against 125 favouring 4-bit, and the loss shrinks with model size at 0.52 points per billion parameters. A calibrated AWQ arm does not recover it.

Code and data

github.com/american-code/ResearchPapersdata/slm-benchmark/

Licensed CC BY 4.0.