Back to news
arXiv cs.CL · 2026-08-12 00:00 UTC
research

The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs

arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a zero-shot multilingual evaluation of 4-bit quantization across the Gemma 4 and Qwen 3.5 architectures. Evaluating on eight typo-logically diverse languages using MMLU ProX Lite and GlobalPIQA, we show parameter truncation exposes deep pre-training inequalities. We identify four phenomena: (1) Typological Fragility: low-r

Why it matters

Quantization may disproportionately harm multilingual edge models, guiding practitioners to budget capacity or use language-aware compression to avoid severe accuracy collapse.

Read the original story

Published to Cognify News · Week 33, 2026