Sub-1-Bit LLM Compression via Latent Factorization
Researchers have introduced a new method that compresses large language models to less than one bit per parameter by using latent factorization techniques. The approach decomposes model weights into low‑rank components, enabling a drastic reduction in storage while preserving performance. Early tests show comparable accuracy to full‑size models on standard benchmarks, suggesting a promising path for deploying LLMs on resource‑constrained devices.