Retrofitting language models to operate over bytes

Retrofitting language models to operate over bytes

Researchers propose modifying large language models to work directly with byte sequences instead of tokenized text. By training on raw byte streams, the approach aims to reduce preprocessing complexity, improve handling of rare or unseen characters, and enable more universal language coverage. Experiments show comparable performance on standard benchmarks while simplifying the model pipeline for downstream tasks.