Dust: Pretraining Transformers Without Backpropagation

Dust: Pretraining Transformers Without Backpropagation

Dust introduces a new approach to pretrain transformer models without relying on backpropagation. The method uses alternative optimization techniques to update weights, potentially lowering memory demands and speeding up training. Researchers claim the technique can match or approach performance of traditional backprop-based pretraining while reducing computational overhead. The approach could make large language models more accessible.