Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation Methods
Benchmarking five optimizers (SGD, Adam, Adagrad, Adadelta, CGD) with three memory strategies across four transformer architectures; gradient accumulation proved most robust, optimizer performance varies by architecture.
Sarthak Mahapatra, Zihan Zhou, Khatoon Khedri et al.