Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
This paper proves that Transformers trained via gradient descent cannot efficiently learn majority Boolean functions, with error growing exponentially with input dimension.
Bo Chen, Zhenmei Shi, Zhao Song et al.