Calling LLMs stochastic parrots is... frankly wrong. It's true that pre-training LLMs is literally matching the empirical distribution of human text, so this would be the only place where "stochastic parrot" is a coherent description of the objective. The capabilities that emerge when training for this objective are anything but parrotry.
We've empirically proved that when learning high dimensional datasets, the interpolation that "stochastic parrotry" implies empirically never happens, and extrapolation actually is the result (arXiv:2110.09485 [cs.LG]). Other research also shows out of distribution generalisation of transformers (Othello-GPT arXiv:2210.13382 [cs.LG] shows that a transformer trained purely on move sequences develops a linear, causally-manipulable representation of board state).
#AI #AIResearch #Transformers #LLM #MechanisticInterpretability