Early generative AI models, circa 2017-2022, were “stochastic parrots.” That is, they generated language by choosing the statistically most likely next word based on patterns in their training data, rather than truly understanding what they were saying.
Beginning in 2023, artificial intelligence labs began releasing large language models (LLMs) that had evolved beyond autoregressive next-token prediction, as researchers call it.
For many people, the stochastic parrot definition stuck. And lately, the “they’re just stochastic parrots” argument has been used as a way of downplaying the potential risks of huge language models such as OpenAI’s Astra models or Anthropic’s Mythos models.
But LLMs in 2026 can’t properly be called stochastic parrots. Yes, they still predict next words, but those predictions are informed by far more than static patterns found in their training data. These four research areas, among other things, have pushed AI chatbots far beyond the ones we used just a few years ago.
Retrieval Augmented Generation (RAG)
Between 2017 and 2020, AI researchers began giving LLMs access to information outside their training data, retrieving relevant documents and feeding them into the model. In some cases, the model used a web index to find the right document, pull the relevant snippet of information from it, then assemble a number of such snippets into a coherent, conversational answer delivered within an internet search or chatbot setting. This improved factual accuracy by giving the model more recent, relevant, and authoritative material to draw from, rather than forcing it to rely entirely on what it had learned during training. The process is known as retrieval augmented generation, or RAG.
Neurosymbolic systems
If RAG gave language models access to “ground truth” information, new research into neurosymbolic AI gave them a more structured form of computation to better interpret and use it. Imagine asking whether a complicated insurance policy covers a particular procedure. An LLM could retrieve the relevant sections, but a neurosymbolic system could translate the policy, with all its definitions, conditions, and exceptions, into explicit facts and rules, boiling it down to deterministic, flow-chart-like language such as “the procedure is covered if the patient has Plan A, has met the deductible, and has either prior authorization or an applicable exception.” A rules engine or logic solver can then apply those rules systematically and return a result to the LLM, which turns it into a natural-language answer. This symbolic “machinery” could be a conventional, deterministic computer program operating alongside the neural model, or it could be developed as an integrated logic system within the LLM during training.
Chain of thought
In 2022, researchers at Google and the University of Tokyo showed that LLMs, if prompted correctly, have the inherent ability to break down complex problems into smaller steps. Giving the model the simple instruction, “Let’s think step by step,” could bring out a latent capability to reason through problems, without showing the model examples or changing any of its weights. The model was still an autoregressive next-token predictor, but generating a sequence of intermediate tokens effectively gave it something like a scratchpad for showing its work. The chain-of-thought research set the stage for the next major step in the evolution of generative AI: reasoning models.
Reasoning models and reinforcement learning
Soon, researchers were redesigning models, and training them differently, to enable them to “reason” for longer periods while working to formulate an answer in real time, during live inference after being prompted by a user. For example, a reasoning model might break a problem down into parts, devise a number of competing approaches, and check its own work and correct errors. OpenAI’s o1 model, released in 2024, was the first reasoning model from a major lab. OpenAI researchers used a combination of pretraining, including human-written examples, and reinforcement learning to teach o1 how to reason.
Then, in early 2025, the Chinese lab DeepSeek showed that it could teach a regular LLM how to reason without explicit training, using only reinforcement learning, in which the model is rewarded for getting the answer right. While striving to earn the reward, the model began producing longer chains of thought, checking itself, reconsidering approaches, and exploring alternatives.
Note that much of this research was going on concurrently, not in successive stages. Retrieval, neurosymbolic approaches, and chain-of-thought research overlapped considerably between 2019 and 2023, while reasoning models came shortly after. AI labs are still building on these ideas, using techniques such as reinforcement learning and test-time computation to make models better at reasoning.
The point is that generative AI models no longer just run prompts through their webwork of parameters, the billions of little dials that were set while the model processed mountains of content during training. That’s just the start. The model does a lot more after that to create a better, more reliable answer. The stochastic parrot is now well connected and has a PhD.