When ChatGPT launched in November 2022, a lot of people assumed large language models (LLMs) launched with it. But researchers and software developers had been building software that predicts words, recognizes speech, and translates sentences for decades, and they were using neural networks long before Google researchers published the 2017 paper that introduced the transformer. The timeline above covers the milestones in between.
More than 75 years ago, researchers were already describing language with statistics. Claude Shannon was generating text from word-pair statistics in 1948, and in 2000, counting word sequences was still the standard way to model language. The challenging part was context. Those early models only looked a few words back, so they could finish a stock phrase but couldn’t hold onto the meaning of a full sentence or paragraph. Neural networks could eventually learn from much longer stretches of text, but they needed to be trained on far more text than hardware could handle at the time.
A typical desktop in 2000 had a single processor running at under 1 GHz and somewhere between 64 and 256 megabytes of memory, nowhere near enough to train a large network on millions of documents. Technologies often stall this way. The idea is ready, but the hardware to run it isn’t. The computer scientist Sara Hooker calls this the hardware lottery: an idea’s success can depend on whether it happens to suit the hardware of its day, and neural networks spent years losing that lottery.
I started in this field in 2000. My first natural language processing (NLP) project was a text classifier that sorted customer support emails into categories, built with a support vector machine (SVM) and a pile of hand-written Perl scripts to clean up the text first. Neural networks came up now and then, usually as something a professor had tried in the ’90s and dropped because the networks took days to train, needed more data than anyone had, and gave different results every time you ran them.
The short version: Large language models didn’t start with ChatGPT. Researchers were modeling language with statistics as early as 1948, published the first well-known neural language model in 2003, and added attention to translation models in 2014. The transformer arrived in 2017, and scaling it up with huge amounts of text and human feedback led to ChatGPT in 2022.
Table of Contents
ToggleBefore attention, there were language models
A language model estimates which words are likely to come next in a sequence. Give it “Please close the” and it will rate “door” as far more likely than “volcano.” The model doesn’t know what a door is. It has learned which words tend to appear together.
Early statistical models worked by counting. You’d feed a program a large collection of text, and it would tally how often each short sequence of words appeared. These sequences are called n-grams, where n is the number of words. A trigram model looks at the previous two words to guess the next one. If “the cat sat” showed up 500 times in the training text and “the cat flew” showed up twice, the model would predict “sat” after “the cat.” This simple approach was used for spell checkers, early predictive text on phones, and a lot of speech recognition software, and it worked well for common phrases.
Neural networks changed how those predictions were made. A counting model treats every word as an unrelated symbol, so to it, “cat” and “dog” have nothing in common. A neural network instead learns a list of numbers for each word (called an embedding), adjusting them during training. You can think of those numbers as coordinates that place each word somewhere on a map, and words used in similar ways end up close together. That means what the model learns about one word carries over to similar ones. If it has seen “the cat sat on the mat” many times, it can guess that “the dog sat on the mat” is a reasonable sentence too, even if that exact phrase never appeared in its training text.
In 2003, Yoshua Bengio and his colleagues published a neural probabilistic language model that learned these embeddings while predicting the next word. Like the counting models, it still only looked at a few previous words. It had no chat window and would look tiny next to today’s models, but it was a neural network trained to predict the next word, which is the same basic job an LLM does. For the bigger story of how machine learning got to this point, see From Mendel’s Peas to ChatGPT.
When neural networks were out of fashion
Bengio’s paper came out while neural networks were still out of favor, years after what many researchers call the second AI winter. Neural networks had a reputation problem when I started out. Many researchers saw them as an old idea that had never delivered, and papers built around them had a hard time getting into major conferences.
Part of the problem was technical. Networks with more than a few layers were very hard to train. Training works by sending an error signal backward through the network to adjust each layer, and in deep networks that signal shrank as it traveled, until the earliest layers barely changed. Researchers call this the vanishing gradient problem. Support vector machines and random forests came with cleaner math, gave more predictable results, and ran well on the computers of the time. They were the safer pick for most projects.
A small group stuck with neural networks anyway. Geoffrey Hinton, Yann LeCun, and Yoshua Bengio (the same Bengio behind the 2003 language model) kept arguing that deeper networks would eventually work. Hinton later joked that they were a kind of conspiracy. The three went on to share the 2018 Turing Award for this work.
The first break came in 2006. Hinton and his colleagues showed that a deep network could be trained one layer at a time, with each layer first learning patterns from unlabeled data (examples with no human-written labels or answers attached) before the whole network was tuned. Around the same time, the group started calling the work deep learning, a name that carried less baggage than “neural networks.” People started paying attention again.
Better hardware helped too. Graphics processing units (GPUs), built to render video game graphics, turned out to be very good at the math neural networks need, because they run thousands of simple calculations at once. A 2009 study from Andrew Ng’s group at Stanford reported training up to about 70 times faster on a GPU than on a regular processor. Speech recognition was the first big commercial payoff. Hinton’s students worked with researchers at Microsoft, Google, and IBM on neural network models for speech, and the groups reported large drops in error rates. Google put the approach into Android voice search in 2012.
That same year brought two results that got attention well outside speech research. At Google, a team including Andrew Ng and Jeff Dean trained a large network on 10 million unlabeled still images taken from YouTube videos, and one part of it learned to respond to cat faces without ever being told what a cat was. A few months later, Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton entered a network called AlexNet in the image recognition challenge built on Fei-Fei Li’s ImageNet dataset, a collection of more than a million photos labeled by hand. AlexNet’s top-5 error rate (the share of images where the correct label wasn’t among the model’s five best guesses) was 15.3 percent. The next best entry came in at 26.2 percent.
I remember a colleague forwarding me the AlexNet results in late 2012 with a one-line note: “Looks like the neural net people were right.” Within a couple of years, most new projects on my team started with a neural network instead of an SVM. The big tech companies were competing for the people who knew how to build these models. In 2013, Google acquired Hinton’s startup, DNNresearch, which brought Hinton, Krizhevsky, and Sutskever to the company, and Facebook hired Yann LeCun to lead a new AI research lab. Language researchers, who had been working with neural networks on a smaller scale, now had the hardware and the attention to go much further.
Embeddings had their own breakthrough in 2013. A team at Google led by Tomas Mikolov released word2vec, a fast way to learn embeddings from billions of words of text. It became famous for one result: take the numbers for “king,” subtract “man,” add “woman,” and you land close to “queen.” Word2vec only learned word vectors and never generated text, but its vectors became a standard starting point, and nearly every language model since has begun by turning words into lists of numbers like these.
Reading one word at a time
Go back to Bengio’s 2003 model for a moment. It looked at a fixed window of a few words, and the word you need might be ten words back. Recurrent neural networks, or RNNs, handle language as a sequence. They read one word, update an internal state (a running summary of what came before), then read the next word. In principle, that state carries information forward through the whole sentence. Researchers were using RNNs as language models well before transformers. A 2010 paper, for example, used one to improve speech recognition. LSTMs, a type of recurrent network introduced in 1997, were designed to hold on to information across longer stretches of text.
Translation makes the limits easy to see. A translation model has to read a sentence in one language and write it in another. Early neural translation models read the whole input sentence and squeezed it into a single fixed-size summary, then wrote the translation from that summary alone. That worked for short sentences. Longer ones lost details, because one summary can only hold so much. Think of reading a paragraph, putting your notes away, and then translating it from memory. You’d probably get the gist and lose a detail from the beginning. Attention was designed to fix that.
The ML Advocate Monthly
New ML Advocate posts plus the best machine learning and AI reads, once a month. No hype, no doom.
Attention changes translation
In 2014, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio proposed a way for a translation model to look back at the input sentence while writing each word of the output. Instead of relying on one compressed summary, the model could give more weight to the source words most relevant to the word it was writing at that moment. Take the sentence “The cat slept under the table.” When the model writes the word for “cat,” the part of the input about “cat” is what it needs. When it gets to “under the table,” different words become more useful. Attention lets the model shift its focus as it writes.
Attention improved how a neural model used the information it already had, and it didn’t replace recurrent networks right away. Google’s production translation system in 2016 used both: it read sentences one word at a time and used attention to pick out the relevant parts while writing the translation. Many AI histories jump from early neural networks straight to transformers and leave this part out, even though the transformer was built directly on it.
Then came the transformer
In 2017, a team at Google published Attention Is All You Need. Their transformer architecture used attention without the recurrent layers most translation models relied on. The original paper was about machine translation, and the authors reported strong results. The key piece is self-attention. In the translation example, attention helps the model look back at the input sentence. Self-attention lets each word in a sentence be represented in relation to the words around it. “Bank” means something different next to “river” than next to “loan,” and the model can represent it differently in each case.
That design also changed how training worked. A recurrent network reads one word at a time, so each step has to wait for the one before it. During training, a transformer looks at every word in a sentence at once, which is exactly the kind of work GPUs do well. It still writes its output one word at a time, but training became much faster. In Sara Hooker’s terms, neural networks finally had an architecture that suited the hardware of its day, and that made much bigger models practical to train.
Despite the paper’s title, attention wasn’t the only ingredient. Because a transformer doesn’t read from left to right, it needs a separate way to record where each word sits in the sentence, and its layers can be stacked deep enough to keep growing the model. The transformer gave researchers an architecture they could scale up, and the next few years were spent figuring out how to train and use it.
From language models to assistants
After 2017, research teams tried different ways to train transformers on huge amounts of text and then reuse what they learned. In 2018, OpenAI trained a transformer as a language model and then fine-tuned it for specific tasks, meaning they gave the general model extra training on a smaller set of examples for each task. Google’s BERT, released the same year, learned by filling in words that had been hidden from a sentence, using the words on both sides of the gap. That made it good at tasks like question answering. BERT was built for understanding text (sorting, searching, answering questions), while OpenAI’s GPT models were built to generate it, and that line of work is the one that led to ChatGPT. Which of these counts as the first large language model depends on where you draw the line for “large,” but many people point to these 2018 models.
Then the models got much bigger. The 2018 OpenAI model had about 117 million parameters (the numbers a network adjusts during training, like the embeddings described earlier). In 2020, the GPT-3 paper described a model with 175 billion, and showed it could handle a range of tasks from instructions and a few examples in the prompt, with no separate training for each task. Later work on following instructions made models better at doing what people asked. In training with human feedback, people rated the model’s answers, and those ratings were used to train it toward the responses people preferred. ChatGPT put all of that into a chat window in November 2022.
The scaling surprised a lot of people in the field, me included. Everything I’d learned about overfitting said a model with that many parameters would memorize its training data and fall apart on anything new. The bigger models kept getting better. In 2020, researchers at OpenAI showed that language models improved smoothly and predictably as models, data, and computing power grew, and GPT-3 handled tasks it was never specifically trained for. Researchers are still working out why. Rich Sutton summed up the pattern in a 2019 essay called The Bitter Lesson: over the long run, general methods that scale with computing power tend to beat methods built on human expertise.
No single person invented large language models. They grew out of decades of work by many people, from Shannon’s word statistics in 1948 to the Google team that introduced the transformer in 2017, and OpenAI and Google turned those pieces into the first LLMs.
Twenty-two years later
For most people, ChatGPT was their first open-ended conversation with a language model, so it felt brand new even though the technology behind it had a long history. The first time I used it, I asked it to explain support vector machines, partly as a test. Its answer was clearer than some of the textbooks I learned from in 2000, and I sat there for a minute thinking about everything that had happened in between.
Knowing where LLMs came from also makes it easier to see when they’re the wrong tool. If you’re weighing an LLM against simpler options, start with Do You Need Rules, ML, an LLM, or an AI Agent? or test your instincts with the AI-ML Fit Quiz.


