Deep Learning from Scratch (Part 3) – History 1

Deep Learning Took Almost 70 Years to Become an Overnight Success

Today, it feels like deep learning appeared almost out of nowhere. One day, it was mostly an academic topic, and a few years later, everyone was talking about ChatGPT, Claude, image generation, and AI assistants.

The reality is much less dramatic.

Most of the ideas behind modern deep learning have been evolving for decades. Some of them were proposed long before personal computers even existed.

The story begins in 1943.

Researchers Warren McCulloch and Walter Pitts published one of the earliest mathematical models of an artificial neuron. It was incredibly simple compared to today’s neural networks, but it introduced an idea that still sits at the center of deep learning. Complex behavior could emerge from networks of very simple computational units.  

Professors and students discussing mathematical formulas in a classroom.
Diagram illustrating a neural network with input signals, weights, bias, summing junction, and activ.

McCulloch & Pitts neuron diagram

A few years later, Alan Turing asked a different question.

Instead of asking how machines could learn, he asked whether a machine could behave intelligently enough that humans could no longer distinguish it from another person in conversation.

The idea became known as the Turing Test, and it shaped AI discussions long before modern deep learning existed.  

Black and white photo of a young man with vintage computer hardware.
Illustration of human and AI communication with data flow and a computer.
Illustration of the Turing test.

In 1956, researchers gathered at the Dartmouth Summer Research Project, the conference where the term Artificial Intelligence was officially introduced.

The optimism was remarkable. Many believed that machines capable of learning and reasoning were only a few decades away.

Black and white photo of Dartmouth AI research team from 1956.
Historical photo of Dartmouth Summer Research Project on Artificial Intelligence, 1956, featuring seven researchers.

Reality turned out to be considerably slower.  

Around the same period, Frank Rosenblatt introduced the Perceptron, one of the earliest neural network models.

It could perform simple binary classification and became one of the first practical implementations of ideas that had previously been mostly theoretical. IBM even demonstrated the model on its computers, making neural networks feel less like mathematics and more like software.  

Young man analyzing data charts in a modern tech laboratory.

Then came one of the most important setbacks in AI history.

In 1969, Marvin Minsky and Seymour Papert showed that a single-layer perceptron could not solve problems like XOR. The limitation itself wasn’t fatal, but many people interpreted it as evidence that neural networks had reached a dead end.

Two men smiling and talking in a library with bookshelves in the background.

Funding slowed down.

Interest faded. For a while, symbolic AI became the dominant direction instead. Deep learning wasn’t progressing rapidly. It was mostly waiting for the missing pieces.

Those pieces began to arrive during the 1980s.

Backpropagation made it possible to train multilayer neural networks efficiently, solving one of the biggest obstacles that earlier researchers had struggled with. Suddenly, having more than one hidden layer became computationally useful instead of merely theoretical.  

Looking back, what’s interesting isn’t how quickly deep learning advanced.

It’s how often the field almost disappeared before becoming one of today’s most influential technologies.

Leave a Reply

Create a website or blog at WordPress.com

Up ↑

Discover more from Writing my way through ideas.

Subscribe now to keep reading and get access to the full archive.

Continue reading