The Memory Problem and Continuity of LLMs

I was wondering why memory prices are increasing so much. I noticed that memory companies suddenly became market darlings, with their share prices rising sharply around their quarterly results. Their profits were increasing, their order books were full, and everyone seemed to be pointing towards AI. 

At the same time, the cellphone and laptop manufacturers have to deal with the same rising memory costs while competing with deep-pocketed AI companies for memory capacity. 

So I started wondering: why does AI need so much memory? 

I don’t have a definitive answer. This is just where my curiosity took me. 

I was chatting with an LLM about behavioural economics, particularly Daniel Kahneman’s Thinking, Fast and Slow. The conversation was actually quite useful. The system gave me perspectives that helped me understand some of the concepts more deeply. 

A few days later, I went back to the LLM for something else. The conversation naturally drifted towards behavioural economics again, so we started discussing it. 

Then I noticed something strange. 

The system was giving me answers that were correct, but we had essentially had the same conversation before. 

It had no recollection of that earlier discussion. 

When I reminded it about our previous conversation, it tried to move the discussion in a different direction. When I asked why it had not remembered our earlier discussion about the book, it explained that it did not have access to that previous conversation unless the relevant information had been stored in some form of long-term memory. 

That got me thinking about the difference between how humans and LLMs remember things. 

Humans don’t remember every conversation verbatim. We don’t need to store every sentence someone has spoken to us. Instead, we retain certain important ideas, associations and experiences. 

I may not remember the exact words of a conversation I had three months ago. But I might remember that I discussed Kahneman, that we disagreed about something, or that I had an interesting insight about System 1 and System 2. 

In other words, human memory is selective. 

We compress enormous amounts of experience into relatively small collections of concepts and associations and retrieve them when they become relevant. 

That made me wonder whether this is one of the fundamental challenges for AI as it becomes more useful in our daily lives. 

An LLM doesn’t simply need to generate a good answer. If we expect it to become a genuine long-term assistant, it also needs to know what from our previous interactions is worth remembering, what can be forgotten, and what needs to be retrieved when we need it. 

And that is not simply a question of adding more RAM. 

There are several different kinds of memory involved in an AI system. The model itself contains its learned information. During a conversation, it needs working memory to process the information currently available to it. And then there is the longer-term memory that could store useful information about previous interactions. 

The last one can potentially be stored outside the model and retrieved when necessary. 

That sounds surprisingly similar to what humans do. 

We don’t need to remember everything. We need to remember the right things. 

Of course, as AI systems become larger and conversations become longer, the physical memory requirements also increase. Modern AI accelerators use specialised high-bandwidth memory to move enormous amounts of data quickly between the processor and memory. Longer contexts and increasingly demanding AI workloads put further pressure on both memory capacity and memory bandwidth. 

So perhaps the question isn’t simply: How much memory does AI need? 

It is also: 

How efficiently can AI remember, retrieve and use information? 

And I suspect this problem is not going to disappear simply by adding more hardware. 

More memory can postpone the problem. But eventually another bottleneck appears. 

The hardware becomes more expensive. The hardware also becomes obsolete. Training increasingly capable models requires enormous amounts of computation and energy. Inference—the process of actually running these models for users—also consumes resources at a massive scale. 

Then there is another problem. 

What happens when AI-generated data becomes an increasingly large part of the data used to train future AI systems? 

Perhaps synthetic data will become extremely useful. Perhaps it will also introduce new limitations if models increasingly learn from information generated by earlier models rather than from independent sources. 

I don’t know what the next bottleneck will be. 

But one thing seems increasingly clear to me: AI scaling is not going to have a single bottleneck. 

Solve computation and memory becomes important. 

Solve memory and bandwidth becomes important. 

Solve those and energy and infrastructure become important. 

Solve the infrastructure and economics become important. 

And eventually, the question may not be whether we can build a more capable model. 

It may be whether we can build one that is economically sustainable. 

This is where the investor in me becomes interested. 

AI companies are competing aggressively to capture market share while their infrastructure expenses are increasing rapidly. If the technology continues to require enormous amounts of capital before the economics have fully matured, companies may have a strong incentive to raise large amounts of capital while investor enthusiasm remains high. 

And that brings us to the public markets. 

Some of the biggest technology companies and potentially some of the biggest IPOs of the coming years may come from this AI boom. 

The question is not whether AI is going to change the world. 

I suspect it will. 

The question is how much of that future success is already reflected in today’s valuations. 

There is a big difference between believing in a technology and believing that its current price is justified. 
I’m definitely not saying that AI will fail. I am saying that I don’t yet know how many of these bottlenecks will be solved, how much capital will be required to solve them, or how much of that cost can ultimately be passed on to customers. 

So for now, I’m happy to be a spectator and watch the story unfold. 

If I miss the bandwagon, let it be. 

I’d rather miss a spectacular opportunity than buy into one at a valuation that assumes everything will go perfectly. 

Because in technology, the most interesting question is often not what is possible. 

It is what it will cost to make it possible.

Popular posts from this blog

How AI leverages our confirmation bias?

The Freedom of Constraints

Morality, Power, and Choice: A Systems View