Artifact 14
Prefill & Decode made Crystal Clear

These two terms are very common. And it's very possible that you may have heard them before, or maybe you even know what they mean. But I feel that even though earlier I knew about prefill and decode, I didn't actually understand them correctly.
What I was thinking about them when I was starting with inference was pretty different from what they actually are.
So that's why I think a good, correct explanation can surely save you from a lot of bad thinking, wrong imagination, or simply applying the wrong logic later because you never properly understood what these two things actually are.
I will not go to any definitions of prefill or decode. You can find it anywhere or with any AI. I just want to make you understand where these two things come into the whole picture (of inference). Once that is clear, I think you will automatically understand what prefill and decode actually are.
( Disclaimer - This is going to be a really long read, because I will go step by step from first principles, with no hurry. And I hope you expect so. )
Okay, so let's start with a very simple example of chatting with Grok, ChatGPT, Claude, or any LLM of your choice. ( I'll use gpt to keep simple )
You go to the chat interface, click on a new chat, and type your prompt.
Let's say you type:
"Hi, my name is X."
You hit return/enter
And, and after a very short amount of time, maybe one second, 1.5 seconds, or two seconds, it starts giving you the response.
Maybe the first few tokens like:
"Hello X..."
And then it keeps generating the rest of the response. After some time, when it has generated the complete response, it stops. So this is generally what happens whenever we chat with an LLM. We give it some input, wait for some small time, and then it starts giving us its response.
Now let's take a slightly deeper look at what is actually happening inside this whole process. But before going deep, let's understand a few basics so that we can relate things in a better way.
Modern LLMs like GPT, Llama, Claude, etc. are generally autoregressive language models.
What does autoregressive mean?
Basically, the model predicts the next token based on the tokens it already has.
Suppose, just as a very simple example, the model currently has:
"Hello"
Based on this, it predicts the next token. Maybe that token is:
"world"
Now the sequence becomes:
"Hello world"
The model now uses "Hello world" to predict what should come next.
Maybe the next token is "!". And this same process keeps repeating.
So during the whole generation that we talked about earlier, the LLM is basically repeating this process again and again.
- It first gets your input prompt.
- Then it predicts the first new token.
- Now it has your original prompt plus the newly generated token.
- It uses all of that context to predict the next token.
- Then it has the original prompt + first generated token + second generated token.
- It predicts another one.
- Then another one.
And this keeps happening until the model generates an end-of-sequence token, decides that the response is complete, or reaches the maximum number of output tokens it is allowed to generate.
So this was just some basic stuff that we should know.
Now let's come back to our ChatGPT session. You typed:
"Hi, my name is X."
and clicked Enter.
The prompt reaches the servers running the model. (It may be the first intuition that model will understand the prompt, and that happens)
But there is one problem.
Your prompt is a text (string).
The model doesn't directly understands with text like we do. Internally, model is a pile of huge matrices and numbers (not necessarily huge numbers)
So the first thing we do is tokenize the text. Our text gets broken into tokens, and each token is represented using a token ID. You can go to openai tokenizer and see this yourself.

So now we have token IDs.
But we still cannot directly pass these IDs through all the transformer layers.
Each token ID is first mapped to an embedding, which is simply a vector of numbers. A vector is basically a mathematical row containing many numbers (called dimension).

So let's say our prompt contains S tokens, and the model uses an embedding dimension of D.
Each token becomes a vector of size D.
If we stack the vectors for all S tokens together, our input becomes a matrix:
X ∈ R^(S × D)
where:
S = number of tokens in our input sequence D = hidden/embedding dimension of the model
So we started with normal text:
"Hi, my name is X"
Then:
text → tokens → token IDs → embeddings
And now we have this S × D matrix of numbers.
This is where the actual transformer computation can start. And this is also where prefill starts entering the picture.
Okay, so now we have this matrix X.
What actually happens when we pass this through the model ?
This matrix now starts going through the transformer layers.

And if you have seen the basic transformer architecture before, you know that inside every transformer block, there are mainly two big things happening:
Attention and MLP/FFN.
I will not go deep into how attention itself works here because then this article will become an attention article instead of a prefill and decode article.
But there is one thing that is very important for us.
Suppose our input has 10 tokens.
One thing you may naturally imagine is that the model will first process token 1, then token 2, then token 3, then token 4, and so on. Because after all, this is an autoregressive model, right?
But this is where I myself had the wrong picture initially.
During the input prompt, this is not how things happen.
Why?
Because all the input tokens are already known. You already gave the complete prompt to the model. If your prompt is:
"Hi, my name is X"
the model already knows every token in this sentence. so why it should process these input tokens sequentially ? why not simultaneously - it will be much faster, right ?
It doesn't have to generate "my" before it can know "name".
It doesn't have to generate "name" before it can know "Abhinandan".
You have already given all of these tokens to it.
And because all of these tokens are already available, we can process them together in parallel.
This entire processing of the input prompt is what we call prefill.
Basically, the first model pass which processes the whole input prompt, is called prefill.

Probably a definition can be,
Prefill is the phase where the model processes the tokens that are already available in your input prompt.
Now there is one thing you may ask.
If all the tokens are processed together, then isn't token 1 able to see token 10?
And if you are wondering why that would even be a problem, let's make this clear first.
Remember, this is an autoregressive model. Its whole job is to predict the next token using only the tokens that came before it. Suppose we have:
The sky is blue
When the model is at:
The sky is
and trying to predict the next token, it should only be allowed to use:
The sky is
It should not already be able to look at:
blue
because that's literally the token it is supposed to predict.
If it can already see "blue" while trying to predict "blue", then we are basically letting the model cheat.
It's like going into an exam where the correct answer is already ticked, and then saying:
"Yes, I think this is the answer."
Of course you do. You already saw it.
So when we say that all prompt tokens can be processed together during prefill, we do not mean that every token is allowed to look at every other token.
The computation can happen in parallel, but the information each token is allowed to use still has to follow the autoregressive rule:
a token can look at itself and the tokens before it, but not the tokens after it.
This is what we call causal mask.
So token 1 can only attend to token 1.
Token 2 can attend to tokens 1 and 2.
Token 3 can attend to tokens 1, 2, and 3.
And so on.
The GPU can still calculate all of these positions together. The causal mask simply makes sure that no token gets to peek into the future.
Or, in even simpler words:
**parallel computation is allowed; future information is not. **
This distinction is actually very important.
I think one wrong imagination people may have is:
"Autoregressive means everything has to happen token by token."
That is true during generation.
But it is not true in the same way during the processing of the prompt.
The prompt already exists.
The output does not.
This small difference is basically the reason why prefill and decode become two very different phases.
Now let us look slightly deeper into what is happening during this prefill.
Inside the attention layer, for every token, we calculate three things:
Query, Key, and Value | Q, K, V
Again, I will assume you already have a basic idea of attention here.
So if we have S prompt tokens, each of these tokens gets its own K and V at every transformer layer.
And here comes something extremely important.
After calculating those Keys and Values, we do not throw them away.
We store them.
This stored collection of Keys and Values is what we call the KV cache.
Why are we storing it?
At this moment it may feel like okay, we calculated something and stored it. But the real reason becomes very clear once decode starts.
So for now, just remember:
During prefill, we process (basically attention + ffn) all the prompt tokens, and along the way, we create the K and V for those tokens and store them in the KV cache.
Now the prompt goes through layer 1, layer 2, layer 3, and so on, all the way through the model.
At the end, we use something called **LMhead **to produce logits.
Basically, logits means raw scores over the complete vocabulary.
And from these scores, we select what the next token should be.
So let us come back to our example.
We had:
"Hi, my name is X."
After prefill, maybe the model decides that the next token should be:
"Hello"
And this is an important.
The first output token comes at the end of the prefill phase.
This is why, when you click Enter in ChatGPT, there is generally some amount of waiting before you see anything.
The model first has to process your input prompt. Only after that can it produce the first new token.
This waiting time is commonly called:
TTFT - Time To First Token.
You send your request.
Then some time passes.
Then the first token appears.
Of course TTFT is affected by many other things too: scheduling, networking, batching, server load, and so on.
But from the model execution side, prefill is a very important part of that waiting.
And this also gives you one intuition.
If your input prompt becomes very large, there is simply more prompt to process before the model can start generating the response.
Okay.
Now something has changed.
Until this point, everything the model was processing was given by you.
All those input tokens already existed.
But now the model has generated:
"Hello"
And what does an autoregressive model do now?
It has to predict what comes after "Hello".
Maybe:
**"X" **(I am ignoring the (space) before token for simplicity)
But here's the problem.
Unlike the input prompt, the model does not already know "X".
It first has to generate "Hello".
Only after "Hello" exists can it use it to predict the next token.
And only after the next token exists can it predict the token after that.
Now we cannot process the complete output in parallel because the future output tokens simply do not exist yet.
This is where decode starts.
Decode is basically the generation phase.
One new token is generated.
Then that token is used to generate the next one.
Then that one is used to generate the next one.

And this keeps happening until the response is finished.
So if I have to put the difference in the simplest possible way:
During prefill, the tokens are already known.
During decode, the tokens are being created.
I think this one sentence already gives a much better idea of why they behave differently.
But now there is another question.
Suppose after prefill we had:
"Hi, my name is X."
and the model generated:
"Hello"
To generate the next token, do we again process:
"Hi, my name is Abhinandan. Hello"
from the beginning?
And then for the next token, process everything again?
If we did this, it would be a huge waste.
We already spent all that compute processing the original prompt.
Why would we calculate the same things again and again for every generated token?
This is exactly where the KV cache that we stored during prefill becomes useful.
Remember that during prefill, for every prompt token, we calculated its K and V and stored them.
So now when "Hello" enters the model, we only need to calculate the new Query Q, Key K, and Value V for this new token.
For all the previous tokens, their K and V are already available inside the KV cache.
The Query of the new token can attend to those cached Keys and Values.
Then we take the new K and V generated for "Hello" and append them to the KV cache.
So now the KV cache contains the information for:
the original prompt + "Hello".
When the next token comes, the same thing happens again.
We process the new token. Basically pass it though model
Its Query attends to the K/V already sitting in the cache.
We generate a new K and V for this token.
Append them to KV cache.
Calculate logits of last token & enerate the next token.
Repeat.
This is basically the decode loop.
There is one thing here that is worth making very clear because I also had confusion about this.
Sometimes when we hear:
"KV cache saves computation during decode"
we may start imagining that during decode, maybe only the attention layer runs, or perhaps the rest of the transformer somehow gets skipped.
That is not what happens.
The new token still goes through the complete transformer.
It goes through attention.
Then MLP.
Then the next transformer block.
Attention.
MLP.
And this keeps happening through all the layers.
The saving is that we are not sending all the previous tokens through those layers again.
Their computation has already happened.
For the new token, we still do the full forward pass.
So, if your context currently has 2,000 tokens, during decode we are not running the MLP again for all 2,000 old tokens.
We only run the new token through the MLP.
For attention, the new token still needs to look back at the previous context, and the KV cache makes that possible without recalculating the old Keys and Values.
This distinction is important because otherwise it is very easy to build the wrong mental model of what KV cache is actually saving.
Now here comes something slightly more interesting.
Prefill and decode are running the same model.
Same transformer.
Same weights.
Same attention.
Same MLP.
But from the GPU's perspective, these two workloads look very different.
Let us again assume our prompt has 2,000 tokens.
During prefill, we have a matrix containing representations for all these tokens.
So when we do something like a linear projection inside attention or inside the MLP, we can perform fairly large matrix multiplications.
There is a lot of work that can happen together.
And GPUs are extremely good at this.
They have a huge number of cores (CUDA & tensor cores), and when you give them large enough matrix operations, they can keep a lot of those cores busy doing useful computation.
So prefill generally has a lot of parallelism because the number of tokens is very high, and for every token we are doing attention + MLP across every transformer layer.
That means there is a very large amount of computation happening at once - potentially billions or trillions of operations.
This is why you will often hear that prefill is more compute-bound.
But what does compute-bound actually mean?
It simply means that the main thing limiting how fast this work can finish is the compute power of the GPU. We already have enough data to keep the GPU busy, and now there is just a huge amount of math to perform. So having more/faster compute can directly help prefill finish faster.
Now look at decode.
For one request, what do we have?
One new token.
That's it.
We have one vector that has to go through this huge model, which may contain billions/trillions of parameters.
For every decode step, the GPU still needs to access those model weights, but now we are doing relatively little computation with them because we only have one new token.
So compared to prefill, there is much less computation happening for the amount of data we need to move.
And this is why decode often becomes memory-bandwidth bound.
Memory-bound basically means that the GPU has enough compute power to do the math, but the main bottleneck is how quickly we can bring the required data - mainly model weights, and also KV-cache data - from HBM to the compute units.
So during decode, a lot of the time the GPU is essentially waiting for data to arrive from memory.
The faster the HBM bandwidth, the faster we can feed those weights and data to the GPU cores, and therefore the faster decode can become.
So very roughly:
Prefill → lots of tokens + lots of math → compute is usually the bottleneck.
Decode → very few new tokens + lots of weights to read → memory bandwidth is usually the bottleneck.

Obviously real systems have batching, different kernels, different architectures, and many other optimizations.
But this is the basic reason why prefill and decode behave so differently on GPUs.
And now something that may have looked random before starts becoming logical.
Why do inference engines care so much about batching during decode?
Suppose I have only one request decoding.
At every step, I have just one new token to process.
That is not a lot of parallel work for a massive GPU.
But suppose I have 100 users currently generating responses.
Now every decode step may contain one new token from each of those 100 requests.
So instead of processing just one token, we may process a batch of many decode tokens together.
Now the GPU gets more work at once.
This is one of the reasons batching is so important in LLM serving.
And modern inference engines do something even better called continuous batching, where requests can enter and leave the batch dynamically instead of waiting for one fixed batch to completely finish.
But this is probably a separate topic by itself.
The important thing for us here is simply:
Decode for one request is very sequential.
But across many requests, the inference engine can still create parallelism.
Now let us come back to the experience you actually see in ChatGPT.
You type a prompt.
You press Enter.
Nothing appears immediately.
During that time, among other things, your prompt is being processed.
That is the prefill side.
Then suddenly the first token appears.
After that:
token...
token...
token...
token...
and the response starts streaming.
That is decode happening again and again.
This also gives us two different latency metrics that you will see very commonly in inference.
The first one we already talked about:
TTFT - Time To First Token.
How long does it take from sending the request until the first token appears?
Then we have something called:
ITL - Inter-Token Latency.
This is basically the time between two generated tokens during decode.
If the first token takes a long time to appear, you may have bad TTFT.
If the first token appears quickly, but after that the response generates very slowly, then your decode or inter-token latency may be poor.
So even from the user's experience, you can kind of feel these two phases separately.
Now let us take two slightly extreme examples because this makes the distinction even clearer.
Suppose request A has:
20,000 input tokens
but only needs:
20 output tokens.
This request has a huge amount of prompt to process.
So there is a lot of prefill work.
Once generation starts, though, only 20 decode steps may be needed.
Now imagine request B.
It has:
20 input tokens
but asks the model to generate:
2,000 output tokens.
Its prefill is tiny.
There is barely any input to process.
So the first token may come quickly.
But then the model has to keep decoding:
one token,
then another,
then another,
for potentially thousands of steps.
So even if the total number of tokens looks comparable in some situation, the workload can be very different depending on whether those tokens belong to the input or to the output.
This is one of those things which becomes obvious once you understand prefill and decode, but before understanding them, input tokens and output tokens can just look like "tokens".
They are not exactly the same from the serving perspective.
And this difference between prefill and decode creates some interesting problems for inference engines.
Imagine you already have a few requests decoding on the GPU.
They are generating tokens smoothly.
Now suddenly a new request arrives with a 50,000-token prompt.
If you simply let this massive prefill consume the GPU for a long time, what happens to the requests that were already decoding?
Their next tokens may get delayed.
And from the user's perspective, their streaming response may suddenly become slower.
This is why inference engines need a scheduler.
The scheduler basically has to keep making decisions like:
How much prefill work should I run?
How many decode requests should I run?
Should I process this complete long prompt right now?
Or should I process only a part of it and come back to it later?
This is where ideas like chunked prefill start coming into the picture.
Instead of processing one massive prompt entirely in one go, the engine may divide its prefill into smaller chunks.
This allows it to mix prefill work with decode work more carefully.
Again, this is a full topic of its own.
But now at least you know why such a thing even needs to exist.
It is not some random optimization.
And it comes from the fact that prefill and decode have very different behavior, and both of them may be competing for the same GPU.
And people have gone one step further.
If prefill and decode are so different, do they even need to run on the same GPU?
Not necessarily.
This leads to something called prefill-decode disaggregation.
The rough idea is that you can have some workers mainly handling prefill and some other workers mainly handling decode.
A request first goes through a prefill worker.
The prompt gets processed.
The KV cache gets created.
Then that KV cache is transferred to a decode worker.
And the decode worker continues generating the output tokens.
So conceptually:
Prompt → Prefill Worker → KV Cache Transfer → Decode Worker → Output Tokens
Now you can scale prefill and decode resources somewhat independently.
Maybe your workload has extremely long prompts.
You may need more prefill capacity.
Maybe your workload generates very long responses.
You may need more decode capacity.
Of course, doing this creates new problems too.
The KV cache can be huge.
Now you have to transfer it between machines or GPUs efficiently.
Networking becomes important.
Scheduling becomes harder.
There are many trade-offs.
But again, you can now see where the whole idea comes from.
It starts with the simple fact that:
processing known input tokens and generating unknown output tokens are two fundamentally different kinds of work.
So now if we go back to the very beginning of this article, I think the whole chat interaction looks slightly different.
When you type:
"Hi, my name is X."
and press Enter, the LLM does not simply "start generating".
There is a phase before generation.
Your text becomes tokens.
Those tokens become embeddings.
All the input tokens go through the transformer.
Their Keys and Values are calculated and stored in the KV cache.
This is prefill.
At the end of that process, the model predicts the first new token.
And now the nature of the work changes.
The model takes that newly generated token, runs it through the transformer, uses the KV cache to attend to the previous context, adds the new K and V to the cache, and predicts another token.
Then another.
Then another.
That repeated one-token-at-a-time generation is decode.
So if I have to keep only one picture in my mind, it would probably be this:
Prefill = process what we already know.
Decode = generate what we don't know yet.
During prefill, the complete input prompt is already available, so there is a lot of parallelism.
During decode, every future token depends on the token generated before it, so generation is fundamentally sequential for a single request.
And somewhere between these two phases sits the KV cache, carrying the useful attention state from everything we already processed into everything we are about to generate.
That's basically prefill and decode.
And once I understood them this way, instead of just memorizing two definitions, a lot of other things in inference started making much more sense.
Hope you have read till last (which you are doing right now :), and found it a real worth your time.