Time to First Token
You click “Send” and stare at an empty chat. For the user's comfort, it matters when the beginning of the response appears, even if writing the rest takes much longer. You need to measure that initial wait.
Time to First Token (TTFT) is the time from a specified starting point in request handling to the first token, a piece of the response. In the client measurement described here, we start when the request is sent and stop at the first nonempty content.
If the question was sent at 0 ms and the first text arrived at 800 ms, TTFT is 800 ms. This can include network time, queuing, input preparation, and computation. Measuring only the model's time would use different boundaries.
A low TTFT does not guarantee a smooth continuation: a long pause may follow the first piece. Inter-token Latency [Polski] describes this. Comparisons must always state where the measurement starts and ends.
Mechanism source: NVIDIA NIM Benchmarking documentation, “Time to First Token”.
Mechanism and details

TTFT covers more than Prefill: the client may wait for the network, input preparation, the queue, computation, and result transmission. A measurement starting only inside the server has different boundaries. Such numbers should not be compared without checking where their timestamps come from.
The first character arrives quickly, the full response late
An original event trace: sending at 0 ms, server arrival at 30 ms, 10 ms of preparation, a 200 ms queue, 120 ms of Prefill, 5 ms to prepare the first result, and 35 ms to deliver it. Client TTFT is 400 ms; the time from request arrival to result preparation on the server is 335 ms.
After the first token, we receive four more, every 20 ms. The full response arrives at 480 ms. Increase their intervals to 100 ms: completion moves to 800 ms, but the first token still arrives at 400 ms. Then increase the queue and see which results change together.
The times are synthetic, the stages do not overlap, and each message contains one token. Real streaming can group tokens. TTFT describes neither the rhythm of Decode nor the time to receive the complete response; the documentation separates these metrics in “Inter-token Latency” and “End-to-End Request Latency”. Inter-token Latency [Polski] describes intervals and pauses after the first token. Prefix Caching can reduce prompt processing, while Continuous Batching affects when the scheduler admits a request for execution.
I use AI-generated content as part of my daily learning process.