How HTTP/3.0 and QUIC Solve Head-of-Line Blocking

How HTTP/3.0 and QUIC Solve Head-of-Line Blocking

You probably have heard the history of how HTTP was created and you might even know about the new HTTP/3 version. In this article, I want to focus on some of the lesser-known details of HTTP and why we need a new version of this ubiquitous protocol.

In the following sections, we will explore the evolution of HTTP and how each iteration tackled the challenges of performance and scalability, culminating in the resolution of Head-of-Line (HoL) blocking with HTTP/3 and QUIC. We’ll begin by examining the origins of HTTP and its initial limitations with the short-lived connection model in HTTP/1.0. From there, we’ll dive into the enhancements introduced in HTTP/1.1, including persistent connections and pipelining, and discuss their shortcomings. Next, we’ll explore HTTP/2’s advances and how these improvements shifted but didn’t eliminate the HoL blocking issue. Finally, we’ll uncover how HTTP/3, powered by the QUIC protocol, overcomes TCP’s inherent limitations and provides a robust solution for modern web performance.

But first, we need to start from the beginning…

HTTP/1.0: Order Amidst Chaos#

In the early years of HTTP, we saw the emergence of web browsers and the rapid growth of consumer-oriented public Internet infrastructure. However, because this one-line protocol was designed to serve only hypertext documents, it was extremely limited. This simple protocol eventually became unofficially known as HTTP/0.9. With the need for new features, web developers followed an ad hoc process: implement, deploy and see if people adopt them. With that process, a set of best practices and patterns emerged and the informational RFC 1945 documented the “common usage” of the many HTTP/1.0 implementations at the time. Take a look at this excerpt from the RFC:

This specification reflects common usage of the protocol referred too as
“HTTP/1.0”. This specification describes the features that seem to be
consistently implemented in most HTTP/1.0 clients and servers.

Even with all this effort, there was a huge problem with the protocol: By default, you should always open a new connection to request a resource and close it immediately after you receive it.

This approach is known as a short-lived connection model and it’s generally undesirable because HTTP is built on top of TCP, which requires a Three-Way Handshake to establish a connection, introducing latency. This delay becomes especially problematic when loading numerous resources on a webpage.

Moreover, TCP employs a slow-start mechanism to manage network congestion. This means it begins with a conservative use of available bandwidth, gradually ramping up as the connection persists and stabilizes. Frequent opening and closing of connections, as in the short-lived model, disrupts this ramp-up process, preventing TCP from fully utilizing the network’s potential and further degrading performance.

Leaked image of a TCP server’s protocol in action (there is still the final ACK btw)
Leaked image of a TCP server’s protocol in action (there is still the final ACK btw)

HTTP/1.1: Internet Standard#

This standard was designed to solve some ambiguities and also introduced some performance optimizations like: keep-alive and request pipelining.

Keep-alive connections can remain open, meaning that we won’t suffer the overhead of creating and destroying connections every time a resource needs to be fetched. In the image below you can see the short-lived and persistent connection model which represents the default way of handling connections in HTTP/1.0 and HTTP/1.1 respectively.

— Models of connections HTTP/1.x
Figure 1 — Models of connections HTTP/1.x

Persisting the connection is great for performance, but can we make it even better?

Theoretically we can, and this is what HTTP pipelining is for. With it, before you even get the response of your first request, you issue a subsequent request (see Figure 1 above). Therefore you save time because the server can answer those as soon as possible, right? The problem is that with HTTP/1.1 there is no way to link the response with the request, so the only way to implement this pipelining is to respect the ordering of requests. This means that even if the first request is the slowest and the server already has the response for the other requests, it must respect the ordering and thus the first one will block the others since it is way slower. This problem is called Head-of-Line blocking.

Respecting the ordering is difficult since there are transparent proxies between a client and a server, and those proxies don’t implement HTTP pipelining, or they implement them in the wrong way. Such problems are so real that by default browsers disable this feature; as a result, pipelining is rarely used in practice.

HTTP/2: Optimized Transport#

HTTP/2 supports all of the core features of previous versions but operates more efficiently than HTTP/1.1 by introducing the concept of frame and streams that would allow multiplexing of request/response.

Multiplexing of requests is achieved by having each HTTP request/response exchange associated with its own stream (Section 5). Streams are largely independent of each other, so a blocked or stalled request or response does not prevent progress on other streams. — RFC 9113

This means that over a single TCP connection, multiple HTTP/2 requests can be made in parallel. This process is different than the HTTP/1.1 pipelining where request/responses are sequentially ordered and solves the HoL blocking at the application level.

— HTTP/2: Request/Response Multiplexing
Figure 2 — HTTP/2: Request/Response Multiplexing

Although the stream concept is effective to some extent, we can still suffer from HoL blocking at the transport level (TCP). HTTP/2 is built on top of TCP, which is a protocol that guarantees the reliability and ordering of segments. TCP does not know what it is transporting, it just does the job, and because of this it can cause the HoL blocking.

Figure 3 below shows an example where the browser is trying to request the files script.js and style.css from the server. The script.js and style.css are being transmitted in streams 1 and 2 respectively. Notice that since TCP is agnostic about what is being transmitted; it could potentially transmit pieces of two files in the same segment.

— HTTP/2 packetization and multiplexing for 2 files
Figure 3 — HTTP/2 packetization and multiplexing for 2 files

TCP guarantees order of delivery between its packets and in a scenario where packet 1 is lost, even though it only contains pieces of the script.js file, it will block the other packets from being delivered from stream 2 which relates to styles.js and thus causing the HoL blocking.

TCP Head-of-Line Blocking
Figure 4— TCP Head-of-Line Blocking

Let’s take a look at another example from Figure 4, focusing on the syscalls (recv in this case). Imagine that each segment — P1, P2, and P3 — is from a different stream. Each stream is independent, allowing them to be delivered to the application individually. However, because TCP is unaware of these separate streams, it won’t deliver any segments to the receiver until they’re in the correct order. If the first segment is lost, TCP will initiate a retransmission, delaying all subsequent segments and negatively impacting performance.

You can think of this scenario as if three images were requested in the same connection using HTTP/2 streams, but because of TCP HoL blocking, one of the image’s missing pieces will affect the other images, preventing them from being rendered until the lost piece arrives.

You can you see the same problem arising. First we had a problem with HTTP/1.1 pipelining, which tried to solve on HTTP/2 with streams, but TCP has the same problem. How do we solve it then? Simple — don’t use TCP and go with UDP.

HTTP/3: Breaking Free from the HoL#

HTTP/2 did not solve the HoL problem completely, it was just pushed down the stack (transport layer). To solve this issue we could create a new version by:

A) Extend TCP to allow streams over it

B) Use the Stream Control Transmission Protocol (SCTP)

C) Create a new protocol that uses UDP

Let’s talk about each of those possible options.

Option A — Extending TCP wouldn’t work because middleboxes (the machines that make the internet infrastructure) could inspect the TCP segments, think they are malformed and drop them. Updating these middleboxes to recognize new protocols would solve this problem, but in practice, they’re rarely updated. This inability to extend or deploy new protocols on the internet is known as protocol ossification.

Option B — The Stream Control Transmission Protocol (SCTP) could leverage streams to eliminate HoL blocking. Still, SCTP is not a good transport protocol alternative since some firewalls only allow TCP or UDP segments to pass through (A Comparison between SCTP and QUIC). For that reason, protocol designers must ensure that solutions to the HoL blocking issue are middlebox-proof, which has led to the QUIC protocol.

Option C — To put it simply, QUIC is a UDP-based multiplexed and secure transport protocol with ambitious goals. One of them is multiplexing without HOL blocking. The fact that QUIC is a UDP-based protocol gives it an advantage over SCTP since most middleboxes already implement UDP. Another aspect of the protocol is that it doesn’t have a clear-text version, which prevents ossification since boxes cannot inspect the transport header (more here). Isn’t that brilliant?

As you can see, the QUIC protocol is the natural replacement for TCP and this led to the HTTP/3. Although HTTP/3 uses QUIC, the latter protocol can be used for other use-cases, which is intentional since new protocols can be built on top of it and benefit from its features.

See the image below where we compare HTTP/2 and HTTP/3:

— Basic HTTP/3 vs HTTP/2 comparison — Layering Focus by Marx, R.
Figure 5 — Basic HTTP/3 vs HTTP/2 comparison — Layering Focus by Marx, R.

As you can see, the QUIC protocol is responsible for the reliability and congestion control provided by TCP and also for stream multiplexing of HTTP/2. So even using UDP we will still have ordering of the same streams as QUIC will ensure it.

Figure 6 shows us how the packetization process for one file happens for the protocols discussed in this article. Notice that QUIC uses UDP under the hood, so if we had the exact same scenario as in Figure 4 where all the segments were in the receiver buffer, they would be delivered to the QUIC layer, which would be responsible for figuring out what could really be delivered to the application. Others from different streams would be delivered to the application layer.

— HTTP/1.1 vs HTTP/2 vs HTTP/3 packetization for 1 file
Figure 6 — HTTP/1.1 vs HTTP/2 vs HTTP/3 packetization for 1 file

With this setup, we solved the TCP HoL blocking issue on HTTP/3 for different streams and we still maintain the original features of HTTP/2, all because of UDP.

Did you notice that I said that we solved HoL blocking for different streams? That is correct! QUIC still can’t deliver out-of-order packets for the same stream, which makes total sense, right? Why would you even deliver a text totally out of order? You can deliver in chunks, but they should be in order. So yes, we still have HoL blocking, but only for the same stream, and this is fine.

And with that, now you know why we need QUIC and HTTP/3!


Bonus#

Another thing I want to shed light on is the fact that QUIC is a user-space protocol, meaning it doesn’t have the same privileges as TCP, which operates in kernel space. While this design allows for rapid evolution and easier updates to QUIC implementations, it also introduces performance challenges, especially on high-speed networks. This is due to receiver-side processing overhead and limited offloading techniques, as highlighted in “QUIC is not Quick Enough over Fast Internet”. By November 2024, 28.8% of all websites were using HTTP/3 and in my opinion, this percentage will increase when QUIC makes it into kernel space because it will make things even faster.

Originally published on Loka Engineering on Medium.

Tags