Skip to content

gzip/Brotli Compression Basics — How HTTP Responses Get Smaller, and Why the Server Gets to Choose

Fast-loading websites often share a trick you never see: the server compresses HTML, CSS, and JavaScript right before sending them, using a format like gzip or Brotli, and the browser decompresses them before rendering the page. None of this exchange shows up on screen, but it has a large effect on how much data actually crosses the wire. Following the last two posts on OGP and HTTP cache headers — both things that live in HTTP response headers rather than <head> — this one covers how compression works.

Note: gzip has been around since 1992 and is supported almost everywhere, by both browsers and servers. Brotli is newer, published by Google in 2015, and tends to compress the same content more tightly than gzip. Think of them as two different implementations of the same problem — making data smaller — and today’s major browsers can decompress both.

What Compression Actually Does — Replacing Repetition With Shorter Codes

Both gzip and Brotli are built on the ideas behind compression algorithms like DEFLATE (which gzip uses directly). Simplified, the process has two stages:

  1. Replace a repeated string with a reference back to where it appeared before (a dictionary-style substitution)
  2. Assign shorter bit sequences to frequently occurring patterns, and longer ones to rare patterns (the idea behind Huffman coding)

HTML, CSS, and JavaScript are full of repetition — tag names, function names, indentation whitespace — so these two steps alone shrink them considerably. Data that’s already close to random (compressed images and video, covered below) has little repetition left to exploit, so compressing it again barely helps.

Brotli adds one more thing on top of this same foundation: a built-in dictionary of words and HTML/CSS boilerplate that’s common across the web. That built-in dictionary is why Brotli tends to beat gzip specifically on web content.

The Server and Browser “Agreement” — Accept-Encoding and Content-Encoding

Sending compressed content is pointless if the receiver can’t decompress it. To avoid that mismatch, the browser and server check compatible formats on every request using two headers:

Header Sent by Role
Accept-Encoding Browser → server Lists the formats the browser can decompress (e.g. gzip, deflate, br)
Content-Encoding Server → browser States which format was actually used (e.g. br)

The server picks one of the formats it supports from the Accept-Encoding list and returns the response compressed that way. The server makes the final call on which format gets used — the browser only offers candidates.

Sending different Accept-Encoding values to wpmm.jp/blog makes this selection visible in practice.

# Offer both gzip and br → br is chosen
curl -s -D - -o /dev/null -H "Accept-Encoding: gzip, br" https://wpmm.jp/blog/ \
  | grep -i content-encoding
# content-encoding: br

# Offer gzip only → gzip is chosen
curl -s -D - -o /dev/null -H "Accept-Encoding: gzip" https://wpmm.jp/blog/ \
  | grep -i content-encoding
# content-encoding: gzip

# Request no compression (identity) → no Content-Encoding header at all
curl -s -D - -o /dev/null -H "Accept-Encoding: identity" https://wpmm.jp/blog/ \
  | grep -i content-encoding
# (no output)

This blog runs on Xserver, and the web server in front of it (nginx) dynamically picks between Brotli, gzip, and no compression depending on Accept-Encoding, exactly as this exchange shows.

What Happens Without Vary: Accept-Encoding

When a CDN or reverse proxy cache sits between the server and the browser, there’s one more header that matters: Vary: Accept-Encoding.

It tells the caching layer, “treat the same URL as a different cache entry when Accept-Encoding differs.” Without it, a cache can serve an uncompressed response — cached from an earlier client that asked for no compression — to a later client that supports gzip, or the reverse: a compressed response ending up in front of a client that can’t decompress it, which shows up as garbled text. This is the same underlying issue covered in the OGP post — a crawler’s own cache hanging onto a stale result. A caching layer assumes “same URL means same content,” so any axis that can actually vary — language, compression format, and so on — needs to be declared explicitly. That principle applies here too.

The Same Mechanism, Already in This App’s Own Code

HTTP compression might sound unrelated to a desktop app, but this app’s own codebase already uses the same family of algorithm. The backup export feature (site_manager_web.py::backup_export), which bundles settings and site data into a downloadable file, builds that file with Python’s zipfile module and specifies the compression method like this:

with zipfile.ZipFile(buf, 'w', zipfile.ZIP_DEFLATED) as zf:
    # bundles sites_*.json / settings.json / server_profiles.json
    ...

zipfile.ZIP_DEFLATED tells the ZIP writer to compress its contents using the DEFLATE algorithm. gzip is DEFLATE with a header and checksum wrapped around it — so compressing an HTTP response with gzip and compressing the contents of a ZIP file are, underneath a different container format, nearly the same algorithm. The use cases couldn’t be more different, but the underlying idea — replace repetition with shorter codes — is shared.

When Compression Doesn’t Help

Compression isn’t universally useful. Compressing data that’s already compressed rarely shrinks it further, and it still costs CPU time for essentially no benefit.

  • Images (JPEG, PNG, WebP): the image format already applies its own compression
  • Video (MP4 and similar): same story
  • Files already gzipped or zipped: little headroom left to compress further

For this reason, a web server’s compression settings typically target text-based MIME types only — HTML, CSS, JavaScript, JSON, SVG — and exclude images, video, and already-compressed files.

Dynamic vs. Pre-Compression: A Trade-off

There’s one more design choice worth knowing about: compress on the fly, for every request (dynamic compression), or compress once ahead of time and just serve the result (pre-compression)?

Brotli exposes a “quality” setting — higher quality means better compression but more CPU time. A server compressing dynamically on every request usually favors a lower quality setting to keep response times fast. Static assets that rarely change (a distributed CSS or JS bundle, for instance) can instead be compressed once, at maximum quality, at deploy time, and served as-is afterward — paying the compression cost once instead of on every request. This is the same “compute it once, reuse the result” idea covered in the earlier post on HTTP cache headers, applied to a different problem.

Common Pitfalls

  • Compressing already-compressed formats (images, video): burns CPU for essentially no size reduction
  • Forgetting Vary: Accept-Encoding: a caching layer ignores the difference in compression format and serves the wrong one
  • Confusing gzip and Brotli compression-level numbers: gzip runs 1–9, Brotli runs 0–11 — the scale and meaning differ between the two algorithms
  • Brotli not being offered over plain HTTP: major browsers often omit Brotli from the Accept-Encoding candidates on non-HTTPS connections for security reasons, leaving gzip as the only option

Summary

gzip and Brotli are both HTTP response compression formats built on the same foundation — replacing repetition with shorter codes — with Brotli’s built-in dictionary giving it an edge on typical web content. Which one gets used is decided by the server, chosen from the candidates the browser lists in Accept-Encoding, and adding Vary: Accept-Encoding keeps caching layers from serving the wrong version. Skipping compression on data that’s already compressed, and choosing between dynamic and pre-compression based on how often the content changes, are the two judgment calls that matter most in practice.