bytes & pixels
Starter draft. This grew out of the garage encoding study, where one photo runs through every encoder and the bytes get counted. The conversation is a first pass; the demo fetches the study's real output files and measures their actual sizes.
the garage page runs one photo through jpeg, zenjpeg, and avif and prints three different sizes. before I trust those numbers, tell me why one image even gives three sizes.
An encoder throws away the detail your eye skips, then packs what survives. The three differ in how hard they push each step.
JPEG (1992) cuts the photo into 8x8 blocks, runs a cosine transform, and rounds off the high-frequency parts. jpegli (Google, 2024) keeps the exact JPEG format but rounds smarter, so it fits roughly 25% more quality into the same bytes as the old mozjpeg encoder. AVIF borrows AV1's video tricks, larger blocks and sharper prediction, so it usually wins on bytes-per-pixel outright.
so jpegli writes a normal .jpg that any browser opens, just smaller?
Right. jpegli stays inside the JPEG bitstream, so a browser from 1995 still decodes it. AVIF needs a modern decoder (Safari 16+, Chrome 85+). That trade, universal but larger against modern but smaller, is why the photo grid ships both and lets the browser pick.
ok, stop describing it. show me the bytes on one photo.
The same 400×266 photo (106,400 pixels), run through every encoder. These sizes are fetched live from the files the garage study actually produced, so the bytes are real. The last column is size against the lossless PNG.

| encoder | size | b/px | vs PNG |
|---|---|---|---|
| measuring real files… | |||
so on this photo jpegli is about half the baseline JPEG, and avif is smaller still. what's the catch with avif then?
Decode support and encode time. AVIF leans on the AV1 codec, which is slower to encode and only decodes on recent browsers. A tuned JPEG is fast and universal. That is why the site ships a tuned JPEG encoder for the universal fallback and uses AVIF as the primary <picture> source: the browser reaches for AVIF, and anything that cannot decode it falls back to the .jpg. That fallback encoder was jpegli for a long stretch, built from source; since 2026-07 it is zenjpeg, which the grids below measure against it directly.
and the grayscale shots? why are those so much smaller?
A color image carries one luma plane (brightness) plus two chroma planes (color). Drop to grayscale and you delete both chroma planes, roughly two-thirds of the color information, so the file falls hard. The garage study has the side-by-side counts.
so color is two of the three planes. JPEG and AVIF don't even keep color at full resolution, right? that's the subsampling thing?
Right, chroma subsampling. Your eye resolves brightness far better than color, so codecs keep luma at full resolution and shrink the two chroma planes. 4:4:4 keeps everything; 4:2:2 halves chroma horizontally; 4:2:0, the web default, stores one chroma sample per 2x2 block, a quarter of the color resolution. The luma carries the sharpness, so you barely notice.
The top band is fine luma detail (black and white lines); the bottom is fine chroma detail (red and green lines) at the same spacing. Switch the mode: luma stays razor sharp at every setting, while the color detail dissolves as you drop chroma resolution.
wild, the black-and-white lines stay razor sharp at 4:2:0 but the red-green lines just dissolve. and that's half the samples gone.
Exactly, and for photographs it is almost free, because real scenes rarely put fine high-contrast color edges right next to each other. Where it shows is red text on a dark background or saturated line art, which is why screenshots and graphics keep 4:4:4 and 4:2:0 became the default for photos.
earlier you said jpegli just "rounds smarter." smarter how, if it writes the same JPEG format?
It models your eye instead of trusting the fixed 1992 quantization tables. jpegli works in the XYB color space, which spaces colors the way you perceive them, sets the quantization adaptively per block by what you would actually notice, and rounds with adaptive dead-zones. The output is still a standard JPEG that any decoder reads, but more bits land where your eye looks and fewer on detail you would never see. That is the rough 25% it gains over the old mozjpeg encoder.
The site rode jpegli for exactly that reason, then moved to zenjpeg in 2026-07. It reaches the same goal by a different road: hybrid trellis quantization plus a search across 64 candidate progressive scan scripts, which lands a few percent under jpegli at matched quality.
ok, I want to SEE these tradeoffs, not just read byte counts. show me one zoomed crop across formats and qualities.
Here is a 96-pixel slice of the car's front wheel, silver spokes over a red caliper, run through every format and blown up so the artifacts show. Start with format against quality.
One 96-pixel slice of the front wheel (spokes, bolts, a red caliper), run through three formats at three quality tiers and shown pixel-zoomed. Read across a row to compare formats at one quality; read down to watch a format fall apart as quality drops.

…

…

…

…

…

…

…

…

…
The slice has real detail in it, so the per-pixel efficiency shows even at 96 pixels: AVIF comes out smallest at all three tiers and JPEG largest, the same order the full-photo table above measures. What the zoom adds is the artifact style: JPEG breaks the spokes into 8x8 blocks, while WebP and AVIF smear and smudge them instead.
jpeg goes blocky, the others go smeary. now the jpeg encoders you mentioned, mozjpeg vs the google one?
This one uses a darker crop of the same car, the body edge, because its jpegli cell was encoded once before that encoder left the site's toolchain and cannot be re-cut. Same quality knob, four encoders in the order the site adopted them. The bytes under each tell the story.
One crop of the body edge, one quality setting (q72), four JPEG encoders in the order the site adopted them. All four write the same JPEG format any decoder reads; the only difference is how cleverly each spends its bits. Watch the byte counts.

…

…

…

…
Same quality knob, real bytes measured live, and each step buys its win a different way. mozjpeg trellis-quantizes inside the 1992 rules. jpegli throws those rules out and models your eye instead (XYB color, adaptive per-block quantization), which is the psychovisual win in one crop. zenjpeg pairs a hybrid trellis with a search across 64 candidate progressive scan scripts. jpegli is the encoder that proved a standard JPEG could be halved; zenjpeg is the one the site ships, because it landed a few percent under jpegli at matched quality.
so each one squeezes a little harder than the last, same setting throughout. and chroma, up close on a real photo?
A different slice for this one. The body-edge crop is nearly all black and gray, and subsampling only touches color, so all three came out the same. These are bare branches against red glass from a night frame, cut at thumbnail scale and saved as a JPEG at the three chroma samplings. Luma holds; the color is what softens.
Chroma subsampling on a real photo instead of stripes: 96 pixels of bare branches lit against red glass, saved as a JPEG at 4:4:4, 4:2:2, and 4:2:0 at one quality setting. The branch shapes hold at every setting, because their brightness lives in luma. What goes is their color: at 4:2:0 the thin orange twigs turn brown.

…

…

…
4:2:0 comes out 42% smaller than 4:4:4 here (the sizes under each cell are measured live). Nearly everything in this slice is color detail, and the bytes subsampling saves are exactly the detail it throws away.
hold on. the garage page says the AVIF thumbnails here ship at 4:4:4, and the camera only records 4:2:2. how does storing more color than the camera kept help anything?
Because 4:2:2 describes the camera's frame, and a thumbnail is a much smaller picture. The Fuji frame is 5152 pixels tall; the 600px tile is its centre square shrunk 8.6x, so every tile pixel is an average of an 8.6 by 8.6 patch of the frame. Even at half-width color, that patch holds about four color samples across and nine down.
So at its own size the tile carries full color, measured separately for every pixel. 4:2:0 would average it back down to one color per 2x2 block, throwing away half in each direction.
Each dot is one color sample the camera recorded at 4:2:2 (one per two pixels across). Each cell is one pixel of the finished tile, filled with the average of the dots inside it. At 1x, pairs of cells share a dot. Drag past 2x and every cell owns its own color; then switch the tile to 4:2:0 and watch the 2x2 blocks flatten it again.
The shipped tiers sit at 8.6x (600px), 12.9x (400px) and 25.8x (200px) of the Fuji frame, so all three are far past the 2x line.
so under 2x there's nothing to gain, since two tile pixels still share one camera sample. past 2x every pixel owns its color, and 4:2:0 is the step that shrinks it again.
That's the model, and the garage page measured it. With no compression at all, halving a tile's color and restoring it scores 93.1 to 93.5 on SSIMULACRA2, against 96.3 for the same round trip on a crop at native size. The downscale is what makes subsampling expensive.
Keeping it costs almost nothing, because color in a shrunken photo is smooth and AV1 codes smooth planes cheaply. At the same file size, 4:4:4 scored higher on 18 of 24 photos at the 600px tier; 4:2:0 needs a median 2.2% more bytes to catch up.
[src]where does it matter most?
Under red light. Red carries only 0.299 of the brightness in the BT.601 matrix these files use, so in a red scene the fine detail lives in the color planes rather than in luma. One night frame of bare branches against red windows, the one the real-photo grid above is cut from, needs 35% more bytes at 4:2:0 to match 4:4:4.
the garage page keeps saying every jpg here is progressive. if the pixels come out the same in the end, what actually changes?
The order the bytes arrive in. Each 8x8 block becomes 64 numbers: one DC coefficient, the block's average, and 63 AC coefficients for finer and finer detail. Quantization rounds them. Everything after that is bookkeeping about how to write them down.
A baseline JPEG writes each block complete, left to right and top to bottom, in one pass called a scan. Half the file gets you the top half of the picture. A progressive JPEG spreads the same numbers across several scans, each covering the whole image. It slices two ways: by frequency (every block's DC first, which is the picture at 1/8 scale, then the coarse AC bands, then the fine ones) and by precision (the high bits of each coefficient first, the low bits in later refinement scans). The decoder adds every scan into one coefficient grid, so when the last scan lands it holds exactly the numbers the baseline file held.
[src]show me. same photo, cut off partway.
Three files holding the same quantized coefficients. jpegtran rewrote the shipped zenjpeg thumbnail in two other orders without decoding it, so all three finish as identical pixels. Drag the slider to choose how many bytes have arrived, and your browser decodes each cut-off file live. The table reads one file's scans out of its own headers.
16,324 B
16,609 B
16,234 B
| # | channels | carries | starts at |
|---|---|---|---|
| 1 | Y Cb Cr | DC (block averages), top bits only | 2% |
| 2 | Y | AC 1–5, top bits only | 12% |
| 3 | Cr | AC 1–63, top bits only | 23% |
| 4 | Cb | AC 1–63, top bits only | 27% |
| 5 | Y | AC 6–63, top bits only | 29% |
| 6 | Y | AC 1–63, refines bit 1 | 36% |
| 7 | Y Cb Cr | DC (block averages), refines bit 0 | 56% |
| 8 | Cr | AC 1–63, refines bit 0 | 58% |
| 9 | Cb | AC 1–63, refines bit 0 | 63% |
| 10 | Y | AC 1–63, refines bit 0 | 66% |
DC is each 8x8 block's average; AC numbers pick a band of the 63 detail coefficients, low numbers coarse and high numbers fine. Top bits only sends each coefficient at reduced precision, and refines bit adds one more bit of it later. Y is brightness, Cb and Cr the two color planes. The starred row is the first scan with every channel in play: Chrome paints nothing from a progressive JPEG before that scan's first data byte, exact to the byte. Other engines may draw sooner, and these panes show what yours does.
the zenjpeg one sits blank until nearly the end. isn't that the file the site ships?
It is, and that is the finding this demo was built to show. zenjpeg searches 64 scan orders for the smallest file, and on this photo the winner sends every bit of brightness first and both color planes last. Chrome waits for every channel to start before it paints anything, so the cutoff lands on the first data byte of the Cr scan, 84.8% of the way through. The order that wins on bytes paints once, at the end, like a slow baseline. Across the library, 156 of 255 color thumbnails wait past their halfway mark.
On a thumbnail that costs almost nothing: the whole file arrives in a round trip or two, and most browsers get the AVIF anyway. It matters on the 20 MB full-size downloads, which the garage study measures: 53 of 256 color originals there hold their color until the last few percent.
and the sizes? does progressive make the file smaller or bigger?
It depends on the script, and the three sizes above show it. A progressive AC scan can code a long run of all-zero blocks as one symbol, and each scan can carry Huffman tables fitted to its own statistics, which usually beats one table set for everything. Refinement scans pay some of that back in correction bits. On this photo libjpeg's standard 10-scan script comes out 1.7% larger than baseline, while zenjpeg's searched 7-scan order comes out 0.6% smaller. So a smaller progressive file is a property of a good script, and the search that found this one scored bytes alone.
so what does the site actually ship per thumbnail?
Each thumbnail is dual-encoded: an AVIF primary at 4:4:4 color and a zenjpeg JPEG fallback inside one <picture>, plus a 400px AVIF tier for phones. The browser loads the smallest format it can decode and never downloads the others.
got it. one photo, several encoders, and the grid hands each browser the cheapest thing it can read. thanks.
→ the full garage encoding study · back to Learning With Errors