Instant Large-Image Decoding

The first thing you notice in an image viewer is first-screen latency — the tens to hundreds of milliseconds between double-clicking and seeing the image. Almost all of that time goes to "decoding": turning the compressed bytes on disk back into pixels the screen can draw.

GuoheView doesn't just plug in one off-the-shelf decoding library and call it done. It picks the fastest path for each format, and has in-house parallel optimizations for the most common ones, JPEG and TIFF. This article is only about decoding.

How an image becomes pixels

Take JPEG as an example. Decoding happens in roughly three steps, like a pipeline:

① Entropy decoding Huffman bitstream ② Dequant + IDCT frequency → spatial ③ Color conversion YCbCr → RGB Inherently serial SIMD-parallel SIMD-parallel

The last two steps (IDCT and color conversion) are independent per block of pixels and naturally suit SIMD parallelism. The hard part is step one, entropy decoding: it is a variable-length bitstream, where the position of the next value depends on how long the previous one was, so it seems it can only be read bit by bit in order. As we'll see, this is exactly where decoders differ.

JPEG: three decoding paths

WIC, built into Windows

WIC (Windows Imaging Component) is the image codec framework built into Windows, provided as COM components. Its strength is zero dependencies and it decodes nearly everything (JPEG / PNG / GIF / BMP / TIFF…) — one system API call gets you pixels.

The price is being "general-purpose": it has to cover every format and every scenario, so it can't be tuned to the limit for any one format, and it won't make full use of your multi-core CPU on its own. That's why in GuoheView WIC is only the fallback — formats without a dedicated fast path (BMP, GIF, ICO, etc.) go to it, where it's simple and reliable.

libjpeg-turbo: speed through SIMD

libjpeg-turbo is the industry-standard JPEG library. Its key improvement over the original libjpeg is rewriting per-pixel hot spots such as IDCT, upsampling and color conversion with SIMD instructions (SSE2 / AVX2) — one instruction processes 8 or 16 values at once instead of one at a time.

Scalar one at a time, N iterations SIMD one instruction, a whole batch

But libjpeg-turbo's SIMD only speeds up the last two steps of the pipeline. Entropy decoding is still single-threaded and serial — unless the JPEG file contains RST (restart) markers that split the bitstream into independent segments. The problem is that the vast majority of JPEGs (from phones, cameras, screenshots) have no RST markers, so entropy decoding becomes a bottleneck that multi-core CPUs can't share.

In-house pjdec: parallelizing serial entropy decoding too

This is the core of GuoheView's decoding engine. We wanted to answer one question: for an ordinary JPEG without RST markers, can entropy decoding use all the cores?

Intuitively, no — the bitstream is variable-length, and each block's brightness (DC coefficient) is stored as a difference from the previous block, so everything is chained together. But we use a "speculative parallel" approach:

Traditional: one thread, in order . One core reads start to finish; the other cores sit idle pjdec: split into segments, each core decodes "blind" from its start Core 1 Core 2 Core 3 Core 4 Dashed = byte-aligned blind entry point (position may be wrong) ● = sync point where decoding "locks on" after a few blocks; bit-exact from here on Verified → keep results after the sync point; if it never matches → decode that part in order

It relies on a remarkable property of Huffman coding — self-synchronization: even if you start decoding at the wrong position, after a few codewords the decoder automatically "lines up" with the real codeword boundaries. So each core starts decoding blindly from the byte boundary at the start of its segment, producing results whose "position may be wrong" at first, and records where each data block starts (checkpoints).

Then a lightweight serial check follows: if the real end position of the previous segment is exactly equal to one of the next segment's checkpoints, the determinism of the encoding guarantees that from that point on the two are bit-for-bit identical, so the next segment's results are used directly. As for the chained DC differences, blind decoding starts counting from 0 and differs from the true value only by a constant; after synchronization one correction is added across the board (using 16-bit integer wrap-around, at zero cost).

If a segment never lines up (very rare), that small part is simply decoded in order — correctness always has a safety net, just with a little less parallelism. The result: even for an ordinary large JPEG, entropy decoding uses all the cores, and the first screen appears noticeably faster.

TIFF: two decoding paths

libtiff

libtiff is the de facto standard library for TIFF, with the most complete format coverage (every kind of compression, CMYK, multi-page). GuoheView uses it as the foundation for compatibility — any TIFF libtiff can read, we can open. Its limitation is the same: it processes data blocks (strips / tiles) one by one on a single thread.

In-house tiff_mt: parallel by block

TIFF's structure is actually made for parallelism — the image is cut into many independent strips or tiles. Our tiff_mt decoder adds multithreading on top of that:

TIFF tiles / strips are independent Core 1Core 2Core 3Core 4 Core 1Core 2Core 3Core 4 Adaptive thread count: decided by core count, compression and workload One large JPEG block → reuse pjdec in-stream parallelism

It goes two steps further than "just throw more threads at it":

  • Adaptive thread count: it doesn't blindly use every core. Uncompressed TIFF is limited by memory bandwidth, where more threads just fight over bandwidth; and when the workload is too small, threads aren't worth starting — the count is decided dynamically from the core count, compression type and data size.
  • Two kinds of parallelism working together: when a TIFF uses JPEG compression but has only a few large blocks, plain "block-level parallelism" would leave most cores idle. In that case tiff_mt switches to in-stream parallelism — directly reusing pjdec's speculative parallelism described above to split up that large JPEG stream. This is where the two in-house decoders connect.

Our strategy

Inside GuoheView there is a decoder registry. Each decoder declares the format signatures it recognizes (file header magic) and a priority. When a file is opened, the engine matches the file header precisely and picks the highest-priority dedicated decoder: JPEG goes to pjdec, TIFF to tiff_mt, and only when nothing matches does it fall back to WIC.