Performance

The previous article covered how compressed bytes are decoded into pixels quickly. This one continues: once decoded, how can an image with hundreds of millions of pixels still browse smoothly — with dragging and zooming that keep up with your hand?

How big can an image be

A casual phone shot is 12 megapixels and a DSLR 24 megapixels, but a satellite image can reach 18641 × 18641 ≈ 350 megapixels. If you dutifully decoded all of it into memory:

Decode everything into memory ≈ 1.4 GB 350 MP × 4 bytes What the screen can show ≈ 8 MP 4K = 3840 × 2160 ≫

1.4 GB of memory and seconds of waiting — just to show an image shrunk to fit the screen, where the eye can't make out individual pixels anyway. Most pixels never get a chance to be seen. GuoheView's approach is to spend effort only at the detail level and in the area actually needed. Three things work together for this: the pyramid, tiles, and on-demand loading.

Pyramid: decode at the detail level you need

Halve the original image again and again and store the results as a stack of layers from large to small — that's an image pyramid. The bottom layer is the original size, each layer above has half the side length, until it fits into a 256×256 thumbnail. An 18641² image has about 8 layers.

Level 0 — 18641² original Level 1 — 9320² Level 2 … Top level ≤ 256² The smaller you zoom out → the higher the layer used The more you zoom in → the lower the layer used

When you open an image with "fit to window", the whole image is shrunk to the screen and there is no need for the 18641² original data at all — one of the middle layers is displayed directly, which is both fast and light. When you zoom in to look at details, it switches to a lower, sharper layer. Always use the layer that is just sharp enough — that's the point of the pyramid. Because each extra layer is only a quarter the size of the one below it, all of them together add only about a third on top of the original, a small price.

Tiles: decode only the part you can see

A pyramid alone isn't enough — when you zoom in for detail you use the bottom layer, which still has hundreds of millions of pixels. So every layer is further cut into 256×256 small squares (tiles), and only those that fall on the screen are processed.

Screen viewport Tiles in the viewport: decoded Tiles outside: left for later 256×256 each; a full 4K viewport is only about 135 tiles

Filling a whole 4K viewport is only about 135 tiles of work — a completely different order of magnitude from "hundreds of millions of pixels". Tiles can also be decoded in parallel: the hundred-odd tiles are spread across several CPU cores at once, so they land faster.

On-demand loading: decode wherever you drag

Put the pyramid and tiles together and you get a "viewport-driven" loading mechanism. Each time you zoom or pan, the canvas works out which layer and which tiles the current viewport covers, hands the not-yet-decoded tiles to background threads, and draws each tile as soon as it's ready.

Zoom / pan viewport changes Find visible tiles layer + intersect Background decode multi-core parallel Draw when ready tile by tile keep interacting → repeat

All decoding happens on background threads, so the interface never freezes; tiles that have already been decoded are cached, so dragging back doesn't decode them again. However large the image, what's actually being decoded at any moment is only that small patch of the screen.

See it first, then sharpen

There's one last experience optimization. The moment you switch to a large image, the visible tiles aren't decoded yet, so the canvas first lays down a low-resolution preview of the whole image as a fallback — you see the complete picture immediately, and then sharp tiles cover it one by one. Sharpness only ever increases.

There's an easy-to-miss detail here: the preview and the tiles are two separately drawn layers, and if the switch-over timing is wrong you can get a regression of "sharp preview → briefly blurrier → sharp again". We specifically guarantee this monotonically sharpening path, so the whole process never flickers backwards.

Putting it together

  • The pyramid decides "how detailed the data is" — the layer is chosen by display zoom, so nothing is wasted on detail you can't see;
  • Tiles decide "how large an area is decoded" — only the patch on screen is touched;
  • On-demand loading + background parallelism decide "when to decode" — decode wherever you drag, without freezing the interface;
  • The low-res preview fallback decides "what you see first" — an image right away, then progressively sharper.

Together, memory and time are spent only on the pixels you can actually see. That's why an image with hundreds of millions of pixels can still be dragged and zoomed freely, keeping up with your hand.