Markus Oberlehner

The Fastest Website I Can Build


Recently, I had a “I Know Kung Fu” moment: I know Go. Actually, I don’t know Go. Yet, I recently realized that I don’t write any code by hand anymore and, with my personal projects, I barely read code anymore, too. So if I don’t read and write Node.js code anymore, I might very well not read and write Go code while reaping all its benefits in terms of speed and memory footprint.

So as a little PoC, I let Claude Code loose building a website rendering engine in Go to render the official Storyblok website, also supporting the Visual Editor Live Preview feature. What can I say: one prompt and 20 minutes later, I had a fully working website in Go. But not so fast. Tuning everything to the max still required pouring in a lot of time and knowledge, but as you’ll see: it was worth it.

If you’re not so much into details and the journey, here is the TL;DR:

Why not Go?

With writing code by hand becoming a thing of the past, choosing a more efficient language over a less efficient one should more and more become the norm.

All three stacks render the same pages from the same Storyblok content, as production builds on the same machine (an M2 Pro laptop), with Storyblok’s responses already in memory.

GoAstro 7Nuxt 2
Render time, one request at a time (median)1.5–3.0 ms4.3–4.4 ms12.9–19.1 ms
Requests per second, 50 at a time1,400–2,600240–25647–73
99th percentile at 50 concurrent requests62–137 ms247–456 ms4.8–9.0 s
Memory after start36 MB80 MB155 MB
Memory at peak, 50 concurrent requests208–225 MB435–455 MB1,476–1,644 MB
Start to first rendered page0.4–0.6 s0.5 s0.8–1.0 s

It turns out that the Go version of my website, petermax.at, renders a page 1.6 to 2.8 times faster than Astro and 4 to 6 times faster than my legacy Nuxt 2 stack. Under load the gap widens: Go serves 6 to 11 times as many requests per second as Astro and 22 to 32 times as many as Nuxt, while using half the RAM of Astro and a seventh of Nuxt’s.

But to be completely honest with you, except for the RAM savings, which are nice in a world of tripling RAM prices, for the website visitors this does not matter much. Why? Because of the following optimization.

Caching with nginx and stale-while-revalidate

Actually, when you hit petermax.at right now, you’ll not get a response from the Go server. Instead, most likely you’ll get a response right from the cache of nginx.

Caveat: as petermax.at is mostly targeting an Austrian audience, I’ve decided that we don’t need a CDN. This means that if you’re reading this from outside of Central Europe, you might not experience the full speed of this stack.

Yet caching is a double-edged sword: yes, users might get a super fast response, but what about cache busting in order to ensure what they see is not stale information? What pattern you use to deal with this depends on the freshness requirements of your content and if you expect to get regular hits on content that needs to be fresh or not. In our case the perfect solution is the stale-while-revalidate (SWR) pattern: a page is fresh for 10 seconds. The first request after that still gets the stale copy, at cache speed, and makes nginx fetch a fresh one in the background. The next request gets the fresh page. This also means that almost every request is a fast cache hit.

proxy_cache pages;
proxy_cache_valid 200 10s;
proxy_ignore_headers Cache-Control Expires;
proxy_cache_use_stale error timeout updating http_500 http_502 http_503 http_504;
proxy_cache_background_update on;
proxy_cache_lock on;

The Go server tells browsers to always revalidate a page (Cache-Control: no-cache), so nginx has to ignore that header and apply its own 10 seconds. The error timeout http_5xx part is a nice bonus: while the Go server restarts for a release, nginx keeps answering from its cache.

RequestResponse time (median)
Go server through nginx, not cached5.4–5.7 ms
nginx cache hit0.9–1.1 ms
First request after expiry (stale, refreshing)as fast as a hit
The request after thathit, with the new page

I measured this locally with nginx in a container, which is why the uncached request is slower here than the 3 ms above: the hop into the container alone costs about 2.4 ms. To see what happens when a page expires under load, I sent 50 requests at once at an expired page: one got the stale copy and triggered the refresh, and the other 49 got the stale copy while the refresh ran. Nobody waited for the Go server.

On the real site, from my desk, the HTML starts to arrive after 115 to 130 ms (median of 20 requests). About 110 ms of that is the connection and TLS handshake. 19 of the 20 requests were cache hits.

A few more things on the server side:

Serving assets, fast

Now we’ve optimized our server response times and caching. That’s great, but as everybody who has ever worked on improving the loading performance of a website knows: this only scratches the surface. The most important factor when it comes to website loading performance are assets. Mostly images and JavaScript, to some degree also CSS.

Image optimizations

The Storyblok Image Service is already a good start, but unfortunately its format(auto) mode does not support AVIF yet: a browser that announces AVIF support still gets WebP. So this is one of the more impactful optimizations I made: my own proxy service that detects AVIF support from the browser’s Accept header and asks Storyblok for format(avif) explicitly.

The same images as the site requests them for a phone with a 2x display, at the same quality setting:

ImageJPEGWebP (format(auto) today)AVIFAVIF vs. WebP
Hero photo, 960×54035.0 KB23.3 KB15.9 KB−32%
Kitchen photo, 960×58242.6 KB37.6 KB26.6 KB−29%
Kitchen tile, 1500×1080112.4 KB76.5 KB48.2 KB−37%
Graphic (PNG), 640×42411.1 KB5.6 KB3.8 KB−33%

Because the proxy decides the format, we don’t need <picture> with a <source> per format, just one <img>:

<img
  src="/media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/666x404/filters:quality(80):format(jpg)"
  srcset="
    /media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/240x146/filters:quality(40):format(auto)    240w,
    /media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/666x404/filters:quality(40):format(auto)    666w,
    /media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/1332x808/filters:quality(40):format(auto)  1332w
  "
  sizes="auto, (min-width: 64rem) 50vw, 100vw"
  width="666"
  height="404"
  alt=""
  loading="lazy"
  decoding="async"
/>

(The real srcset has six candidates.) sizes="auto" lets the browser pick the candidate from the size the image really has in the layout, which works for lazy-loaded images.

Low quality at double density

Next, we want to make sure to squeeze the quality as much as possible while also ensuring that the images are crisp on high resolution displays.

I found that, with AVIF, you can go as low as quality 40 at 2x pixel density. An image at twice the resolution and quality 40 costs about the same bytes as the regular 1x image at quality 80, and looks sharper. Three photos, shown 640 CSS pixels wide:

Image1x, quality 802x, quality 802x, quality 40
Hero photo28.9 KB, SSIM 0.9390.6 KB, SSIM 0.9925.6 KB, SSIM 0.96
Kitchen tile40.6 KB, SSIM 0.89155.5 KB, SSIM 0.9837.3 KB, SSIM 0.95
Kitchen photo78.1 KB, SSIM 0.81326.9 KB, SSIM 0.99102.4 KB, SSIM 0.94

SSIM compares each to the uncompressed 2x image. 1 means identical. 2x at quality 80 is three to four times the bytes for a difference that is hard to see.

Bonus: remove metadata

While measuring, I found out that Storyblok does not remove metadata from AVIF. Exif, XMP, color profiles. Depending on the images, this can have severe consequences:

Logo, 214 px wideAVIF from StoryblokMetadata in itAVIF without metadata
Gorenje49.9 KB98%1.2 KB
Glas Berger27.7 KB96%1.0 KB
Schachermayer28.9 KB91%2.7 KB
Kaindl16.8 KB93%1.2 KB

So my proxy also strips metadata.

Optimizing videos

Unfortunately, Storyblok does not do video optimization. We could use a service like Cloudinary, and I did in the past. But I prefer not having to declare bankruptcy, so using Cloudinary is out of the question.

Instead, I’ve built my own video optimization service which squeezes the last byte out of a video while keeping its quality. It does the conversion in three passes:

  1. A very quick H.264 encode, scaled to the requested size (320p for phones, 720p otherwise). The first visitor waits for it, and it takes seconds, so nobody ever gets the gigantic original.
  2. A slow, better H.264 encode in the background.
  3. An AV1 encode in the background. It takes a while, but the results are worth it.

Passes 2 and 3 only start when the service has no visitor waiting for a first pass, and they run at the lowest CPU priority, so a first pass arriving in the meantime gets nearly the whole CPU.

The page offers both, AV1 first:

<video preload="none" playsinline muted loop poster="…">
  <source src="/media/videos/f/46288/x/3bb4f75c33/wagner-headervideo_1.mp4/720.av1.mp4" type='video/mp4; codecs="av01.0.08M.08"' />
  <source src="/media/videos/f/46288/x/3bb4f75c33/wagner-headervideo_1.mp4/720.mp4" type='video/mp4; codecs="avc1.64001F"' />
</video>

As long as the AV1 file does not exist, the service answers that request with a “not yet” and the browser moves on to the H.264 source.

The hero video of the home page, 18 seconds at 720p, on my laptop:

Pass720p: encode time720p: file size320p: encode time320p: file size
Original upload2.26 MB
1: H.264, fast2.4 s1.45 MB0.6 s390 KB
2: H.264, slow5.1 s2.07 MB1.2 s563 KB
3: AV13 min 24 s0.88 MB52 s334 KB

Pass 2 is larger than pass 1 on purpose: the fast pass trades quality for speed, the slow one buys the quality back. This upload was already a web-sized 720p file, so H.264 cannot gain much on it. The gain comes from AV1, which is 61% smaller than the original, and from the 320p variant phones get, which is 85% smaller. On the server the encoder has a single CPU core instead of my laptop’s ten, so it takes longer there: the 720p AV1 pass used over 7 minutes of CPU time, the first pass 10 seconds.

So users visiting a site immediately after a new video was uploaded have bad luck, they see an only lightly optimized version. But hey, still better than the original that was uploaded by the content editor. The vast majority of people will see the perfectly optimized AV1 video.

PDF compression

Optimizing images and videos for websites is a classic. But often overlooked, while still important, at least for petermax.at, are PDFs.

While there are existing services for optimizing images and videos for websites via URL parameters, I could not find anything for PDFs. So again, I built it on my own, using Ghostscript behind the scenes. Instead of linking the Storyblok asset, the site links the same path on my service:

https://a.storyblok.com/f/46288/x/a0f961a02a/winterrezept_01_2018.pdf
https://www.petermax.at/media/pdfs/46288/x/a0f961a02a/winterrezept_01_2018.pdf

Ghostscript downsamples the embedded images to 144 dpi (96 dpi for sources above 100 MB) and recompresses them, then qpdf rewrites the file so the first page shows before the rest has arrived. A PDF link has no fallback like a video has, so the very first request gets the original while the optimization runs, and every request after that the small file.

The results are even more impressive:

PDFOriginalOptimizedSaved
Catalogue, print export, 16 pages75 MB2.4 MB97%
Catalogue, print export289 MB14 MB95%
Sustainability report, 2 pages10.5 MB756 KB93%
Recipe, 2 pages2.6 MB511 KB80%
Care instructions, 5 pages3.2 MB1.27 MB61%
Brochure, 80 pages9.4 MB5.8 MB38%

Before, I required content editors to optimize PDFs by hand. Which they mostly didn’t do. So this is a significant improvement for content editors and users.

JavaScript and CSS

When I took over petermax.at I built it with what I was most hyped about at the time: Nuxt.js 2. On the one hand, this was a great choice for my personal development because it made me build an open source project that got somewhat famous and even mentioned at one Google I/O, but from a technological standpoint, it was the completely wrong choice. Even current Nuxt, Next.js, and similar might be great choices for building web apps, but for anything that is more a marketing page than an app, I think they’re a disaster.

The big reason why: JavaScript. No matter how hard I tried, I never got decent Lighthouse numbers out of the Nuxt.js 2 setup because of the amount of JavaScript it ships and all the client-side hydration. For apps this all might be worth it, but for content-heavy websites it is not.

Disclaimer: CPUs got a lot faster, so JavaScript execution is less of an issue than it was a couple of years ago. Yet, normal people still have cheap, old phones surfing the web. And even more importantly: why make it slow when you can make it fast?

So my first choice for a replacement was Astro, which serves no JS at all by default. And with my Go stack I do the same: everything works without JavaScript. Yet, we add some JavaScript in the form of htmx and lightweight custom scripts for progressive enhancement. Each page gets one small bundle with only the scripts of the components it renders, and htmx is loaded only on pages that use it.

PageStackJavaScriptScript startup timePerformance score
HomeNuxt266 KB, 20 files305 ms92
HomeGo13 KB, 4 files34 ms97
Kitchen pageNuxt265 KB, 23 files324 ms88
Kitchen pageGo28 KB, 5 files32 ms99
Service pageAstro31 KB, 19 files35 ms99
Service pageGo26 KB, 5 files53 ms99
Colors pageAstro29 KB, 14 files35 ms99
Colors pageGo12 KB, 4 files59 ms99

(Lighthouse on mobile settings, median of five runs, JavaScript as transferred with brotli)

Compared to Nuxt, this is a different world: a twentieth of the JavaScript on the home page and a tenth of the startup time. Compared to Astro, it is a tie. Go ships a bit less JavaScript in fewer files, while Astro spends a bit less time running it (I did not investigate this further, but if I had to guess, my money would be on htmx v2 vs. v4 in Go).

On the CSS side I do nothing special except for inlining CSS, which pretty much has become the norm.

Wrapping it up

With AI we have to rethink our technology choices: it is not very important anymore how ergonomic we find it to work with a particular tech stack. What is more important is the result. What can a particular tech stack do for us and how well does it perform, not during development but when deployed on a server.

Furthermore, we should always think twice if we actually need a ready-made framework or if a solution custom-tailored to our requirements is the better choice. In my case, it is.