The Fastest Website I Can Build
- #go ,
- #performance ,
- #ai
Recently, I had a “I Know Kung Fu” moment: I know Go. Actually, I don’t know Go. Yet, I recently realized that I don’t write any code by hand anymore and, with my personal projects, I barely read code anymore, too. So if I don’t read and write Node.js code anymore, I might very well not read and write Go code while reaping all its benefits in terms of speed and memory footprint.
So as a little PoC, I let Claude Code loose building a website rendering engine in Go to render the official Storyblok website, also supporting the Visual Editor Live Preview feature. What can I say: one prompt and 20 minutes later, I had a fully working website in Go. But not so fast. Tuning everything to the max still required pouring in a lot of time and knowledge, but as you’ll see: it was worth it.
If you’re not so much into details and the journey, here is the TL;DR:
- Go renders a page in 1.5 to 3 ms. That is 1.6 to 2.8 times faster than Astro and 4 to 6 times faster than Nuxt 2, with half the memory of Astro and a seventh of Nuxt’s.
- Caching with the stale-while-revalidate pattern, using nginx as a reverse proxy: a cache hit takes about 1 ms.
- No JavaScript hydration, 100% progressively enhanced, htmx loaded only on pages that use it: 13 to 28 KB of JavaScript per page instead of Nuxt’s 266 KB.
- Images:
- A proxy in front of the Storyblok Image Service for automatic AVIF support and metadata stripping. One partner logo went from 50 KB to 1.2 KB, because 98% of the file was metadata.
- Quality tuning: quality 40 at 2x for crisp images at the size of a regular 1x image.
- Videos: what the Storyblok Image Service does for images, but for videos. The first request gets a quick H.264 encode within seconds; in the background the service encodes AV1 for eventually much smaller videos.
- PDFs: automatic PDF minification, similar to the Storyblok Image Service. A 10.5 MB PDF becomes 756 KB.
Why not Go?
With writing code by hand becoming a thing of the past, choosing a more efficient language over a less efficient one should more and more become the norm.
All three stacks render the same pages from the same Storyblok content, as production builds on the same machine (an M2 Pro laptop), with Storyblok’s responses already in memory.
| Go | Astro 7 | Nuxt 2 | |
|---|---|---|---|
| Render time, one request at a time (median) | 1.5–3.0 ms | 4.3–4.4 ms | 12.9–19.1 ms |
| Requests per second, 50 at a time | 1,400–2,600 | 240–256 | 47–73 |
| 99th percentile at 50 concurrent requests | 62–137 ms | 247–456 ms | 4.8–9.0 s |
| Memory after start | 36 MB | 80 MB | 155 MB |
| Memory at peak, 50 concurrent requests | 208–225 MB | 435–455 MB | 1,476–1,644 MB |
| Start to first rendered page | 0.4–0.6 s | 0.5 s | 0.8–1.0 s |
It turns out that the Go version of my website, petermax.at, renders a page 1.6 to 2.8 times faster than Astro and 4 to 6 times faster than my legacy Nuxt 2 stack. Under load the gap widens: Go serves 6 to 11 times as many requests per second as Astro and 22 to 32 times as many as Nuxt, while using half the RAM of Astro and a seventh of Nuxt’s.
But to be completely honest with you, except for the RAM savings, which are nice in a world of tripling RAM prices, for the website visitors this does not matter much. Why? Because of the following optimization.
Caching with nginx and stale-while-revalidate
Actually, when you hit petermax.at right now, you’ll not get a response from the Go server. Instead, most likely you’ll get a response right from the cache of nginx.
Caveat: as petermax.at is mostly targeting an Austrian audience, I’ve decided that we don’t need a CDN. This means that if you’re reading this from outside of Central Europe, you might not experience the full speed of this stack.
Yet caching is a double-edged sword: yes, users might get a super fast response, but what about cache busting in order to ensure what they see is not stale information? What pattern you use to deal with this depends on the freshness requirements of your content and if you expect to get regular hits on content that needs to be fresh or not. In our case the perfect solution is the stale-while-revalidate (SWR) pattern: a page is fresh for 10 seconds. The first request after that still gets the stale copy, at cache speed, and makes nginx fetch a fresh one in the background. The next request gets the fresh page. This also means that almost every request is a fast cache hit.
proxy_cache pages;
proxy_cache_valid 200 10s;
proxy_ignore_headers Cache-Control Expires;
proxy_cache_use_stale error timeout updating http_500 http_502 http_503 http_504;
proxy_cache_background_update on;
proxy_cache_lock on;
The Go server tells browsers to always revalidate a page (Cache-Control: no-cache), so nginx has to ignore that header and apply its own 10 seconds. The error timeout http_5xx part is a nice bonus: while the Go server restarts for a release, nginx keeps answering from its cache.
| Request | Response time (median) |
|---|---|
| Go server through nginx, not cached | 5.4–5.7 ms |
| nginx cache hit | 0.9–1.1 ms |
| First request after expiry (stale, refreshing) | as fast as a hit |
| The request after that | hit, with the new page |
I measured this locally with nginx in a container, which is why the uncached request is slower here than the 3 ms above: the hop into the container alone costs about 2.4 ms. To see what happens when a page expires under load, I sent 50 requests at once at an expired page: one got the stale copy and triggered the refresh, and the other 49 got the stale copy while the refresh ran. Nobody waited for the Go server.
On the real site, from my desk, the HTML starts to arrive after 115 to 130 ms (median of 20 requests). About 110 ms of that is the connection and TLS handshake. 19 of the 20 requests were cache hits.
A few more things on the server side:
- Versioned assets: content-hash URLs, immutable for a year, but only when the version in the URL matches the file.
- Pages: an
ETag, so a browser revalidating a page gets an empty 304 response. - Storyblok responses and decoded stories: kept in memory, keyed by the space’s cache version. That is why a render takes 1.5 to 3 ms: no request leaves the server.
Serving assets, fast
Now we’ve optimized our server response times and caching. That’s great, but as everybody who has ever worked on improving the loading performance of a website knows: this only scratches the surface. The most important factor when it comes to website loading performance are assets. Mostly images and JavaScript, to some degree also CSS.
Image optimizations
The Storyblok Image Service is already a good start, but unfortunately its format(auto) mode does not support AVIF yet: a browser that announces AVIF support still gets WebP. So this is one of the more impactful optimizations I made: my own proxy service that detects AVIF support from the browser’s Accept header and asks Storyblok for format(avif) explicitly.
The same images as the site requests them for a phone with a 2x display, at the same quality setting:
| Image | JPEG | WebP (format(auto) today) | AVIF | AVIF vs. WebP |
|---|---|---|---|---|
| Hero photo, 960×540 | 35.0 KB | 23.3 KB | 15.9 KB | −32% |
| Kitchen photo, 960×582 | 42.6 KB | 37.6 KB | 26.6 KB | −29% |
| Kitchen tile, 1500×1080 | 112.4 KB | 76.5 KB | 48.2 KB | −37% |
| Graphic (PNG), 640×424 | 11.1 KB | 5.6 KB | 3.8 KB | −33% |
Because the proxy decides the format, we don’t need <picture> with a <source> per format, just one <img>:
<img
src="/media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/666x404/filters:quality(80):format(jpg)"
srcset="
/media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/240x146/filters:quality(40):format(auto) 240w,
/media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/666x404/filters:quality(40):format(auto) 666w,
/media/images/f/46288/1332x808/2a4ad2008a/wald.jpg/m/1332x808/filters:quality(40):format(auto) 1332w
"
sizes="auto, (min-width: 64rem) 50vw, 100vw"
width="666"
height="404"
alt=""
loading="lazy"
decoding="async"
/>
(The real srcset has six candidates.) sizes="auto" lets the browser pick the candidate from the size the image really has in the layout, which works for lazy-loaded images.
Low quality at double density
Next, we want to make sure to squeeze the quality as much as possible while also ensuring that the images are crisp on high resolution displays.
I found that, with AVIF, you can go as low as quality 40 at 2x pixel density. An image at twice the resolution and quality 40 costs about the same bytes as the regular 1x image at quality 80, and looks sharper. Three photos, shown 640 CSS pixels wide:
| Image | 1x, quality 80 | 2x, quality 80 | 2x, quality 40 |
|---|---|---|---|
| Hero photo | 28.9 KB, SSIM 0.93 | 90.6 KB, SSIM 0.99 | 25.6 KB, SSIM 0.96 |
| Kitchen tile | 40.6 KB, SSIM 0.89 | 155.5 KB, SSIM 0.98 | 37.3 KB, SSIM 0.95 |
| Kitchen photo | 78.1 KB, SSIM 0.81 | 326.9 KB, SSIM 0.99 | 102.4 KB, SSIM 0.94 |
SSIM compares each to the uncompressed 2x image. 1 means identical. 2x at quality 80 is three to four times the bytes for a difference that is hard to see.
Bonus: remove metadata
While measuring, I found out that Storyblok does not remove metadata from AVIF. Exif, XMP, color profiles. Depending on the images, this can have severe consequences:
| Logo, 214 px wide | AVIF from Storyblok | Metadata in it | AVIF without metadata |
|---|---|---|---|
| Gorenje | 49.9 KB | 98% | 1.2 KB |
| Glas Berger | 27.7 KB | 96% | 1.0 KB |
| Schachermayer | 28.9 KB | 91% | 2.7 KB |
| Kaindl | 16.8 KB | 93% | 1.2 KB |
So my proxy also strips metadata.
Optimizing videos
Unfortunately, Storyblok does not do video optimization. We could use a service like Cloudinary, and I did in the past. But I prefer not having to declare bankruptcy, so using Cloudinary is out of the question.
Instead, I’ve built my own video optimization service which squeezes the last byte out of a video while keeping its quality. It does the conversion in three passes:
- A very quick H.264 encode, scaled to the requested size (320p for phones, 720p otherwise). The first visitor waits for it, and it takes seconds, so nobody ever gets the gigantic original.
- A slow, better H.264 encode in the background.
- An AV1 encode in the background. It takes a while, but the results are worth it.
Passes 2 and 3 only start when the service has no visitor waiting for a first pass, and they run at the lowest CPU priority, so a first pass arriving in the meantime gets nearly the whole CPU.
The page offers both, AV1 first:
<video preload="none" playsinline muted loop poster="…">
<source src="/media/videos/f/46288/x/3bb4f75c33/wagner-headervideo_1.mp4/720.av1.mp4" type='video/mp4; codecs="av01.0.08M.08"' />
<source src="/media/videos/f/46288/x/3bb4f75c33/wagner-headervideo_1.mp4/720.mp4" type='video/mp4; codecs="avc1.64001F"' />
</video>
As long as the AV1 file does not exist, the service answers that request with a “not yet” and the browser moves on to the H.264 source.
The hero video of the home page, 18 seconds at 720p, on my laptop:
| Pass | 720p: encode time | 720p: file size | 320p: encode time | 320p: file size |
|---|---|---|---|---|
| Original upload | 2.26 MB | |||
| 1: H.264, fast | 2.4 s | 1.45 MB | 0.6 s | 390 KB |
| 2: H.264, slow | 5.1 s | 2.07 MB | 1.2 s | 563 KB |
| 3: AV1 | 3 min 24 s | 0.88 MB | 52 s | 334 KB |
Pass 2 is larger than pass 1 on purpose: the fast pass trades quality for speed, the slow one buys the quality back. This upload was already a web-sized 720p file, so H.264 cannot gain much on it. The gain comes from AV1, which is 61% smaller than the original, and from the 320p variant phones get, which is 85% smaller. On the server the encoder has a single CPU core instead of my laptop’s ten, so it takes longer there: the 720p AV1 pass used over 7 minutes of CPU time, the first pass 10 seconds.
So users visiting a site immediately after a new video was uploaded have bad luck, they see an only lightly optimized version. But hey, still better than the original that was uploaded by the content editor. The vast majority of people will see the perfectly optimized AV1 video.
PDF compression
Optimizing images and videos for websites is a classic. But often overlooked, while still important, at least for petermax.at, are PDFs.
While there are existing services for optimizing images and videos for websites via URL parameters, I could not find anything for PDFs. So again, I built it on my own, using Ghostscript behind the scenes. Instead of linking the Storyblok asset, the site links the same path on my service:
https://a.storyblok.com/f/46288/x/a0f961a02a/winterrezept_01_2018.pdf
https://www.petermax.at/media/pdfs/46288/x/a0f961a02a/winterrezept_01_2018.pdf
Ghostscript downsamples the embedded images to 144 dpi (96 dpi for sources above 100 MB) and recompresses them, then qpdf rewrites the file so the first page shows before the rest has arrived. A PDF link has no fallback like a video has, so the very first request gets the original while the optimization runs, and every request after that the small file.
The results are even more impressive:
| Original | Optimized | Saved | |
|---|---|---|---|
| Catalogue, print export, 16 pages | 75 MB | 2.4 MB | 97% |
| Catalogue, print export | 289 MB | 14 MB | 95% |
| Sustainability report, 2 pages | 10.5 MB | 756 KB | 93% |
| Recipe, 2 pages | 2.6 MB | 511 KB | 80% |
| Care instructions, 5 pages | 3.2 MB | 1.27 MB | 61% |
| Brochure, 80 pages | 9.4 MB | 5.8 MB | 38% |
Before, I required content editors to optimize PDFs by hand. Which they mostly didn’t do. So this is a significant improvement for content editors and users.
JavaScript and CSS
When I took over petermax.at I built it with what I was most hyped about at the time: Nuxt.js 2. On the one hand, this was a great choice for my personal development because it made me build an open source project that got somewhat famous and even mentioned at one Google I/O, but from a technological standpoint, it was the completely wrong choice. Even current Nuxt, Next.js, and similar might be great choices for building web apps, but for anything that is more a marketing page than an app, I think they’re a disaster.
The big reason why: JavaScript. No matter how hard I tried, I never got decent Lighthouse numbers out of the Nuxt.js 2 setup because of the amount of JavaScript it ships and all the client-side hydration. For apps this all might be worth it, but for content-heavy websites it is not.
Disclaimer: CPUs got a lot faster, so JavaScript execution is less of an issue than it was a couple of years ago. Yet, normal people still have cheap, old phones surfing the web. And even more importantly: why make it slow when you can make it fast?
So my first choice for a replacement was Astro, which serves no JS at all by default. And with my Go stack I do the same: everything works without JavaScript. Yet, we add some JavaScript in the form of htmx and lightweight custom scripts for progressive enhancement. Each page gets one small bundle with only the scripts of the components it renders, and htmx is loaded only on pages that use it.
| Page | Stack | JavaScript | Script startup time | Performance score |
|---|---|---|---|---|
| Home | Nuxt | 266 KB, 20 files | 305 ms | 92 |
| Home | Go | 13 KB, 4 files | 34 ms | 97 |
| Kitchen page | Nuxt | 265 KB, 23 files | 324 ms | 88 |
| Kitchen page | Go | 28 KB, 5 files | 32 ms | 99 |
| Service page | Astro | 31 KB, 19 files | 35 ms | 99 |
| Service page | Go | 26 KB, 5 files | 53 ms | 99 |
| Colors page | Astro | 29 KB, 14 files | 35 ms | 99 |
| Colors page | Go | 12 KB, 4 files | 59 ms | 99 |
(Lighthouse on mobile settings, median of five runs, JavaScript as transferred with brotli)
Compared to Nuxt, this is a different world: a twentieth of the JavaScript on the home page and a tenth of the startup time. Compared to Astro, it is a tie. Go ships a bit less JavaScript in fewer files, while Astro spends a bit less time running it (I did not investigate this further, but if I had to guess, my money would be on htmx v2 vs. v4 in Go).
On the CSS side I do nothing special except for inlining CSS, which pretty much has become the norm.
Wrapping it up
With AI we have to rethink our technology choices: it is not very important anymore how ergonomic we find it to work with a particular tech stack. What is more important is the result. What can a particular tech stack do for us and how well does it perform, not during development but when deployed on a server.
Furthermore, we should always think twice if we actually need a ready-made framework or if a solution custom-tailored to our requirements is the better choice. In my case, it is.