The self-hosted infrastructure behind this site
The page you are reading runs on an ARM board at my home: tag-triggered CI, rolling zero-downtime releases, self-healing and autoscaling.
- Docker
- GitHub Actions
- Traefik
- nginx
- Cloudflare Tunnel
This portfolio is not on Vercel or a VPS. It runs on an Armbian single-board computer at my home — four Cortex-A53 cores, about 1.8 GB of RAM, shared with a handful of other services. That hardware budget drove nearly every decision below.
Nothing is exposed to the internet#
The machine has no static IP and I did not want to forward ports on my router. Traffic
arrives through a Cloudflare Tunnel: a cloudflared process opens an outbound
connection, and Cloudflare pushes requests back down it. No open ports, no home IP in
public DNS, and TLS terminates at Cloudflare.
Behind the tunnel, Traefik routes by hostname into an nginx caching layer, and only then into the application containers.
Why nginx sits in the middle#
Waking Node to re-render a page on every view is expensive on this CPU. nginx keeps a micro-cache in RAM: HTML and RSC payloads for 60 seconds, hashed build assets for seven days. When several people request the same page at once, exactly one request reaches the app and the rest share its result:
location / {
proxy_pass http://portfolio_app;
proxy_cache portfolio_pages;
# Next.js serves HTML and the RSC payload from the same URL, distinguished
# only by a request header — the cache key has to keep them apart, or a
# browser expecting a page receives a component payload.
proxy_cache_key "$scheme$host$request_uri|$enc|$http_rsc|$http_next_router_prefetch";
proxy_cache_lock on;
proxy_cache_background_update on;
proxy_cache_use_stale error timeout updating http_500 http_502 http_503 http_504;
}proxy_cache_use_stale is my favourite line here: if the app dies outright, visitors
still get the cached page instead of an error.
Releasing with a single tag#
I write code on a Mac, but builds happen on a Fedora box on the same network running a
GitHub self-hosted runner inside Docker. Pushing a v* tag runs the whole chain: lint
and type checks, an arm64 image build, then zstd compression and a direct stream over
SSH into the board's Docker daemon — no registry involved. A release takes about two
minutes from git push to serving the new version.
Swapping versions without dropping a request#
The site used to run as a single container, so every release meant roughly 30 seconds of downtime. Now nginx only forwards to the containers listed in a file that the deploy script owns:
# New replicas are healthy: point nginx at them, then WAIT for the old workers
# to finish the requests they are still holding before stopping anything.
write_upstream "${new_ids[@]}"
reload_edge || { write_upstream "${old_ids[@]}"; reload_edge; exit 1; }
retire "${old_ids[@]}"New replicas start alongside the old ones and must pass their health checks first. Only then does nginx switch lists and reload gracefully, and only after its old workers have exited — meaning every in-flight request has been answered — are the old containers stopped. A release that never becomes healthy is removed, and it never received a single request.
I tested the whole thing on my Mac against a real Traefik with three continuous request loops (pages, an uncached API endpoint, and a slow three-second download deliberately in flight during the switch): the architecture migration, a normal release, an unhealthy build, a broken nginx config, replacing nginx itself, and scaling up and down — not one failed request.
Healing and scaling itself#
Docker only restarts containers that crash; one that hangs while still "running" sits there forever. A systemd timer runs every 30 seconds: it pulls an unhealthy container out of the nginx rotation before restarting it, adds a replica when CPU stays high, and gives the replica back when things go quiet — always subject to a free-memory floor so the box never has to swap.
Outcome#
- Releases are one
git tag; nothing is done by hand on the server - No downtime on version changes, verified against thousands of live requests
- 300 requests at 20 concurrent: all returned 200, with CPU barely moving
- A broken build never reaches visitors, and rolling back is a single command
- All of it on hardware that costs about as much as a dinner, drawing a few watts