Infrastructure behind Lykhari

Lykhari hosts about 500 blogs. Some are on a lykhari.com subdomain, some on their own domain. All of it runs on one rented server and one database.

I built it to be cost effective and as lean as possible. It can still get leaner but good enough for now.

The short version

  • Nextjs, Hetzner server, postgres on AWS RDS, traefik for handling ssl, github actions for ci/cd, promethus, loki, grafana for monitoring

It costs me about $30 a month.

One app for every blog

There is a single Next.js app. When a request comes in, the app looks at the host name and works out which blog it belongs to.

If the host is something like kashif.lykhari.com, it looks up the blog by its slug. If the host is anything else, it looks up the blog by custom domain. If someone visits the subdomain of a blog that has its own domain, they get redirected to the domain.

That one lookup is most of what "multi-tenant" means here. Every blog uses the same code, the same database tables and the same server. A blog is just a row.

The app uses Prisma for the database and Lexical for the editor. Posts are stored as the editor's JSON, which means the same content can be turned into HTML for the page, text for search engines, or an email for subscribers.

I try to keep dependencies low. The app has around 30 of them, and most of those are the obvious ones: Next, React, Prisma, Stripe, the AWS SDK.

One server

Everything runs on a single Hetzner machine using Docker Compose. The compose file has these services:

  • traefik, the front door. It takes every request on ports 80 and 443 and sends it to the app.
  • blue and green, two copies of the app. Only one is live at a time. More on that below.
  • cron, a tiny container that runs curl in a loop.
  • prometheus, loki, grafana, node-exporter and blackbox-exporter for monitoring.

That is the whole production environment. No Kubernetes, no load balancer, no queue, no Redis.

The database

Postgres runs on AWS RDS. It is the one managed piece, because backups and upgrades are the thing I least want to get wrong by hand.

The database only accepts connections from the app server. When I want to look at production data, I open an SSH tunnel through the server and connect through it. It is a little slower, but nothing else in the world can reach the database.

Getting there meant moving one thing. Database migrations used to run during the image build on GitHub's machines, which meant the database had to accept connections from them. Now migrations run on the server itself, right before a deploy. Once that moved, I could lock the database down to a single address.

HTTPS for every domain

This is the part that took the most learning.

Blogs on *.lykhari.com share one wildcard certificate. Traefik gets it from Let's Encrypt using a DNS challenge, where it proves ownership by adding a record in Route 53.

Custom domains work differently. The writer points their domain at the server. When they save it in their dashboard, the app first checks that the DNS is actually pointing at us. Then it writes a small YAML file that Traefik watches. Traefik notices the change, asks Let's Encrypt for a certificate using an HTTP challenge, and starts serving the domain. No restart.

The app and Traefik share that file through a Docker volume. It is a simple way for the app to tell Traefik about new domains.

Three things I learned the hard way:

One certificate per domain. At first, every custom domain sat in one Traefik router. Traefik then requests a single certificate covering all of them. The problem is renewal. A certificate with many names renews all together or not at all. When one writer let their domain lapse, renewal failed for everyone. Now each domain gets its own router and its own certificate, so one dead domain only affects itself.

Traefik throws away the whole file if the router list is empty. If there are no custom domains, the routers key has to be missing entirely. An empty one makes Traefik reject the file, which took the live app with it.

Removing a domain does not remove its certificate. The old certificate stays in Traefik's storage and keeps failing to renew in the logs. I clean those up by hand.

Deploys

I deploy by creating a release on GitHub. That starts a workflow that does this:

  1. Build the app image and push it to Docker Hub
  2. Build a second, smaller image that only holds the database migrations
  3. Copy the compose file and a deploy script to the server over SSH
  4. Run the deploy script

The deploy script does the interesting part:

  1. Work out which color is live, blue or green
  2. Run the migrations. If they fail, stop here. The live app keeps serving and nothing changed.
  3. Start the other color with the new image
  4. Wait for its health check to pass, up to a minute
  5. Point Traefik at the new color
  6. Stop the old one
  7. Clean up old images and logs

Step 5 is a one-line edit to Traefik's config file. Traefik picks it up immediately and new requests go to the new container. Nobody sees downtime.

A small money saving detail: Docker Hub's free plan allows one private repository. So the migration image is a tag inside the app's repository instead of a repository of its own. A second repository would have been public and shown the whole database schema to anyone.

Cron

Scheduled posts and a couple of background jobs need something to run them every minute. I use a container running curl in a loop:

while true; do
curl -H "Authorization: XYZ" https://lykhari.com/api/cron/publish
curl -H "Authorization: XYZ" https://lykhari.com/api/cron
sleep 60
done

The endpoints check the secret and do their work. One publishes posts whose scheduled time has passed. The other picks one new post from a paying writer and asks an AI model to find a single strong line in it, which the writer can share. The model is told to return nothing more often than not, and it does.

Images, email and payments

Images go straight from the writer's browser to S3. The app hands out a short-lived signed upload URL, so the file never passes through the server.

Email to subscribers goes through SendGrid. SendGrid sends back events like bounces and spam reports to a webhook, which checks the signature before trusting them. Sign-in links go out over plain SMTP.

Payments are Stripe. A webhook tells the app when someone subscribes or cancels, and the app flips their blog to PRO or back.

Monitoring

Prometheus checks both app containers every 15 seconds through the blackbox exporter, so I know if the live one stops answering. Once an hour it also checks the HTTPS certificates of the custom domains, which is how I would notice a renewal problem before a writer does.

Node exporter watches the server itself: disk, memory, CPU. App logs go to Loki and are kept for a week. Grafana shows all of it.

Prometheus keeps 90 days of data under a 2 GB cap. Early on it could not fit 90 days, and the reason was host metrics being collected every 5 seconds. They made up three quarters of everything stored. Nobody needs server memory at five second resolution, so it went to 15 and the problem disappeared.

Testing

There is a Playwright test suite that runs the real app in a browser: signing in, writing, publishing, changing settings. A couple of the test files only check that one blog can never see another blog's drafts or data, since every blog shares the same tables. I run it before releases and after big changes. For a one person project, this is what lets me deploy without being scared.

What is missing

The honest weak spot is that it is one server. If the machine dies, every blog is down until I bring up a new one. The database would be fine, and the whole server is described in one compose file and one script, so rebuilding is a matter of time. For a project of this size, I think that is the right trade. I will change it when it stops being.

There is also no CDN in front of the blogs and no cache layer. So far nothing has needed one.

Why like this

I wanted something I could understand completely, that costs little, and that will still work in ten years without me chasing upgrades. Docker, Postgres, a reverse proxy and a shell script have been around a long time and will be around for a long time.

If you need help with your infra email me at ya.kashif@gmail.com