Skip to content
Skip to the article
In NixOS: 13 articles
NixOS

Reference architecture for self-hosting business systems on NixOS

A reference pattern for hosting ERP, CMS and similar apps on bare-metal NixOS with Nomad, systemd-nspawn, Cloudflare Tunnel, ZFS and JuiceFS, no Kubernetes.

Updated
Tags
  • nixos
  • architecture
  • hosting
  • infrastructure
Reading time
16 min

This document describes the hosting pattern Avunu recommends for business systems such as ERP, CMS and identity servers, and how to build it with public tools. Read it when you want a small, declarative, Kubernetes-free platform on a few bare-metal servers. If you are still deciding whether Nix suits your team, start with Why Nix for Business Systems.

Note

This is a reference pattern, not a product and not a description of any one deployment. Avunu does not sell, support or distribute this platform as software. Nothing here is a turnkey installer: you assemble the layers from the public projects linked below, and the code that glues them together is yours to write. Every value below is a starting point to adjust, not a record of a live system.

The pattern in one page

The design rests on a handful of decisions that reinforce each other.

  • Every node is a bare-metal NixOS machine that builds its own configuration from one git repository.

  • Every application runs in its own systemd-nspawn container that is built on every node but stays off until a scheduler turns it on.

  • Nomad only starts and stops those containers. It never ships images or code.

  • Every site reaches the internet through its own outbound Cloudflare Tunnel, so serving a site needs no inbound port.

  • Latency-sensitive data lives on local ZFS. Shared files live on JuiceFS backed by S3-compatible object storage.

  • Backups are per-app, logical, off-provider and rehearsed.

There is no Kubernetes, no image registry and no load balancer. That is the point: three or four servers do not need a control plane built for thousands.

git repository (one flake, one topology file)
        |  every node pulls on a timer
        v
bare-metal NixOS nodes (ZFS, Nomad, etcd, JuiceFS client)
        |  Nomad starts a dormant container where it should run
        v
systemd-nspawn container  --cloudflared (outbound only)-->  Cloudflare edge
   app + web server on a unix socket

Declare the fleet in one file

Keep one data file, a topology, as the single source of truth. It lists the nodes, their private VLANs and addressing, the object-storage endpoint, the backup target and schedule, and a list of role tags per node.

Derive everything else from it: network units, firewall rules, Nomad server counts, etcd membership and Prometheus scrape targets. When a node is added, you edit one file and the rest follows.

Assign singleton roles by tag

Some services must run on exactly one node: the JuiceFS metadata engine, the job registrar, Prometheus and the autoscaler. Do not guard them with "the first node alphabetically". Renaming or adding a node would silently move them.

Use tags instead, such as juicefs-meta, nomad-leader and metrics, and make evaluation fail when a singleton tag matches zero or two nodes. Two metadata engines over one bucket would be a split-brain filesystem, so that case must be an error and not a runbook warning.

clan is a good fit for the machine registry, role tags, secrets and bare-metal install. Use it for those jobs only, and keep the container, Nomad and storage orchestration in your own code, where you can read and change it.

Install by push, run by pull

Install each machine once from your workstation with nixos-anywhere, using disko for the disk layout. After that, nodes pull.

Each node runs a timer that builds its own system closure from the git flake and switches only when the result differs from the running system. A timer that fires every ten minutes with a small random delay is a reasonable starting point, so nodes do not rebuild in lockstep.

nix build --refresh --no-link --print-out-paths "<FLAKE_URL>#nixosConfigurations.<NODE_NAME>.config.system.build.toplevel"
nixos-rebuild switch --flake "<FLAKE_URL>#<NODE_NAME>"

The first command builds and prints the new closure path. Compare it with readlink -f /run/current-system, and run the second command only if they differ. --refresh bypasses the flake tarball cache so a fresh push is seen within the interval.

systemd.timers.cluster-upgrade = {
  wantedBy = [ "timers.target" ];
  timerConfig = {
    OnBootSec = "5min";
    OnCalendar = "*:0/10";
    Persistent = true;
    RandomizedDelaySec = "90s";
  };
};

Keep a hold file for manual pushes

If you ever push a build from your working tree, the next timer tick rebuilds from the last pushed commit and silently reverts you. Give the pull service a condition that skips it while a hold file exists, and put the file under /run so a forgotten hold clears on reboot.

systemd.services.cluster-upgrade.unitConfig.ConditionPathExists = "!/run/cluster-upgrade.hold";

Treat manual push as break-glass only. Backups and restores should also create a hold, so a deploy never reloads a container while its database is being dumped.

Run each app in its own container

Define every app as a NixOS container that is built on every node but has autoStart = false. The nodes already contain the compiled closure in /nix/store, so starting an app anywhere is a cold start, not an image transfer.

containers.<APP_NAME> = {
  autoStart = false;
  privateNetwork = true;
  hostBridge = "br-app";
  bindMounts = { /* secrets, shared files, local state */ };
  config = { ... }: { /* the app's own NixOS module */ };
};

See the NixOS manual on declarative containers and the systemd-nspawn reference.

Let Nomad start and stop, nothing more

Nomad runs one task per app using the raw_exec driver. The task is a small wrapper that starts the container, stays alive while its systemd unit is active, and stops it cleanly on SIGTERM.

cleanup() { nixos-container stop <APP_NAME> || true; }
trap cleanup TERM INT
nixos-container start <APP_NAME>
while systemctl is-active --quiet "container@<APP_NAME>"; do sleep 2; done
exit 1

"Task alive" and "container running" are then the same thing. If the unit fails, the task exits non-zero and Nomad reschedules. raw_exec has to run as root and the Nomad Docker driver is not needed.

Run Nomad as both server and client on every node, turn ACLs on from the start, and bind its traffic to a private VLAN. Register jobs declaratively: one tagged node waits for Raft leadership, then runs nomad job run for every rendered jobspec. Nomad diffs jobs itself, so an unchanged job is a no-op.

Handle the failure cases deliberately

  • Orphans. A SIGKILL or host reboot can leave a container running that Nomad no longer owns. Run a fencing timer on every node (a short interval such as 30 seconds is a sensible start) that asks the local Nomad agent which allocations run here and stops any managed container it does not claim. Make it fail safe: if the API is unreachable, do nothing.

  • Kill timeouts. Nomad silently caps a task's kill_timeout at the client's max_kill_timeout. Its default is 30 seconds, which cuts a slow app's shutdown short, so raise it and assert that every app's kill timeout fits under it.

  • Deploys. Because containers are not autoStart, a host switch does not restart them. Make the container closure a reload trigger so an active container gets the new system activated inside it, and only changed units restart. Bind mounts, interfaces and capabilities are fixed at nspawn start, so changing those needs a Nomad restart of the job.

  • Slow starts. A reload and a start may include a database migration. Raise timeoutStartSec per app, because the nixos-containers default of one minute will kill a long migration.

  • Isolated failures. If a container's own activation fails, the host's switch fails too, which blocks deploys for every other app on the node. Replace the reload so it reports the guest's failed units (journal, a metric, an alert) without failing the host switch.

Scale from resource pressure

Run prometheus-node-exporter on every node with the systemd collector. Each container is a container@<name>.service unit, so the collector reports its state, restart count and task count with nothing installed inside. It does not report per-unit CPU or memory. For those, add a host-side exporter that reads the cgroup tree, such as cAdvisor, and check that its metrics include the container@<name>.service cgroups. The Nomad Autoscaler can then read those metrics from Prometheus and scale a job's count against a target.

Give each app a tier: monolith, split, replica or scale. Monoliths and split apps are pinned to a home node because their state is local. Treat the tier as a platform setting so the app's own module never changes as it grows.

Put one tunnel in front of each site

Run cloudflared inside each container and point it at a unix socket, so no TCP port listens anywhere in the request path. Cloudflare connectors dial out, so the firewall needs no inbound rule for the sites. Keep the hosts' own administrative access to the minimum you need, key-only, and off the path the sites use.

cloudflared tunnel create <APP_NAME>
cloudflared tunnel route dns <APP_NAME> app.example.com

The tunnel credentials arrive as a read-only bind mount. The ingress target is a path, unix:/run/<APP_NAME>/<APP_NAME>.sock, which is correct on whichever node runs the container, so there are no static addresses, no ACME and no host-side proxy.

Warning

The tunnel unit is the site, not the container or the cluster. Cloudflare requires every connector on a tunnel to serve identical routes and sends each request to the nearest one with no route awareness. Two different sites sharing a tunnel would silently serve each other's traffic. Two containers of the same app may share one.

Each hostname needs its own CNAME to <TUNNEL_ID>.cfargotunnel.com. A wildcard record can point at only one tunnel, and every site has its own. Make declaring a hostname without a tunnel id an evaluation error, so ingress can never silently evaluate to nothing.

Three things that fail quietly on unix sockets

  1. Socket permissions gate nothing. nginx sets its listen socket to mode 0666, and gunicorn's follows its umask. The directory is the access control: put each socket in its own directory, never directly in /run.

  2. X-Forwarded-Proto is wrong. The NixOS nginx option recommendedProxySettings hardcodes $scheme, which over a unix socket is http. The apps then build http:// URLs and drop Secure on cookies. Set the header explicitly in each proxying location.

  3. The client IP is lost. A unix socket has no peer address. With nginx use set_real_ip_from unix:; and real_ip_header CF-Connecting-IP. Caddy has no equivalent, so recover it in the application.

Know the limits

Cloudflare fails over between connectors, not origins, so a live connector in front of a dead container returns 502. Proxied requests have a body size cap that depends on your plan. A container restart drops its tunnel briefly. Accept these deliberately; they suit business workloads with modest uptime needs.

Choose storage by access pattern

  • Local ZFS for databases and other latency-sensitive data. Mirror the two NVMe drives, use zstd on / and /nix, set a 16k recordsize on database datasets, enable auto-snapshots, scrub weekly, and cap the ARC so database memory is not starved. Keep a reservation so the copy-on-write pool never fills. See OpenZFS.

  • JuiceFS for shared, follow-the-app state such as document roots, uploads and attachments. Mount it on every node and bind-mount the app's path into its container. See JuiceFS.

For JuiceFS, put the metadata in PostgreSQL on one tagged node and the data in an S3-compatible bucket. Mount without --writeback, so a write is acknowledged only after it reaches S3 and is not lost if the node dies first.

Warning

The metadata engine is a singleton and a single point of failure until you make it highly available. Its loss strands the mount on every node. JuiceFS dumps metadata hourly into the bucket, so make the format step restore the newest dump on a rebuilt node, refuse to format over a bucket that holds objects but no dump, and never pass --force.

Gate Nomad on storage health: a unit checks that the mount is writable, then sets a jfs_ready meta value on the local agent, and jobs that use JuiceFS carry a constraint on it. Without the gate they start where the mount is missing.

Databases and etcd

Run a three-member etcd cluster on the nodes tagged for it. Its job is to be the shared distributed configuration store for PostgreSQL HA, a small, low-churn keyspace that suits etcd well. It is not the JuiceFS metadata store. Derive its membership from tags and bind it to the private control VLAN.

For high availability, the intended design is Patroni for PostgreSQL, with a primary and a streaming replica on that etcd, and MariaDB Galera with a garbd arbitrator for MariaDB. Keep data on local ZFS, never on JuiceFS, and give each client its own cluster so one client's migration stalls only that client.

Note

Treat Patroni and Galera as the target design, not as something to switch on. A first build can run a local PostgreSQL or MariaDB inside each container, which has no high availability. Quorum needs at least three voters, so a single-node bring-up has none either. Build and test failover before you rely on it.

Manage secrets without a key on the container

Generate secrets in an offline sandbox, commit them encrypted with age, and decrypt on the host in an activation script. clan's vars does this. The decrypted file is bind-mounted read-only into the container, so a container never holds a key.

Two rules save you pain:

  • Anything issued by a third party is a prompt-once value, stored after the first entry.

  • A referenced but ungenerated secret must be a build error, not an empty file.

Deliver each node's decryption key once, at install. You cannot encrypt to a host key before the host exists, and the pull loop uploads nothing.

Back up each app and rehearse the restore

Back up on each app's home node, into its own restic repository on an off-provider target, so a backup does not live with the provider it protects. Each snapshot holds a logical dump and the app's shared files together.

  • PostgreSQL: pg_dumpall --globals-only, plus pg_dump -Fc for each database.

  • MariaDB: mariadb-dump --databases for each database, plus its users.

Do not ship raw database directories. A dump survives a major-version upgrade and a new node; a directory copied from a live server survives neither. For apps with state a dump cannot capture, such as a directory server, take the files from a ZFS snapshot instead.

A sound starting schedule is nightly backups with retention of 7 daily, 4 weekly and 12 monthly snapshots. Each week, prune and run restic check --read-data-subset=5%, then run a restore rehearsal that loads the latest dump into a throwaway database server and compares table lists. That is the difference between "the backup ran" and "the backup can be restored".

Tip

On an S3 target, enable bucket versioning and give the backup user a policy that cannot delete object versions or change bucket settings. A compromised cluster can then hide objects but cannot destroy history.

Backups are nightly, not point-in-time. ZFS snapshots narrow local loss to 15 minutes, but only for losses on the same node.

Observe it

  • Metrics. Every node reports node-exporter (with systemd and textfile collectors), smartctl_exporter and a storage probe that writes to the JuiceFS mount under a timeout. The metrics node runs Prometheus and a blackbox exporter that checks each app's public hostname through Cloudflare, which is the path a user takes.

  • Alerts. Prometheus sends to Alertmanager, a bridge turns alerts into email, and critical alerts repeat hourly. Cover failed container units, failed deploys, stale backups, unhealthy ZFS pools, and an unwritable shared mount.

  • Dead-man. Alerting on the cluster cannot report the cluster's death. Send an always-firing heartbeat to something off-cluster, such as a small Cloudflare Worker, that emails when heartbeats stop.

  • Logs. Vector on each node reads the host journal and each container's persistent journal directly and pushes to Loki, with chunks in object storage. Grafana is the reading surface. Nothing runs inside the container.

Label logs by node, app, site, service and level, and have your apps write to the journal, not log files. Add a weekly positive report as well: if it stops arriving, that tells you something even when nothing alerted.

Typed app kinds

An app is one short file. Its kind is a typed namespace on that entry, such as odoo, frappe, wordpress, freeipa or grafana, and it fills in the container, paths, sockets, secrets, backups and ingress. The file carries only what differs, such as a package, a hostname, a worker count, a tunnel id and resource limits.

{ inputs, ... }:
{
  homeNode = "<NODE_NAME>";
  odoo = {
    package = inputs.<PROJECT>.packages.x86_64-linux.default;
    hostname = "erp.example.com";
    workers = 4;
  };
  tunnel.id = "<TUNNEL_ID>";
  resources = { cpu = 2000; memoryMB = 4096; };
}

Kinds are where platform knowledge accumulates: which units to stop around a restore, where the socket lives, which secrets exist. Adding an app is then a file plus a push, and adding a new kind is the only time you write platform code.

Adapt it

  • Start with one node and no HA, and add voters as hardware arrives. Quorum for Nomad and etcd derives from the tags, so a placeholder node must not carry voter tags until it is real, or the cluster waits forever for a machine that does not exist.

  • Keep a console open for changes to a node's WAN networking. A mistake there is a lockout, not a failed request.

  • Pin your configuration framework by revision and bump it alone, in its own commit, because every node applies the lock file within minutes.

  • Verify what you copy. This is a pattern, not a record of any one deployment, and your hardware, provider and workloads will differ.