Self-Hosted Backup Architecture: Databases, Volumes, Configs and Restore Testing


A useful backup is not a copy of a server. It is a tested path back to a working service.

That distinction matters for self-hosted systems because one application may span several kinds of state at once: a PostgreSQL database, uploaded media, Docker volumes, Compose files, environment variables, certificates, API keys and reverse-proxy configuration. Copy only the obvious data directory and you may discover during an outage that the database is inconsistent, the secrets are missing, or nobody remembers the restore order.

For most Docker-based homelabs and small self-hosted deployments, the practical design is:

  1. back up databases with a database-aware method;
  2. back up durable files and application data separately;
  3. preserve the configuration needed to recreate the stack;
  4. protect secrets and encryption keys without casually duplicating them everywhere;
  5. keep at least one backup outside the primary failure domain;
  6. verify repositories and perform real restore drills.

The objective is not maximum backup complexity. It is a recovery process you can execute when the original server is unavailable.

The five things a self-hosted service may need restored

State Examples Preferred backup treatment Main failure if omitted
Database PostgreSQL, MariaDB/MySQL, SQLite Database-consistent dump/base backup or application-supported method Service starts with missing or inconsistent records
User files Photos, documents, media, attachments File-level backup with version history Irreplaceable content is lost
Persistent app state Docker named volumes, bind-mounted app data File/volume backup, coordinated with the application Settings, indexes or service state disappear
Configuration compose.yaml, reverse proxy, schedules, service settings Versioned config backup; Git for non-secret text where appropriate Rebuild becomes guesswork
Secrets and keys .env, passwords, tokens, TLS/private keys, backup repository keys Encrypted, access-controlled backup with separate recovery instructions Restored data cannot be decrypted or services cannot authenticate

The categories overlap. The important step is to inventory them explicitly for each service rather than treating /var/lib/docker or one storage pool as a complete recovery plan.

Start with recovery objectives, not backup software

Before choosing restic, Borg, snapshots or cloud storage, decide what loss and downtime are acceptable.

Two concepts are useful even in a home lab:

  • Recovery point objective (RPO): how much recent data can be lost. If losing one day of photos is unacceptable, a nightly backup is not enough.
  • Recovery time objective (RTO): how long the service can remain unavailable. A backup that takes two days to retrieve may be acceptable for an archive but not for a service the household depends on.

Different data deserves different schedules. A mostly static media library may need infrequent incremental backup, while a busy database may justify frequent dumps or continuous archiving.

Do not treat a live database like an ordinary folder

One of the most consequential mistakes is copying database files while the database is actively changing and assuming the result is recoverable.

PostgreSQL documents three distinct backup approaches: SQL dumps, file-system-level backups, and continuous archiving with point-in-time recovery. A pg_dump produces a logically consistent snapshot and can generally be restored into newer PostgreSQL versions. For higher-reliability deployments, a base backup combined with archived write-ahead log (WAL) records enables point-in-time recovery.

That leads to a useful rule:

Back up a database using a method that the database itself documents as consistent, then back up the resulting dump/archive as part of the broader backup system.

For a small service, this can be as simple as scheduling a PostgreSQL dump before the backup tool snapshots the dump directory. For a busier system with stricter RPO requirements, WAL archiving and tested base backups may be more appropriate.

Do not assume that copying the underlying database volume is equivalent to pg_dump or a valid PostgreSQL base backup. PostgreSQL's documentation explicitly separates logical dumps from file-system backups and continuous archiving because they have different consistency and recovery properties.

SQLite is different because it is an embedded database, but the same principle applies: use the application's documented backup method or SQLite's supported backup mechanisms rather than inventing a file-copy procedure around an actively written database.

Docker volumes persist, but persistence is not backup

Docker volumes exist outside an individual container's writable layer. Removing a container does not automatically remove its named volume, which is exactly why volumes are appropriate for persistent container data.

But that persistence protects mainly against container replacement, not against disk failure, accidental deletion, corruption, ransomware, theft or loss of the entire host.

Docker's documentation includes explicit procedures for backing up and restoring volume data. The broader architectural point is more important than the exact command: identify every persistent mount in the Compose stack and decide whether it contains:

  • authoritative data that must be backed up;
  • cache/index data that can be regenerated;
  • database files that need database-aware handling;
  • configuration that should instead live in a controlled config directory.

This keeps backups smaller and makes restores easier to reason about.

Named volumes versus bind mounts

Docker recommends volumes as the preferred general mechanism for persistent container data, while bind mounts are appropriate when host-side access to the files is required.

From a backup perspective:

  • a bind mount gives the backup system an obvious host path;
  • a named volume is managed by Docker and should be inventoried by volume name and service;
  • neither is automatically backed up merely because it survives container recreation.

Avoid building recovery around undocumented internal paths if a supported volume export or application-level method exists.

Configuration should let you rebuild the machine

A good test is this: if the server vanished, could you reinstall the operating system and reconstruct the service definitions without consulting the old disk?

For a Compose-based system, preserve at least the non-secret parts of:

  • compose.yaml and override files;
  • reverse-proxy configuration;
  • service-specific config files;
  • scheduled jobs and timers;
  • storage mount definitions;
  • firewall/network notes where they are not generated elsewhere;
  • documented image versions or tags;
  • restore order and dependencies.

Plain-text configuration is a good fit for version control when it contains no secrets. Git history is useful for answering what changed?, but a Git repository is not by itself a disaster-recovery strategy: make sure the repository also exists outside the failed host.

Secrets need a different recovery plan

Do not solve secret recovery by scattering plaintext .env files, API tokens and private keys across every backup target.

Instead, decide which secrets are actually required to recover the service and store them in a protected form. Examples include:

  • database credentials;
  • application secret keys;
  • SSO/OIDC client secrets;
  • TLS private keys when they cannot simply be reissued;
  • encryption keys for encrypted application data;
  • backup repository passwords or keys.

The backup repository credential is particularly important. An encrypted backup that survives the disaster is still unusable if its only key was stored on the destroyed server.

Keep recovery credentials under separate access control and document how an authorized operator retrieves them. Do not put the only copy of a backup password inside the backup it unlocks.

Snapshots are useful, but they solve a different problem

Filesystem, ZFS, Btrfs, LVM, VM and NAS snapshots are excellent for fast rollback and creating a stable point from which a backup can be copied. They can dramatically reduce recovery time after a bad update or accidental edit.

They are not automatically independent backups.

If the snapshot lives on the same storage pool and that pool fails, both the active data and snapshot can disappear together. The same applies to replication: synchronizing corruption or deletion to a second machine quickly is still synchronization, not guaranteed historical recovery.

A strong design often uses snapshots plus versioned backups:

  1. create or obtain an application-consistent point;
  2. snapshot where appropriate;
  3. send versioned backup data to another storage target;
  4. retain older recovery points according to the service's needs.

Separate the backup from the primary failure domain

A backup stored only on the same server is vulnerable to the same power event, controller failure, filesystem damage, theft and many administrative mistakes as the source.

NIST guidance on ransomware recovery emphasizes maintaining and testing backups, while its recovery guidance also stresses isolation so destructive events cannot readily reach every copy. CISA guidance similarly recommends multiple copies with at least one in a physically separate, segmented or otherwise secure location.

For a home server, practical second locations include:

  • another machine that is not permanently writable from the application host;
  • an external drive that is disconnected or rotated after backup;
  • a remote NAS at another location;
  • object storage or another cloud backup target;
  • a combination of local fast recovery plus remote disaster recovery.

The goal is independent failure, not collecting backup destinations for their own sake.

Restic and Borg: what they add

Tools such as restic and BorgBackup are useful because they add capabilities that a simple directory copy lacks: deduplication, encrypted repositories, snapshots/archives, retention workflows and integrity checking.

They do not make application state consistent automatically. If the source is a live database, you still need the database-aware step first.

restic

restic supports encrypted repositories and snapshot-based backups across local and remote storage backends. Its documentation provides explicit restore workflows and repository checking commands. That makes it suitable for a design where application-consistent exports and ordinary files are collected first, then stored as versioned snapshots.

BorgBackup

Borg provides encrypted, deduplicated archives and includes borg check for repository/archive verification. Its extraction tooling also supports dry-run extraction, which is useful as one layer of verification before a full recovery exercise.

Choose based on the targets and operational model you actually need. The architecture matters more than whether the repository format is restic or Borg.

Integrity checking is not the same as restore testing

A repository check can tell you that backup structures and stored data pass the tool's integrity checks. It does not prove that the restored application will work.

A restore drill should answer questions such as:

  • Can the repository be accessed without the original server?
  • Are all required credentials available?
  • Can the database dump actually be imported?
  • Do restored file ownership and permissions make sense?
  • Are application and database versions compatible with the backup?
  • Does the service start with networking and reverse proxy configuration restored?
  • Can users log in and access representative data?
  • Are background jobs, thumbnails, indexes or other derived state rebuilding as expected?

NIST's backup guidance repeatedly treats testing as part of backup management, not as an optional final step. That is the right model for a homelab too.

A practical restore-first layout

A small Docker server might use a structure like this conceptually:

recovery-set/
├── databases/
│   ├── immich-postgres.dump
│   └── app2-postgres.dump
├── files/
│   ├── photos/
│   └── documents/
├── app-data/
│   └── service-specific-persistent-data/
├── config/
│   ├── compose/
│   ├── reverse-proxy/
│   └── service-config/
└── recovery-notes/
    ├── inventory.md
    └── restore-order.md

Secrets do not need to live in that plaintext hierarchy; they can be recovered from a separate encrypted secret store or protected recovery package.

The backup tool then captures the recovery set plus any other required persistent data and sends it to one or more independent repositories.

Example recovery order

For many containerized services, a safe conceptual sequence is:

  1. rebuild the host and storage mounts;
  2. restore Compose and infrastructure configuration;
  3. recreate empty volumes/directories with correct ownership;
  4. start only the database service if needed;
  5. restore the database using its supported tool;
  6. restore uploaded files and other authoritative application data;
  7. restore required secrets and keys through the protected recovery process;
  8. start the application stack;
  9. verify application health and representative user data;
  10. re-enable scheduled jobs only after the restored state is confirmed.

Exact steps vary by application. Immich, Nextcloud, Vaultwarden, Paperless-ngx and other services have different databases, storage layouts and compatibility requirements. Their own backup/restore documentation should override a generic sequence when it is more specific.

What to test after every major change

Revisit the recovery plan when you:

  • move from bind mounts to named volumes;
  • migrate databases or major database versions;
  • add application-level encryption;
  • change reverse proxies, SSO or identity providers;
  • move media to a NAS or object store;
  • change backup repository credentials;
  • add a second host or replication layer;
  • substantially change Compose files or storage paths.

A backup configuration that was correct six months ago can silently become incomplete as the service architecture changes.

Common failure patterns

"The RAID is the backup"

RAID/ZFS mirrors and parity improve availability against some device failures. They do not preserve an independent historical copy against deletion, corruption, theft or loss of the whole system.

"The NAS backs itself up"

Snapshots on the same NAS are useful recovery points, but a separate backup target is still needed for failures that affect the whole appliance or pool.

"Docker volumes survive containers, so the data is safe"

They survive normal container lifecycle operations. They do not survive every host or storage failure.

"Everything is in Git"

Git is excellent for configuration history. Databases, uploaded files and secret recovery usually need other mechanisms.

"The backup job says success"

A successful job confirms that a process completed. Only integrity checks and restore exercises provide evidence that the data is recoverable.

One small server

Use application-consistent database dumps, file-level backups, versioned Compose/configuration, and an encrypted backup repository on a physically separate disk or remote target. Test a sample restore regularly.

Server plus NAS

Keep the NAS as shared storage if that fits the workload, but do not assume server-to-NAS placement alone creates a backup. Maintain versioned recovery copies and make sure at least one copy is outside the shared server/NAS failure domain.

Multi-service homelab

Create a service inventory recording database type, persistent paths, backup method, RPO, retention and restore steps. Centralize backup scheduling where practical, but keep application-specific consistency hooks for databases and stateful services.

High-value or frequently changing data

Use shorter backup intervals, database-native continuous or point-in-time recovery where justified, protected off-site copies, stronger monitoring and scheduled full recovery drills.

The decision rule

A self-hosted backup system is ready when you can answer four questions without touching the failed server:

  1. What exact data does each service require?
  2. Where is each required copy stored?
  3. Which credentials and keys unlock it?
  4. What is the tested sequence that turns those copies back into a working service?

If any answer depends on remembering how the old machine was configured, the backup architecture is incomplete.

Sources