DocsSelf-hosting

Operations

How to run ProjectVerse in production once it's deployed (deploy.md). Commands assume the Docker Compose stack in /opt/projectverse and this alias:

bash
alias pvc='docker compose -f /opt/projectverse/docker-compose.prod.yml --env-file /opt/projectverse/deploy/prod.env'
ServiceWhatState
caddyTLS + reverse proxy, the only public portscaddy-data volume (certificates: keep it)
webNext.js server, proxies /api/*stateless
apiRust API, migrations, background jobsuploads volume (only with STORAGE_BACKEND=local)
dbPostgres 16db-data volume
backupdaily backup loop (profile backup)none

Health and readiness

CheckMeaningHow
GET /healthLiveness: the process answers. Doesn't touch the database. Docker's HEALTHCHECK uses it (pvc ps shows healthy)pvc exec api curl -fsS http://127.0.0.1:8080/health
GET /health/readyReadiness: database and storage backend reachable. 200 or 503pvc exec api curl -fsS http://127.0.0.1:8080/health/ready
WebThe Next.js server renderscurl -fsS -o /dev/null -w '%{http_code}\n' https://<DOMAIN>/login

/health and /health/ready aren't reachable from the internet (the web server only proxies /api/*). For external uptime monitoring, check https://<DOMAIN>/login (expects 200). For deeper checks, run the readiness command from a cron job on the host that alerts on failure.

Logs

bash
pvc ps                                   # state and health of every service
pvc logs -f --tail=200                   # everything (= make prod-logs)
pvc logs api --since 30m                 # API: requests (tower_http), jobs, errors
pvc logs caddy | grep '"status":5'       # 5xx at the edge, as JSON
pvc logs backup                          # backup runs ("backup FAILED" on errors)
  • API verbosity: RUST_LOG in deploy/api.env (default projectverse_api=info,tower_http=info,info); pvc up -d api applies a change.
  • Docker rotates each container's log at 10 MB, 5 files.
  • Errors also go to Sentry when SENTRY_DSN / NEXT_PUBLIC_SENTRY_DSN are set.
  • Caddy redacts cookies, auth headers and one-time tokens in URLs. Logs still contain IP addresses and paths: treat them as personal data.

Backups

What, where, how long

  • What: the whole Postgres database, as pg_dump --format=custom (compressed; restore with pg_restore). This includes users, workspaces, issues, the audit log and the encrypted integration secrets and 2FA seeds.
  • Not included: attachment files. With R2 they stay in the bucket (attachments/); see "Off-site copy" below. With STORAGE_BACKEND=local they're in the uploads volume: back that up too (e.g. restic or rclone of /var/lib/docker/volumes/projectverse_uploads).
  • Not included: secrets. A restored database needs the same APP_ENCRYPTION_KEY, or webhooks, integrations and 2FA stop working (see "Rotating APP_ENCRYPTION_KEY"). Keep deploy/api.env and deploy/prod.env in a password manager.
  • Where: the storage backend, at backups/YYYY/MM/DD/<timestamp>.dump: in the bucket with s3, or under STORAGE_DIR in the uploads volume with local (same disk as the database: not a real backup on its own).
  • Retention: each run deletes backups older than BACKUP_RETENTION_DAYS (default 30).

Schedule

Use exactly one of these:

  1. Compose backup service (default): COMPOSE_PROFILES=backup in deploy/prod.env. It runs backup daily at BACKUP_AT (UTC, default 03:00) using deploy/scripts/backup-loop.sh. A failed run is logged as backup FAILED and retried the next day.
  2. systemd timer: deploy/systemd/ (projectverse-backup.service + .timer). Failures show in systemctl --failed and the journal.
  3. Host cron: deploy/cron/projectverse-backup.

Options 2 and 3 run pvc run --rm --no-deps -T api backup; remove COMPOSE_PROFILES=backup when using them. Whatever you pick, alert on failure: watch for backup FAILED, or check that the newest object under backups/ is less than 26 hours old.

On demand

bash
pvc run --rm --no-deps api backup        # e.g. before an upgrade
make backup                              # development: the database in apps/api/.env

make backup runs cargo run --bin backup with apps/api/.env and needs pg_dump 16+ on your machine. If that .env points at the real R2 bucket, the dump lands next to the production backups; point it at a development bucket.

The API image ships pg_dump 16, matching the db service. If the database moves to another major version, build the image with --build-arg PG_MAJOR=<version>.

Off-site copy

The backups and the attachments share one bucket and one token. If that account or token is compromised, both can be deleted together. Keep a second copy elsewhere, e.g. a nightly rclone sync r2:projectverse-attachments other:projectverse-offsite with credentials kept off the server (a different provider or Cloudflare account, with versioning or object lock).

Restoring

deploy/scripts/restore.sh downloads a dump (with rclone, or the AWS CLI as aws s3 cp --endpoint-url $S3_ENDPOINT, using the S3_* values from deploy/api.env) and runs pg_restore --clean --if-exists --no-owner --no-privileges --single-transaction --exit-on-error inside the db container. It shows what it's about to do and asks you to type restore.

bash
cd /opt/projectverse
deploy/scripts/restore.sh --list                         # available dumps
deploy/scripts/restore.sh --latest                       # newest one into the compose db
deploy/scripts/restore.sh --key backups/2026/10/03/<timestamp>.dump --fresh
deploy/scripts/restore.sh --file ./copy.dump             # a dump you already have

What it does with the compose target:

  1. Stops api, web and backup, so nothing writes during the restore (the site answers 502).
  2. With --fresh, drops and recreates the public schema first. Use --fresh when the backup is older than the running release: without it, tables added since the backup would survive the restore, and re-running their migrations would fail on the next start.
  3. Restores in one transaction. If it fails, nothing changed (unless --fresh already emptied the schema) and the services stay stopped until you fix it and re-run, or start them with pvc up -d.
  4. Starts the stack (the API applies any newer migrations) and checks /health/ready.

Restoring with STORAGE_BACKEND=local: find the dump in the volume, copy it out, and use --file:

bash
docker run --rm -v projectverse_uploads:/data:ro alpine find /data -name '*.dump' | sort | tail -3
docker run --rm -v projectverse_uploads:/data:ro -v "$PWD:/out" alpine cp /data/<path>.dump /out/

After a restore: sign in, open a few issues and attachments, check the audit log. Users who signed up or changed something after the backup was taken have to redo it. Say so in your incident notes.

Monthly restore test

A backup counts only once it has been restored. Once a month, restore the newest dump into a throwaway Postgres on the server (production is untouched):

bash
cd /opt/projectverse
dump=$(deploy/scripts/restore.sh --latest --download-only)
docker run -d --name pv-restore-test -e POSTGRES_PASSWORD=test postgres:16-alpine
until docker exec pv-restore-test pg_isready -U postgres -q; do sleep 1; done
docker exec -i pv-restore-test pg_restore -U postgres -d postgres \
  --no-owner --no-privileges --exit-on-error < "$dump"
docker exec pv-restore-test psql -U postgres -c "
  SELECT max(version) AS last_migration FROM _sqlx_migrations;
  SELECT (SELECT count(*) FROM users) AS users, (SELECT count(*) FROM workspaces) AS workspaces,
         (SELECT count(*) FROM issues) AS issues, (SELECT max(created_at) FROM issues) AS newest_issue;"
docker rm -f pv-restore-test && rm -f "$dump"

Check that the counts look plausible, that newest_issue is from the day before the backup and that last_migration matches the newest file in apps/api/migrations/. Note the date, dump name, duration and result in your ops log. Once a quarter, also start an API container against the test database (docker run --rm --network ... -e DATABASE_URL=... projectverse-api:latest) and confirm /health/ready.

Migrations

  • SQL files in apps/api/migrations/, embedded in the binary at build time and applied by the API on every start, before it listens. sqlx holds a Postgres advisory lock while migrating, so several instances starting together are safe.
  • Forward-only: there are no down migrations. To undo a release, restore the backup taken before the upgrade and run the previous image (deploy.md, "Upgrades").
  • Never edit a migration that has been deployed. sqlx checks checksums and the API refuses to start ("migration ... was previously applied but has been modified"). Add a new migration instead.
  • A failing migration stops the API from starting (pvc logs api; the container restarts in a loop and web waits for it). Each migration runs in its own transaction, so a failed one leaves no partial changes. Fix forward with a new release, or restore.
  • Postgres major upgrades (16 → 17): back up, stop the stack, start a new db volume on the new image, restore with restore.sh --file, and bump PG_MAJOR for the API image.

Rotating APP_ENCRYPTION_KEY

APP_ENCRYPTION_KEY (AES-256-GCM) encrypts the secrets the server has to read back: webhook signing secrets, Slack webhook URLs, GitHub integration secrets and TOTP (2FA) seeds. Passwords, sessions and recovery codes are hashed and don't depend on it.

There is one active key and no re-encryption tool yet, so changing the key makes every stored secret unreadable. Rotate only when the key may have leaked, and plan for the cleanup:

  1. Take a backup.
  2. Generate a key (deploy/scripts/generate-secrets.sh), put it in deploy/api.env and the password manager, and restart: pvc up -d api backup.
  3. Outgoing webhooks: in every workspace, Rotate secret on each webhook (or delete and recreate it) and give receivers the new secret. Deliveries fail until then.
  4. Slack: delete and reconnect each Slack integration with its incoming-webhook URL.
  5. GitHub: delete and recreate each GitHub integration, then update the webhook secret in the GitHub repository settings.
  6. 2FA: TOTP codes can't be checked anymore. Reset two-factor for everyone who had it, and ask them to enable it again:
    sql
    -- pvc exec db psql -U projectverse -d projectverse
    BEGIN;
    UPDATE users SET totp_secret_encrypted = NULL, totp_enabled_at = NULL, totp_last_step = NULL
      WHERE totp_secret_encrypted IS NOT NULL;
    DELETE FROM recovery_codes;
    COMMIT;
    (Check these columns against the newest migration first. For a short window those accounts are protected by their password only, so tell them.)

Restoring an old backup also needs the key that was active when it was taken. Keep retired keys (labelled with dates) in the password manager.

Future work: support APP_ENCRYPTION_KEY_PREVIOUS (decrypt with either key, encrypt with the new one) plus a reencrypt command that rewrites every *_encrypted column in one transaction. That would make rotation routine instead of disruptive.

Incident basics

First: write down the time, what you see, and every action you take. Then:

SymptomLook atUsual fix
Site unreachable / TLS errorpvc ps, pvc logs caddy, DNS (dig +short <DOMAIN>), ports 80/443 openRestart caddy. Certificate errors: DNS must point here and port 80 must be reachable for ACME
502 from Caddypvc ps (web/api unhealthy?), pvc logs web apipvc up -d. API crash-looping on a migration: see Migrations
/health/ready 503pvc logs api, pvc ps db, R2 status page, token validityDatabase down or full disk; R2 token revoked or expired
Everything slowdocker stats, df -h, pvc exec db psql -U projectverse -c "SELECT pid, now()-query_start AS age, state, left(query,80) FROM pg_stat_activity ORDER BY age DESC NULLS LAST LIMIT 10;"Kill runaway queries (SELECT pg_terminate_backend(pid)), add RAM/CPU
Disk fulldf -h, docker system dfdocker image prune -a (keeps running images), trim logs, grow the disk
Emails not arrivingpvc logs api | grep -i smtp, provider dashboardCredentials, sender verification, blocked ports
Backups failingpvc logs backupStorage credentials, pg_dump version mismatch (PG_MAJOR), disk

Suspected compromise: take a backup (evidence), then rotate in this order: database password, R2 token, SMTP password, Google client secret, Sentry DSNs. Sign everyone out with pvc exec db psql -U projectverse -d projectverse -c 'DELETE FROM sessions;'. Review the audit log (workspace settings, and audit_events in the database) for user.signed_in, member.*, webhook.* and integration.* actions. Rotate APP_ENCRYPTION_KEY (above) only if the key itself may have leaked. Afterwards, write a short post-mortem: timeline, impact, cause, follow-ups.

Scaling

One API instance handles a small-to-medium team comfortably. Before adding instances, know what's per-instance:

  • Live events (SSE) are per instance. An event reaches only the browsers connected to the instance that produced it. With two API instances, users miss live updates made through the other one until they reload (sticky sessions don't fix this, since the writer and the reader are different users). Multi-instance live events need a shared bus, e.g. Postgres LISTEN/NOTIFY (future work).
  • Background jobs run in every instance, and are safe: the email digest takes a Postgres advisory lock, so only one instance sends it. The webhook delivery worker claims rows with FOR UPDATE SKIP LOCKED, so a delivery goes out once.
  • Rate limits are in memory, per instance, and reset on restart. With N instances the effective limit is up to N times higher.
  • Database connections: each API instance opens up to 20. Postgres allows 100 by default, so add a pooler (PgBouncer in transaction mode) or raise max_connections past ~4 instances.
  • Uploads stage on the instance's local disk, then go to S3. Use STORAGE_BACKEND=s3 with more than one instance (local storage isn't shared).
  • Web is stateless and scales freely. Every web instance needs API_URL pointing at the API (a load balancer address when there are several).
  • Scale vertically first (CPU/RAM for the VPS and Postgres). That avoids every caveat above.

Routine maintenance

  • Weekly: glance at pvc ps, disk usage, Sentry, and that last night's backup exists.
  • Monthly: restore test (above); pvc pull db caddy && make prod-up for base image patches; OS updates (unattended-upgrades handles security fixes).
  • Automatic: TLS renewals (Caddy), backup pruning (BACKUP_RETENTION_DAYS), audit log pruning (365 days).
  • Yearly: review who holds production credentials; rotate the R2 token and SMTP password.