Operations
How to run ProjectVerse in production once it's deployed (deploy.md). Commands
assume the Docker Compose stack in /opt/projectverse and this alias:
alias pvc='docker compose -f /opt/projectverse/docker-compose.prod.yml --env-file /opt/projectverse/deploy/prod.env'| Service | What | State |
|---|---|---|
caddy | TLS + reverse proxy, the only public ports | caddy-data volume (certificates: keep it) |
web | Next.js server, proxies /api/* | stateless |
api | Rust API, migrations, background jobs | uploads volume (only with STORAGE_BACKEND=local) |
db | Postgres 16 | db-data volume |
backup | daily backup loop (profile backup) | none |
Health and readiness
| Check | Meaning | How |
|---|---|---|
GET /health | Liveness: the process answers. Doesn't touch the database. Docker's HEALTHCHECK uses it (pvc ps shows healthy) | pvc exec api curl -fsS http://127.0.0.1:8080/health |
GET /health/ready | Readiness: database and storage backend reachable. 200 or 503 | pvc exec api curl -fsS http://127.0.0.1:8080/health/ready |
| Web | The Next.js server renders | curl -fsS -o /dev/null -w '%{http_code}\n' https://<DOMAIN>/login |
/health and /health/ready aren't reachable from the internet (the web server only proxies
/api/*). For external uptime monitoring, check https://<DOMAIN>/login (expects 200). For
deeper checks, run the readiness command from a cron job on the host that alerts on failure.
Logs
pvc ps # state and health of every service
pvc logs -f --tail=200 # everything (= make prod-logs)
pvc logs api --since 30m # API: requests (tower_http), jobs, errors
pvc logs caddy | grep '"status":5' # 5xx at the edge, as JSON
pvc logs backup # backup runs ("backup FAILED" on errors)- API verbosity:
RUST_LOGindeploy/api.env(defaultprojectverse_api=info,tower_http=info,info);pvc up -d apiapplies a change. - Docker rotates each container's log at 10 MB, 5 files.
- Errors also go to Sentry when
SENTRY_DSN/NEXT_PUBLIC_SENTRY_DSNare set. - Caddy redacts cookies, auth headers and one-time tokens in URLs. Logs still contain IP addresses and paths: treat them as personal data.
Backups
What, where, how long
- What: the whole Postgres database, as
pg_dump --format=custom(compressed; restore withpg_restore). This includes users, workspaces, issues, the audit log and the encrypted integration secrets and 2FA seeds. - Not included: attachment files. With R2 they stay in the bucket (
attachments/); see "Off-site copy" below. WithSTORAGE_BACKEND=localthey're in theuploadsvolume: back that up too (e.g.resticorrcloneof/var/lib/docker/volumes/projectverse_uploads). - Not included: secrets. A restored database needs the same
APP_ENCRYPTION_KEY, or webhooks, integrations and 2FA stop working (see "Rotating APP_ENCRYPTION_KEY"). Keepdeploy/api.envanddeploy/prod.envin a password manager. - Where: the storage backend, at
backups/YYYY/MM/DD/<timestamp>.dump: in the bucket withs3, or underSTORAGE_DIRin theuploadsvolume withlocal(same disk as the database: not a real backup on its own). - Retention: each run deletes backups older than
BACKUP_RETENTION_DAYS(default 30).
Schedule
Use exactly one of these:
- Compose
backupservice (default):COMPOSE_PROFILES=backupindeploy/prod.env. It runsbackupdaily atBACKUP_AT(UTC, default03:00) usingdeploy/scripts/backup-loop.sh. A failed run is logged asbackup FAILEDand retried the next day. - systemd timer:
deploy/systemd/(projectverse-backup.service+.timer). Failures show insystemctl --failedand the journal. - Host cron:
deploy/cron/projectverse-backup.
Options 2 and 3 run pvc run --rm --no-deps -T api backup; remove COMPOSE_PROFILES=backup
when using them. Whatever you pick, alert on failure: watch for backup FAILED, or check that
the newest object under backups/ is less than 26 hours old.
On demand
pvc run --rm --no-deps api backup # e.g. before an upgrade
make backup # development: the database in apps/api/.envmake backup runs cargo run --bin backup with apps/api/.env and needs pg_dump 16+ on your
machine. If that .env points at the real R2 bucket, the dump lands next to the production
backups; point it at a development bucket.
The API image ships pg_dump 16, matching the db service. If the database moves to another
major version, build the image with --build-arg PG_MAJOR=<version>.
Off-site copy
The backups and the attachments share one bucket and one token. If that account or token is
compromised, both can be deleted together. Keep a second copy elsewhere, e.g. a nightly
rclone sync r2:projectverse-attachments other:projectverse-offsite with credentials kept off
the server (a different provider or Cloudflare account, with versioning or object lock).
Restoring
deploy/scripts/restore.sh downloads a dump (with rclone,
or the AWS CLI as aws s3 cp --endpoint-url $S3_ENDPOINT, using the S3_* values from
deploy/api.env) and runs pg_restore --clean --if-exists --no-owner --no-privileges --single-transaction --exit-on-error inside the db container. It shows what it's about to do
and asks you to type restore.
cd /opt/projectverse
deploy/scripts/restore.sh --list # available dumps
deploy/scripts/restore.sh --latest # newest one into the compose db
deploy/scripts/restore.sh --key backups/2026/10/03/<timestamp>.dump --fresh
deploy/scripts/restore.sh --file ./copy.dump # a dump you already haveWhat it does with the compose target:
- Stops
api,webandbackup, so nothing writes during the restore (the site answers 502). - With
--fresh, drops and recreates thepublicschema first. Use--freshwhen the backup is older than the running release: without it, tables added since the backup would survive the restore, and re-running their migrations would fail on the next start. - Restores in one transaction. If it fails, nothing changed (unless
--freshalready emptied the schema) and the services stay stopped until you fix it and re-run, or start them withpvc up -d. - Starts the stack (the API applies any newer migrations) and checks
/health/ready.
Restoring with STORAGE_BACKEND=local: find the dump in the volume, copy it out, and use --file:
docker run --rm -v projectverse_uploads:/data:ro alpine find /data -name '*.dump' | sort | tail -3
docker run --rm -v projectverse_uploads:/data:ro -v "$PWD:/out" alpine cp /data/<path>.dump /out/After a restore: sign in, open a few issues and attachments, check the audit log. Users who signed up or changed something after the backup was taken have to redo it. Say so in your incident notes.
Monthly restore test
A backup counts only once it has been restored. Once a month, restore the newest dump into a throwaway Postgres on the server (production is untouched):
cd /opt/projectverse
dump=$(deploy/scripts/restore.sh --latest --download-only)
docker run -d --name pv-restore-test -e POSTGRES_PASSWORD=test postgres:16-alpine
until docker exec pv-restore-test pg_isready -U postgres -q; do sleep 1; done
docker exec -i pv-restore-test pg_restore -U postgres -d postgres \
--no-owner --no-privileges --exit-on-error < "$dump"
docker exec pv-restore-test psql -U postgres -c "
SELECT max(version) AS last_migration FROM _sqlx_migrations;
SELECT (SELECT count(*) FROM users) AS users, (SELECT count(*) FROM workspaces) AS workspaces,
(SELECT count(*) FROM issues) AS issues, (SELECT max(created_at) FROM issues) AS newest_issue;"
docker rm -f pv-restore-test && rm -f "$dump"Check that the counts look plausible, that newest_issue is from the day before the backup
and that last_migration matches the newest file in apps/api/migrations/. Note the date,
dump name, duration and result in your ops log. Once a quarter, also start an API container
against the test database (docker run --rm --network ... -e DATABASE_URL=... projectverse-api:latest) and confirm /health/ready.
Migrations
- SQL files in
apps/api/migrations/, embedded in the binary at build time and applied by the API on every start, before it listens. sqlx holds a Postgres advisory lock while migrating, so several instances starting together are safe. - Forward-only: there are no down migrations. To undo a release, restore the backup taken before the upgrade and run the previous image (deploy.md, "Upgrades").
- Never edit a migration that has been deployed. sqlx checks checksums and the API refuses to start ("migration ... was previously applied but has been modified"). Add a new migration instead.
- A failing migration stops the API from starting (
pvc logs api; the container restarts in a loop andwebwaits for it). Each migration runs in its own transaction, so a failed one leaves no partial changes. Fix forward with a new release, or restore. - Postgres major upgrades (16 → 17): back up, stop the stack, start a new
dbvolume on the new image, restore withrestore.sh --file, and bumpPG_MAJORfor the API image.
Rotating APP_ENCRYPTION_KEY
APP_ENCRYPTION_KEY (AES-256-GCM) encrypts the secrets the server has to read back:
webhook signing secrets, Slack webhook URLs, GitHub integration secrets and TOTP (2FA)
seeds. Passwords, sessions and recovery codes are hashed and don't depend on it.
There is one active key and no re-encryption tool yet, so changing the key makes every stored secret unreadable. Rotate only when the key may have leaked, and plan for the cleanup:
- Take a backup.
- Generate a key (
deploy/scripts/generate-secrets.sh), put it indeploy/api.envand the password manager, and restart:pvc up -d api backup. - Outgoing webhooks: in every workspace, Rotate secret on each webhook (or delete and recreate it) and give receivers the new secret. Deliveries fail until then.
- Slack: delete and reconnect each Slack integration with its incoming-webhook URL.
- GitHub: delete and recreate each GitHub integration, then update the webhook secret in the GitHub repository settings.
- 2FA: TOTP codes can't be checked anymore. Reset two-factor for everyone who had it,
and ask them to enable it again:(Check these columns against the newest migration first. For a short window those accounts are protected by their password only, so tell them.)
-- pvc exec db psql -U projectverse -d projectverse BEGIN; UPDATE users SET totp_secret_encrypted = NULL, totp_enabled_at = NULL, totp_last_step = NULL WHERE totp_secret_encrypted IS NOT NULL; DELETE FROM recovery_codes; COMMIT;
Restoring an old backup also needs the key that was active when it was taken. Keep retired keys (labelled with dates) in the password manager.
Future work: support APP_ENCRYPTION_KEY_PREVIOUS (decrypt with either key, encrypt with
the new one) plus a reencrypt command that rewrites every *_encrypted column in one
transaction. That would make rotation routine instead of disruptive.
Incident basics
First: write down the time, what you see, and every action you take. Then:
| Symptom | Look at | Usual fix |
|---|---|---|
| Site unreachable / TLS error | pvc ps, pvc logs caddy, DNS (dig +short <DOMAIN>), ports 80/443 open | Restart caddy. Certificate errors: DNS must point here and port 80 must be reachable for ACME |
| 502 from Caddy | pvc ps (web/api unhealthy?), pvc logs web api | pvc up -d. API crash-looping on a migration: see Migrations |
/health/ready 503 | pvc logs api, pvc ps db, R2 status page, token validity | Database down or full disk; R2 token revoked or expired |
| Everything slow | docker stats, df -h, pvc exec db psql -U projectverse -c "SELECT pid, now()-query_start AS age, state, left(query,80) FROM pg_stat_activity ORDER BY age DESC NULLS LAST LIMIT 10;" | Kill runaway queries (SELECT pg_terminate_backend(pid)), add RAM/CPU |
| Disk full | df -h, docker system df | docker image prune -a (keeps running images), trim logs, grow the disk |
| Emails not arriving | pvc logs api | grep -i smtp, provider dashboard | Credentials, sender verification, blocked ports |
| Backups failing | pvc logs backup | Storage credentials, pg_dump version mismatch (PG_MAJOR), disk |
Suspected compromise: take a backup (evidence), then rotate in this order: database
password, R2 token, SMTP password, Google client secret, Sentry DSNs. Sign everyone out with
pvc exec db psql -U projectverse -d projectverse -c 'DELETE FROM sessions;'. Review the audit
log (workspace settings, and audit_events in the database) for user.signed_in,
member.*, webhook.* and integration.* actions. Rotate APP_ENCRYPTION_KEY (above) only if
the key itself may have leaked. Afterwards, write a short post-mortem: timeline, impact, cause,
follow-ups.
Scaling
One API instance handles a small-to-medium team comfortably. Before adding instances, know what's per-instance:
- Live events (SSE) are per instance. An event reaches only the browsers connected to the
instance that produced it. With two API instances, users miss live updates made through the
other one until they reload (sticky sessions don't fix this, since the writer and the reader
are different users). Multi-instance live events need a shared bus, e.g. Postgres
LISTEN/NOTIFY(future work). - Background jobs run in every instance, and are safe: the email digest takes a Postgres
advisory lock, so only one instance sends it. The webhook delivery worker claims rows with
FOR UPDATE SKIP LOCKED, so a delivery goes out once. - Rate limits are in memory, per instance, and reset on restart. With N instances the effective limit is up to N times higher.
- Database connections: each API instance opens up to 20. Postgres allows 100 by default,
so add a pooler (PgBouncer in transaction mode) or raise
max_connectionspast ~4 instances. - Uploads stage on the instance's local disk, then go to S3. Use
STORAGE_BACKEND=s3with more than one instance (local storage isn't shared). - Web is stateless and scales freely. Every web instance needs
API_URLpointing at the API (a load balancer address when there are several). - Scale vertically first (CPU/RAM for the VPS and Postgres). That avoids every caveat above.
Routine maintenance
- Weekly: glance at
pvc ps, disk usage, Sentry, and that last night's backup exists. - Monthly: restore test (above);
pvc pull db caddy && make prod-upfor base image patches; OS updates (unattended-upgradeshandles security fixes). - Automatic: TLS renewals (Caddy), backup pruning (
BACKUP_RETENTION_DAYS), audit log pruning (365 days). - Yearly: review who holds production credentials; rotate the R2 token and SMTP password.