I made my own social media network federated via ActivityPub on the Fediverse running Mastodon 4.1 a few years ago. I did it with a script that installed everything on an Oracle Cloud VPS. I did not use the Docker method because I thought that might be extra RAM that could be used elsewhere. Then I became lazy about learning how to upgrade Mastodon on an installation built from source. Reading the release notes and documentation for each Mastodon upgrade involved lots of special tweaks and database migrations that sounded pretty dangerous. I was worried about breaking it completely, so I kept putting it off. I thought about paying a friend to help too. Then I decided to give Claude Code SSH access to the server and let the AI evaluate what needed to be done.
The Opus 5 AI model looked at my server, scanned to see how it was configured, and developed the below runbook as a plan of attack. After reading through it and deciding that it all sounded good, I decided to just let Claud Opus proceed. I was using the Claude Code plugin with Visual Studio Code and it all took quite a long time. Several times it seemed like Claude had stalled or frozen as there was no way for it to display progress on some of the procedures that it was running. So I had to “resume” the runbook process several times. This was totally fine as we had built out the runbook plan as an HTML web page ahead of time.
social.lein.us · upgrade runbook · rev. 2 · 5 September 2026, 23:10 UTC
Lein Social Upgrade Runbook
Moving a three-year-old source install of Mastodon to current, by containerising it — because Ubuntu 22.04 can no longer supply the libraries Mastodon 4.7 needs.
→
4.7.1
six release hops · six sets of DB migrations
- Host
- 1 vCPU5.8 GB RAM, aarch64, 4 GB swap
- Disk
- 19 GB free27 of 45 GB used
- Database
- 2.9 GBPostgreSQL 15.19
- Local media
- 7.01 MBirreplaceable
- Remote cache
- ~7.7 GBregenerable
- Accounts
- 4 local28,650 known
00Where this stands
This revision records what has actually been executed, with measured results. Figures throughout have been replaced with real ones; where a first-pass estimate turned out wrong, the correction is called out rather than quietly overwritten.
/home/ubuntu/mastodon-backup-2026-09-05/ — a 446 MB pg_dump -Fc, an 18 MB local-media tarball, Adam’s row exports, all config including .env.production, and a SHA-256 manifest over 22 files.
It was restored into a scratch database and counted: Adam’s 941 statuses, 30 attachments, 795 following, 360 followers and all 4 local users came back exactly, with post text intact. pg_restore exited 0.
It is still on the same disk it was taken from. Until a copy exists on other hardware, this is not yet a backup in the sense that matters.
Corrections to revision 1
| Rev 1 said | Actually |
|---|---|
public/system is roughly 24 GB, mostly orphaned files | It was ~24 GB pre-prune, but not from orphans — a full remove-orphans scan found 3 files, 243 KB. The bulk was live remote cache. Total media now ~7.7 GB. |
| Expect Phase 0 to recover 15–20 GB | True in aggregate — 33 GB used fell to 27 GB — but almost all of it came from one accounts prune. The later passes reclaimed essentially nothing. |
| Gate: >25 GB free before continuing | The gate was wrong. It was sized against the mistaken 24 GB estimate. Phase 2/3 needs ~6 GB (images + a Postgres copy). 19 GB free clears it comfortably. |
| Adam has 359 followers | 360. He gained one during the day; the verified backup reflects 360. |
| Instance holds 35,195 statuses / 55,482 accounts | Post-prune: ~29,000 statuses / 28,650 accounts. Local data unchanged throughout. |
01What is running today
The instance is healthy and serving — /health returns 200 on ports 3000, 4000 and public HTTPS — but every layer of it is three years old, and the checkout is not on a release tag.
The code at /home/mastodon/live sits on branch main at commit 9d75b03ba (22 April 2023) — a development snapshot, not v4.1.2, even though version.rb reports 4.1.2. The database is the reliable signal: its newest applied migration is 20230215074423, which is squarely 4.1-era schema. The upgrade path below starts from that schema, not from the code.
| Component | Installed | Needed by 4.7.1 | Available on Ubuntu 22.04 |
|---|---|---|---|
| Mastodon | 4.1.2 main snapshot | 4.7.1 | — |
| Ruby | 3.2.2 | 4.0.6 | Source build; local ruby-build last updated Apr 2023 stale |
| Node.js | 16.20.2 | 24.19 (engines ≥22) | NodeSource available |
| Yarn | 1.22.19 | 4.18.0 (corepack) | available |
| PostgreSQL | 15.19 | ≥ 14 | meets requirement |
| Redis | 6.0.16 | ≥ 7.0 | Jammy ships 6.0.16 only blocked |
| ffmpeg | 4.4.2 | ≥ 5.1 | Jammy ships 4.4.2 only blocked |
| libvips | not installed | 8.18.x | Jammy ships 8.12.1 blocked |
| ImageMagick | 6.9.11 | removed in 4.6 | — |
Other findings from the inspection
- There were no database backups at all.
/var/backupsheld only routine OS state. addressed — one verified backup now exists, but there is still no recurring backup and no off-box copy. - The media purge job has never run. The mastodon crontab contains
0 5 * * 1 purge-media.sh— a bare filename with no path, which cron cannot resolve. The script also callsbin/tootctlwithout the rbenv shims onPATH. Still unfixed; see section 07. - nginx serves static assets straight off disk.
/etc/nginx/conf.d/social.lein.us.confsetsroot /home/mastodon/live/publicand usestry_files $uri =404for/assets/,/packs/,/emoji/,/sounds/,/avatars/,/headers/,/shortcuts/and/sw.js. Those all need changing for a container deployment — see Phase 4. - Sidekiq runs 25 threads on one core (
-c 25,DB_POOL=25). Worth trimming as part of this work. - The box reboots for kernel updates unattended — it moved from
6.8.0-1058-oracleto-1060mid-session on 5 September. Anything long-running must survive that; run it undersystemd-run, not a login shell. - SSH is being brute-forced (
invalid user testuser from 163.7.3.241). Ordinary background noise, but fail2ban is cheap. - Ubuntu 22.04 standard support ends April 2027, so the OS itself is a fixed deadline sitting behind all of this.
- TLS is fine: Let’s Encrypt for
social.lein.us, valid to 7 November 2026.
02The decision: yes, move to Docker Compose
You asked whether containerising would make upgrades easier. On this host it is not just easier — it is the difference between a routine upgrade and a source-compilation project.
Mastodon 4.6 dropped ImageMagick and made libvips mandatory, and raised the ffmpeg floor to 5.1. Mastodon 4.5 raised the Redis floor to 7.0. Ubuntu 22.04 cannot supply any of those three. Mastodon’s own Dockerfile resolves this by compiling libvips 8.18.5 and ffmpeg 9.0.1 from source inside the image.
So a native upgrade on this box means building Ruby 4.0.6, Node 24, libvips and ffmpeg from source, on a single ARM core with 5.8 GB of RAM, and then doing it again at every future release. The prebuilt image removes all of it — and ghcr.io/mastodon/mastodon:v4.7.1 and mastodon-streaming:v4.7.1 both publish linux/arm64 manifests, verified against the GHCR manifests.
Change one image tag, docker compose pull, up -d, run --rm web rails db:migrate. Minutes, not an evening. Rolling back is changing the tag back.
The shape of the target stack
Now — everything on the host
After — nginx on host, rest in compose
nginx and certbot stay on the host. Compose publishes on 127.0.0.1:3000 and 127.0.0.1:4000 — exactly where the existing backend and streaming upstreams already point, so the proxy blocks need no edits and TLS renewal keeps working untouched.
Postgres and Redis both move into containers. Leaving Postgres on the host would work, but it re-creates the problem this migration exists to solve: Mastodon would stay coupled to what the OS ships, and the April 2027 Ubuntu deadline would still be blocking. Redis has no choice — 6.0.16 is below 4.5’s floor and jammy has nothing newer.
Because Postgres is migrated by dump and restore into a fresh container volume, the host’s PostgreSQL 15 cluster is never written to. It sits there, untouched, still holding the 4.1 schema. Rollback is docker compose down, systemctl start mastodon-*, revert one nginx line. That is worth the ~3 GB of duplicated data.
Alternatives considered
| Option | Verdict | Why |
|---|---|---|
| In-place native upgrade | Not recommended | Requires compiling Ruby, Node, libvips and ffmpeg from source on 1 vCPU. High risk of OOM mid-build, and every future upgrade repeats it. |
| Docker Compose, same host | Recommended | Prebuilt arm64 images, clean rollback, decouples Mastodon from the OS. Plan below assumes this. |
| Rebuild on a new VM | Good, if you’ll pay for it | Cleanest outcome — current Ubuntu, Docker from day one, old box stays live as the rollback. Costs a second VM and a DNS cutover. The phases below transfer almost unchanged if you choose this. |
03The upgrade chain
You cannot go 4.1 → 4.7 in one jump. Each release carries its own migrations and several carry hard gates. Take the latest patch of each branch in order, running db:migrate at every stop. Under Docker each hop is a tag change, so the chain is cheap.
| Hop | Gate you must clear | Risk |
|---|---|---|
| v4.2.29 | Ruby ≥3.0, Node ≥16, PG ≥10. Adds periodic update checks. Drops streaming clustering. | low |
| v4.3.23 | Requires three new encryption secrets (ACTIVE_RECORD_ENCRYPTION_*). Ruby ≥3.1, Node ≥18, PG ≥12. Docker image splits into web + streaming. yarn 1 → 4. | high |
| v4.4.24 | Redis ≥6.2, PG ≥13, Ruby ≥3.2, Node ≥20. Redis namespaces dropped. Adds a fasp Sidekiq queue. | medium |
| v4.5.17 | Redis ≥7.0, PG ≥14, Node ≥20.19. Sidekiq updated — health-check shape changes. | medium |
| v4.6.7 | Ruby ≥3.3, Node ≥22, ffmpeg ≥5.1. ImageMagick removed, libvips required. Custom themes need updating. | medium |
| v4.7.1 | Assets recompilation. Unusually long database migrations — account uri uniqueness constraint, ActivityPub identity rework. | medium |
The 4.3 encryption secrets are generated once and must then be kept forever, alongside SECRET_KEY_BASE and OTP_SECRET. Lose them after the fact and you lose access to everything encrypted with them — including users’ stored 2FA secrets. Generate them in Phase 3, write them into .env.production, and copy that file off the box before continuing.
The database is small — ~29,000 statuses, 28,650 accounts, 2.9 GB — and the measured dump/restore round trip was under six minutes, so the migrations themselves should be minutes rather than hours. On one core, still budget generously: plan a 4-hour window. The instance has four users, so a maintenance window is a social non-event.
04Execution
Reclaim disk and add swap
done · 5 Sep, no downtime
Swap: the clear win. 4 GB file, chmod 600, persisted in /etc/fstab, plus vm.swappiness=10 — the default 60 pages out too eagerly on a database host. It was holding ~200 MB within the hour.
Pruning: mostly already done. An accounts prune run earlier the same day took accounts from 55,482 → 31,294 and disk from 33 GB used → 24 GB. The formal sequence afterwards added little:
media remove --days 7 → 0 files · preview_cards remove --days 30 → 0 files · accounts prune → 2,709 accounts · media remove-orphans → 3 files, 243 KB. All exited 0.
The reason the age-based passes found nothing is worth carrying forward: this instance’s cache is dominated by recent media. Attachments grew from 2.96 GB to 4.35 GB in a single evening of federation. A --days 7 cutoff has almost nothing to bite on, because anything older had already gone. Reclaiming that pool needs --days 1 or --days 3, at the cost of re-downloading as people browse.
# Swap — do this first; 5.8 GB with none is too tight for a restore + migrations
sudo fallocate -l 4G /swapfile && sudo chmod 600 /swapfile
sudo mkswap /swapfile && sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
echo 'vm.swappiness=10' | sudo tee /etc/sysctl.d/99-swappiness.conf
sudo sysctl -q vm.swappiness=10
# tootctl needs the rbenv shims on PATH — they are not there by default
TC="sudo -u mastodon env RAILS_ENV=production \
PATH=/home/mastodon/.rbenv/shims:/usr/local/bin:/usr/bin:/bin \
/home/mastodon/live/bin/tootctl"
$TC media usage # baseline first
nice -n 15 ionice -c2 -n7 $TC media remove --days 7 --concurrency 2
nice -n 15 ionice -c2 -n7 $TC preview_cards remove --days 30 --concurrency 2
nice -n 15 ionice -c2 -n7 $TC accounts prune --concurrency 2
nice -n 15 ionice -c2 -n7 $TC media remove-orphans # slowest, run lastEvery long job here must survive a dropped connection or an unattended kernel reboot — both happened during this work. Wrap them:
sudo systemd-run --unit=mastodon-prune --collect --service-type=exec \
bash -c '/path/to/script.sh > /home/ubuntu/prune.log 2>&1'Custom emoji — the largest remaining pool
79,705 remote custom emoji, 1,865 MB, and zero local ones, so nothing of this instance’s own identity is at stake. None of the four commands above touch emoji; it needs its own pass:
$TC emoji purge --remote-onlyBudget properly for it. Mastodon destroys each record individually along with its file, and the measured rate on this box is roughly 200 records per minute — about six hours for the full set. There is no faster safe path: a raw SQL delete would leave custom_emoji_categories and announcement_reactions references dangling. Remote emoji re-download as new posts reference them; older posts show the shortcode until then.
Back everything up, and verify the backup
done · 4 m 05 s downtime
Captured 21:57:34–22:01:39 UTC with services stopped, verified by restore at 22:03–22:07.
mastodon_2026-09-05.dump 446 MB · local-media_2026-09-05.tar.gz 18 MB (395 files) · Adam’s nine CSV exports · all config · MANIFEST.txt with SHA-256 over 22 files.
Restored into a scratch database: 941 statuses, 30 attachments, 795 following, 360 followers, 4 local users — all exact. Identity intact, post text verbatim, pg_restore rc=0, 904 TOC entries.
An unverified backup is not a backup. The restore test is the step that makes the rest of this plan safe to attempt — it is not optional garnish.
1 · Adam’s own exports, from the web UI
Log in as Adam at https://social.lein.us/settings/export and download the CSVs and the archive. This is the one step that needs his login. See section 05 for what those can and cannot restore. still outstanding
2 · Stop the services and dump
D=$(date +%F)
STAGE=/var/backups/mastodon
sudo mkdir -p $STAGE && sudo chown postgres:postgres $STAGE
# Record counts BEFORE stopping, as the verification baseline
sudo -u postgres psql -d mastodon_production -c \
"SELECT count(*) FROM statuses WHERE account_id=110248000680880933;"
# Let Sidekiq drain — the Redis queue is not migrating
sudo systemctl stop mastodon-web mastodon-streaming
sleep 60
sudo systemctl stop mastodon-sidekiq
sudo -u postgres pg_dump -Fc -Z6 -d mastodon_production -f "$STAGE/mastodon_${D}.dump"
# Local media only — 7.01 MB of it. The remote cache is regenerable, skip it.
sudo tar czf "$STAGE/local-media_${D}.tar.gz" \
-C /home/mastodon/live/public/system accounts media_attachments site_uploads
sudo systemctl start mastodon-web mastodon-sidekiq mastodon-streamingPuma takes ~50 s to preload after restart. A health check straight after systemctl start returns 502 and looks like a failure. It is not — wait for :3000/health before concluding anything. Most of the 4-minute outage was this, not the dump.
Wrap the whole sequence in a trap so a failure cannot leave the site down: trap 'systemctl start mastodon-web mastodon-sidekiq mastodon-streaming' ERR INT TERM. This fired for real when the session was killed mid-dump, and it is why the site came back on its own.
3 · Prove the dump restores
sudo -u postgres createdb mastodon_restoretest
sudo -u postgres pg_restore -d mastodon_restoretest < "$STAGE/mastodon_${D}.dump"
sudo -u postgres psql -d mastodon_restoretest -c \
"SELECT count(*) FROM statuses WHERE account_id=110248000680880933;"
# must return 941
sudo -u postgres dropdb mastodon_restoretestPass the dump on stdin, not as a path. /home/ubuntu is mode 750, so the postgres system user cannot traverse it and pg_restore /path/to.dump fails with a bare “Permission denied”. Streaming it in sidesteps that entirely.
The verified backup currently lives on the same 45 GB disk it was taken from. scp it somewhere else and confirm the SHA-256 values in MANIFEST.txt match at the far end. This is the only genuinely irreversible risk left in the plan.
04bWhat remains
Build the compose stack
~45 min, no downtime
# Docker CE from the official repo (arm64)
sudo install -m 0755 -d /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg \
| sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
echo "deb [arch=arm64 signed-by=/etc/apt/keyrings/docker.gpg] \
https://download.docker.com/linux/ubuntu jammy stable" \
| sudo tee /etc/apt/sources.list.d/docker.list
sudo apt-get update
sudo apt-get install -y docker-ce docker-ce-cli containerd.io \
docker-buildx-plugin docker-compose-plugin
sudo mkdir -p /opt/mastodon/public/system
sudo cp -a /home/mastodon/live/public/system/. /opt/mastodon/public/system/
sudo chown -R 991:991 /opt/mastodon/public # uid/gid the image runs asThat cp -a copies the whole media tree — currently ~7.7 GB and several hundred thousand files, so it is slow on this disk. Run it under systemd-run like everything else, and check df first: it roughly doubles media usage until the old tree is removed.
/opt/mastodon/docker-compose.yml
Adapted from upstream’s file. Two deliberate departures: postgres:15-alpine, not the upstream default of 14 — you are restoring a PostgreSQL 15 dump, and matching versions removes a variable; and password auth on the database rather than POSTGRES_HOST_AUTH_METHOD=trust.
services:
db:
image: postgres:15-alpine
restart: always
shm_size: 256mb
networks: [internal_network]
healthcheck:
test: ['CMD', 'pg_isready', '-U', 'mastodon']
volumes:
- ./postgres15:/var/lib/postgresql/data
environment:
POSTGRES_USER: mastodon
POSTGRES_DB: mastodon_production
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
redis:
image: redis:7-alpine
restart: always
networks: [internal_network]
healthcheck:
test: ['CMD', 'redis-cli', 'ping']
volumes:
- ./redis:/data
web:
image: ghcr.io/mastodon/mastodon:v4.7.1
restart: always
env_file: .env.production
command: bundle exec puma -C config/puma.rb
networks: [external_network, internal_network]
healthcheck:
test: ['CMD-SHELL', "curl -s --noproxy localhost localhost:3000/health | grep -q 'OK' || exit 1"]
ports: ['127.0.0.1:3000:3000']
depends_on: [db, redis]
volumes:
- ./public/system:/mastodon/public/system
streaming:
image: ghcr.io/mastodon/mastodon-streaming:v4.7.1
restart: always
env_file: .env.production
command: node ./streaming/index.js
networks: [external_network, internal_network]
healthcheck:
test: ['CMD-SHELL', "curl -s --noproxy localhost localhost:4000/api/v1/streaming/health | grep -q 'OK' || exit 1"]
ports: ['127.0.0.1:4000:4000']
depends_on: [db, redis]
sidekiq:
image: ghcr.io/mastodon/mastodon:v4.7.1
restart: always
env_file: .env.production
command: bundle exec sidekiq -c 8
networks: [external_network, internal_network]
depends_on: [db, redis]
volumes:
- ./public/system:/mastodon/public/system
networks:
external_network:
internal_network:
internal: truesidekiq -c 8 rather than the 25 you run today — this is one core. Set DB_POOL=10 to match.
/opt/mastodon/.env.production — what changes and what must not
Start from config/env.production in the verified backup. Change only the connection settings; carry the secrets across verbatim.
# --- CHANGE these ---
DB_HOST=db
DB_PORT=5432
DB_NAME=mastodon_production
DB_USER=mastodon
DB_PASS=<the same value as POSTGRES_PASSWORD>
REDIS_HOST=redis
REDIS_PORT=6379
# delete REDIS_PASSWORD — the container Redis is on an internal-only network
# --- KEEP byte-for-byte from the old file ---
LOCAL_DOMAIN=social.lein.us
SECRET_KEY_BASE=...
OTP_SECRET=...
VAPID_PRIVATE_KEY=...
VAPID_PUBLIC_KEY=...
# and the whole SMTP_* block, unchanged
# --- ADD in Phase 3, before the 4.3 hop ---
ACTIVE_RECORD_ENCRYPTION_DETERMINISTIC_KEY=...
ACTIVE_RECORD_ENCRYPTION_KEY_DERIVATION_SALT=...
ACTIVE_RECORD_ENCRYPTION_PRIMARY_KEY=...
# --- sizing for one core ---
DB_POOL=10
WEB_CONCURRENCY=1
MAX_THREADS=5Changing SECRET_KEY_BASE logs every session out and invalidates Adam’s 15 live OAuth tokens. Changing OTP_SECRET breaks 2FA. Changing the VAPID pair breaks push subscriptions. There is no reason to regenerate any of them.
Restore, then walk the chain
~1–2 h, downtime
cd /opt/mastodon
docker compose up -d db redis
# wait for the db healthcheck, then restore into the fresh volume
docker compose exec -T db pg_restore -U mastodon -d mastodon_production \
--no-owner --role=mastodon < /home/ubuntu/mastodon-backup-2026-09-05/mastodon_2026-09-05.dumpThen hop the chain. For each tag in turn — v4.2.29, v4.3.23, v4.4.24, v4.5.17, v4.6.7, v4.7.1 — edit the three image tags and run:
docker compose pull
docker compose run --rm web bundle exec rails db:migrateGenerate the encryption secrets, add all three lines to .env.production, and copy that file off the box before running the 4.3 migrations:
docker compose run --rm web bin/rails db:encryption:initThen continue. Do not lose these.
After the final hop, rebuild what Redis used to hold — the old queue and feed data did not come across:
docker compose run --rm web bin/tootctl cache clear
docker compose run --rm web bin/tootctl feeds buildPoint nginx at the containers
~15 min
Two edits to /etc/nginx/conf.d/social.lein.us.conf. The upstream and proxy_pass blocks stay exactly as they are — compose already listens on the ports they name.
- Change both
root /home/mastodon/live/public;lines toroot /opt/mastodon/public; - Change every
try_files $uri =404;totry_files $uri @proxy;. Assets now live inside the image, not on disk — without this,/assets/,/packs/,/emoji/,/sounds/,/avatars/,/headers/,/shortcuts/and/sw.jsall return 404 and you get an unstyled site. The image setsRAILS_SERVE_STATIC_FILES=true, so Puma serves them; nginx keeps its cache headers. The config file’s own comment at line 66 says to do exactly this.
sudo cp /etc/nginx/conf.d/social.lein.us.conf /root/nginx-preswitch.conf # cheap undo
sudo sed -i 's#root /home/mastodon/live/public;#root /opt/mastodon/public;#' \
/etc/nginx/conf.d/social.lein.us.conf
sudo sed -i 's#try_files $uri =404;#try_files $uri @proxy;#' \
/etc/nginx/conf.d/social.lein.us.conf
sudo nginx -t && sudo systemctl reload nginxThen disable the old units so nothing races for the ports on reboot — but leave them installed, they are your rollback:
sudo systemctl disable --now mastodon-web mastodon-sidekiq mastodon-streaming
docker compose up -dVerify
~30 min
Against the numbers from the verified backup, so a discrepancy is visible rather than assumed.
| Check | Expected |
|---|---|
| curl -s https://social.lein.us/health | 200 / OK (allow ~50 s for boot) |
| curl -s …/api/v1/instance | grep version | 4.7.1 |
| Adam’s post count | 941 |
| Adam’s following / followers | 795 / 360 |
| Local accounts | 4 |
| Adam’s attached media renders | 30 items, 7.01 MB |
| Page styling and JS load | Confirms the nginx try_files change |
| Log in as Adam, post, delete | Session, compose and Sidekiq all working |
| Follow something remote, see it arrive | Federation both directions |
| Trigger a password-reset email | SMTP config carried across |
| docker compose logs –tail=100 | No repeating errors |
Leave the old systemd units and the host Postgres cluster in place for a fortnight before reclaiming the space.
05Adam’s account: backup, and the truth about restore
The account is Adam — capitalised, id 110248000680880933, admin role, created 23 April 2023. Mastodon resolves handles case-insensitively, so [email protected] is this account; there is no separate lowercase adam. It holds 941 statuses, 30 media attachments totalling 7.01 MB, 795 follows, 360 followers, 5 bookmarks, 2 favourites and 15 live OAuth tokens.
Mastodon has no import for the account archive. The .tar.gz you get from “Request your archive” is a portability and readability format — outbox.json, actor.json and media. No Mastodon version, including 4.7.1, can load it back. There is also no tootctl command that restores a single account.
So: the only thing that brings back Adam’s 941 posts is the full database restore. The per-account exports are a genuine but partial second net.
| Artefact | Contains | Restorable? |
|---|---|---|
| mastodon_2026-09-05.dump | Everything — all four accounts, all posts, follows, tokens | Yes — the real restore path |
| local-media_2026-09-05.tar.gz | Adam’s 30 attachments, avatars, headers, site uploads | Yes — untar into place |
| following_accounts.csv | 795 follows | Yes — Import, “overwrite” |
| bookmarks / lists / blocks / mutes / domain_blocks .csv | Those lists | Yes — Import |
| archive .tar.gz | All 941 posts as ActivityPub JSON, plus media | No import exists — readable reference only |
| account-adam/*.csv | Raw rows for Adam, plus handle-resolved follow lists | Manual — for reconciliation, not a restore path |
| Followers (360) | — | Not exportable — they are held on remote servers and re-federate over time |
Do not count records with wc -l. Post text contains newlines, so a single row often spans several physical lines — statuses.csv reads as 1,061 lines for 941 posts. Use a CSV reader, or count in SQL.
If a restore is actually needed
cd /opt/mastodon
docker compose stop web sidekiq streaming
docker compose exec -T db dropdb -U mastodon mastodon_production
docker compose exec -T db createdb -U mastodon mastodon_production
docker compose exec -T db pg_restore -U mastodon -d mastodon_production \
--no-owner --role=mastodon < /home/ubuntu/mastodon-backup-2026-09-05/mastodon_2026-09-05.dump
sudo tar xzf /home/ubuntu/mastodon-backup-2026-09-05/local-media_2026-09-05.tar.gz \
-C /opt/mastodon/public/system
sudo chown -R 991:991 /opt/mastodon/public/system
# the dump is 4.1-schema — re-walk the chain from P3 before starting web
docker compose run --rm web bundle exec rails db:migrate
docker compose run --rm web bin/tootctl feeds build
docker compose up -dAnything federated in or posted since the dump is gone, and remote servers will still hold posts your database no longer knows about. For a four-person family instance that is a curiosity rather than a crisis, but decide deliberately rather than discovering it.
06Rollback
Available at every point up to and including Phase 4, and it is fast, because nothing in the plan writes to the host’s PostgreSQL cluster or to /home/mastodon/live.
cd /opt/mastodon && docker compose down
sudo cp /root/nginx-preswitch.conf /etc/nginx/conf.d/social.lein.us.conf
sudo nginx -t && sudo systemctl reload nginx
sudo systemctl enable --now mastodon-web mastodon-sidekiq mastodon-streaming
# then wait — Puma needs ~50 s before /health answersYou are back on 4.1.2 with the original database, minus anything posted during the window. Keep /opt/mastodon/postgres15 so you can diagnose the failure before retrying.
07Afterwards — closing the gaps that got you here
The upgrade is the occasion; these are what stop the same three years happening again.
- Get the backup off the box, then automate it. One verified backup now exists, but it sits on the disk it protects and nothing recreates it tomorrow. A cron line does the job:
docker compose exec -T db pg_dump -U mastodon -Fc mastodon_productionto a dated file, shipped elsewhere. Keep.env.productionalongside it — the dump is useless without the secrets. - Fix the purge job. It has never once run. Absolute path, and rewrite the body for containers:
0 5 * * 1 cd /opt/mastodon && docker compose run --rm web bin/tootctl media remove --days 14. Drop the--days 0 --include-followsheader purge from the old script; it re-downloads everything it deletes. Add anemoji purge --remote-onlypass on a slower cadence — monthly is plenty, and it takes hours. - Pruning is a holding action, not a fix. The cache regrew 1.4 GB in one evening. Whatever cadence you choose, expect to run it forever; the alternative is moving media to object storage.
- Keep the swap file. It cost 4 GB and was in use within the hour.
- Run long jobs under
systemd-run. This box takes unattended kernel reboots — it did one mid-session — and login shells do not survive them. - Upgrade cadence. Now that it is a tag change, take patch releases as they land. 4.2 onwards will tell you when one is available, and 4.7 warns when your version leaves support.
- The OS is still on the clock. Ubuntu 22.04 loses standard support in April 2027 — but with Mastodon in containers, that becomes an ordinary
do-release-upgraderather than a rebuild.
Decisions still open
- Same host or new VM? The plan assumes same host. A new VM is cleaner and keeps the old box as a live rollback, at the cost of a second machine and a DNS cutover.
- Where do backups go? Unanswered, and now the single largest risk — there is something worth protecting and nowhere off-box to put it.
- Rotate the
mastodonaccount password. It was shared in a chat transcript during this work and was never needed; passwordless sudo was already available.
And then when finished, we updated the Runbook status:
Lein Social Upgrade Runbook
A three-year-old source install is now Mastodon 4.7.1 in containers on Ubuntu 24.04 LTS. This revision is the record of how, and the reference for running it.
00 Status
4 m 05 s to take the verified backup · 10 m 50 s for the cutover to containers · ~90 s for the 24.04 reboot.
The six-hop upgrade chain and the whole OS release upgrade ran with the site live. That is the payoff from containerising first: the host was rewritten underneath a running Mastodon.
01 What runs now
Everything Mastodon lives under /opt/mastodon. nginx and certbot stay on the host and were never rebuilt — compose publishes on the same 127.0.0.1:3000 and :4000 the original upstreams already pointed at.
Host
Containers — /opt/mastodon
| Path | What it is |
|---|---|
| /opt/mastodon/docker-compose.yml | The stack. Image tags come from MASTODON_VERSION in .env. |
| /opt/mastodon/.env | MASTODON_VERSION and POSTGRES_PASSWORD. Mode 600, root. |
| /opt/mastodon/.env.production | Mastodon config and all secrets. Mode 600, root. |
| /opt/mastodon/postgres15/ | Database volume. |
| /opt/mastodon/public/system/ | Media. Owned 991:991 (the image’s uid). |
| /opt/mastodon/redis/ | Redis volume. |
| /usr/local/bin/mastodon-maintenance.sh | Cache maintenance, run by cron. |
| /home/ubuntu/mastodon-backup-2026-09-05/ | The verified backup. Also copied off-box. |
| /root/pre-noble-config/ | Config snapshot taken before the OS upgrade. |
Puma takes 60–105 s to preload. A health check straight after any restart returns 502 and looks broken. Wait for :3000/health before concluding anything. The compose healthcheck has start_period: 180s so it no longer reports false “unhealthy”.
nginx serves /system/ from disk and proxies everything else. All eight static blocks use try_files $uri @proxy. If they ever revert to =404 you get a site that loads with no CSS or JS.
Run long jobs under systemd-run, never a login shell. This box takes unattended kernel reboots and SSH sessions drop. Every long operation in this migration was wrapped: sudo systemd-run --unit=NAME --collect --service-type=exec bash -c '…'.
02 Operating it
Upgrading Mastodon
This is the whole point of the exercise. A release is now a one-line edit.
cd /opt/mastodon
sudo sed -i 's/^MASTODON_VERSION=.*/MASTODON_VERSION=v4.8.0/' .env
sudo docker compose pull
sudo docker compose run --rm web bundle exec rails db:migrate
sudo docker compose up -d
# then wait ~90 s and check
curl -s https://social.lein.us/api/v1/instance | grep -o '"version":"[^"]*"'Rolling back is the same edit in reverse, provided the new version ran no migrations. Once migrations have applied, going back means restoring a dump — so take one first for a major release.
Minimum PostgreSQL, Redis, Ruby and Node versions rise over time, and a release occasionally demands a new secret — 4.3 required three. The image carries Ruby, Node, libvips and ffmpeg, but PostgreSQL and Redis are yours to keep current: they are pinned in docker-compose.yml at postgres:15-alpine and redis:7-alpine.
Backups
# database — takes ~2 min, no downtime needed (pg_dump is a consistent snapshot)
cd /opt/mastodon
sudo docker compose exec -T db pg_dump -U mastodon -Fc mastodon_production \
> ~/mastodon_$(date +%F).dump
# local media — small; the remote cache is regenerable, skip it
sudo tar czf ~/local-media_$(date +%F).tar.gz \
-C /opt/mastodon/public/system accounts media_attachments site_uploads
# the secrets, without which the dump is useless
sudo cp /opt/mastodon/.env.production ~/env.production_$(date +%F)Prove it, into a scratch database, and count something you know:
sudo docker compose exec -T db createdb -U mastodon restoretest
sudo docker compose exec -T db pg_restore -U mastodon -d restoretest < ~/mastodon_DATE.dump
sudo docker compose exec -T db psql -U mastodon -d restoretest -At \
-c "SELECT count(*) FROM statuses WHERE account_id=110248000680880933;"
sudo docker compose exec -T db dropdb -U mastodon restoretestAnd get it off this machine. A backup on the disk it protects only survives the failures that were never going to matter.
Cache maintenance
Runs itself. /usr/local/bin/mastodon-maintenance.sh, logging to /var/log/mastodon-maintenance.log (rotated monthly), both entries flock-guarded so they cannot overlap.
| Schedule | Does | Why that cadence |
|---|---|---|
| 0 3 * * * | media remove --days 2preview_cards remove --days 7 | This instance ingests ~5.6 GB/day of remote media. At 2 days the cache settles near ~11 GB; at 3 days it was ~17 GB. |
| 0 4 1 * * | accounts prunestatuses remove --days 30media remove-orphans | Orphan scan is a full-disk walk (~10 min) that has found 0–1 files in practice, so it is monthly rather than nightly. |
emoji purge --remote-only. Measured 5 Sep: purging all 79,705 remote emoji reclaimed 1.82 GB — and it was fully back within 36 hours as they re-federated. It buys a day or two of space for the cost of re-downloading everything.
media remove --remove-headers --include-follows --days 0, which the original purge-media.sh ran. It deletes avatars and headers for accounts you actively follow, which Mastodon then immediately re-fetches. It spends bandwidth to reclaim space it gives straight back.
Restoring
cd /opt/mastodon
sudo docker compose stop web sidekiq streaming
sudo docker compose exec -T db dropdb -U mastodon mastodon_production
sudo docker compose exec -T db createdb -U mastodon mastodon_production
sudo docker compose exec -T db pg_restore -U mastodon -d mastodon_production \
--no-owner --role=mastodon < /path/to/mastodon_DATE.dump
sudo tar xzf /path/to/local-media_DATE.tar.gz -C /opt/mastodon/public/system
sudo chown -R 991:991 /opt/mastodon/public/system
sudo docker compose run --rm web bundle exec rails db:migrate
sudo docker compose run --rm web bin/tootctl feeds build
sudo docker compose up -d/home/ubuntu is mode 750, so a postgres user cannot traverse it and pg_restore /path/to.dump fails with a bare “Permission denied” that looks like a corrupt file. The < redirect runs as your shell and sidesteps it. This cost real time to diagnose.
Outbound goes through OCI Email Delivery, smtp.email.us-ashburn-1.oci.oraclecloud.com:587, STARTTLS, verified end to end on 7 Sep.
SMTP_SSL and SMTP_TLS must be false with SMTP_ENABLE_STARTTLS=always. The previous provider used 465 with SMTP_SSL=true; carrying that over to 587 silently breaks all mail. Also note OCI only accepts mail from an approved sender — authentication succeeding proves nothing about whether it will deliver.
03 What was done
Chronological, with measured results. Times are UTC.
| When | Step | Result |
|---|---|---|
| 5 Sep 16:56 | Inspection | Found 4.1.2 on main@9d75b03ba — a dev snapshot, not a release tag. No backups existed at all. |
| 5 Sep 18:24 | Swap + prune | 4 GB swap, swappiness=10. Accounts 55,482 → 31,294; disk 33 → 24 GB used. |
| 5 Sep 21:57 | Backup | 446 MB dump + 18 MB media. 4 m 05 s downtime. |
| 5 Sep 22:07 | Restore-verified | pg_restore rc=0, 904 TOC entries. 941/30/795/360 and 4 users exact; post text verbatim. |
| 5 Sep 23:08 | Emoji purge | 79,705 → 0 in 22 m 34 s. Zero local emoji existed to lose. |
| 5 Sep 23:19 | Docker installed | 29.8.0 + Compose v5.5.1, arm64. Container networking verified against ufw. |
| 6 Sep 00:47 | Restore into containers | 4 m 43 s. Counts matched exactly; schema at 4.1-era 20230215074423. |
| 6 Sep 00:48–01:05 | Six-hop chain | 17 minutes, site live throughout. Every hop gated on 941 statuses / 4 users. |
| 6 Sep 01:21 | Cutover | 10 m 50 s downtime. Web up 65 s after container start; public 200 immediately after. |
| 7 Sep 13:43 | Mail moved to OCI | Accepted by relay, delivered. |
| 7 Sep 14:55 | Old install removed | Host PG, Redis, Node 16, ffmpeg, ImageMagick, /home/mastodon, 3 units, the mastodon account. 7 health gates, no drop. |
| 7 Sep 15:12 | Ubuntu 24.04 | rc=0 in 12 minutes, Mastodon live throughout. |
| 7 Sep 15:20 | Reboot | Kernel 6.17.0-1020-oracle. Site answered 200 10 s after boot. 0 failed units. |
The six-hop chain
4.1 → 4.7 is not one jump. Each release carries migrations, and several carry hard gates. Under Docker each hop is a tag change, which is what made it 17 minutes instead of an evening.
| Hop | Schema after | Gate it carried |
|---|---|---|
| v4.2.29 | 20230907150100 | Ruby ≥3.0, Node ≥16, PG ≥10 |
| v4.3.23 | 20241007071624 | Three new encryption secrets. Image splits web/streaming. yarn 1 → 4. |
| v4.4.24 | 20250627132728 | Redis ≥6.2, PG ≥13. Redis namespaces dropped. |
| v4.5.17 | 20251023210145 | Redis ≥7.0, PG ≥14 |
| v4.6.7 | 20260611150940 | ImageMagick removed, libvips required, ffmpeg ≥5.1 |
| v4.7.1 | 20260812154114 | Long migrations: account uri uniqueness, ActivityPub identity rework |
ACTIVE_RECORD_ENCRYPTION_PRIMARY_KEY, _DETERMINISTIC_KEY and _KEY_DERIVATION_SALT live in .env.production and are in the backup. Lose them and everything encrypted with them is unrecoverable, including users’ stored 2FA secrets. They were generated and backed up before the 4.3 migrations ran, deliberately in that order.
Why containerising was the unblocker, not a preference
Mastodon 4.6 dropped ImageMagick and made libvips mandatory, and raised ffmpeg to ≥5.1. Mastodon 4.5 raised Redis to ≥7.0. Ubuntu 22.04 shipped libvips 8.12.1, ffmpeg 4.4.2 and Redis 6.0.16 — none of them qualified, and no newer versions existed in its repositories. Mastodon’s own Dockerfile resolves this by compiling libvips 8.18.5 and ffmpeg 9.0.1 from source inside the image.
A native upgrade meant building Ruby 4.0.6, Node 24, libvips and ffmpeg from source on a single ARM core — and repeating it at every release. The prebuilt image removed all of it. It also meant the 24.04 upgrade two days later was an ordinary do-release-upgrade that never took Mastodon down.
04 What surprised us
Kept because the reasoning matters more than the conclusions, and because several of these were wrong the first time.
| Believed | Actually |
|---|---|
public/system held ~24 GB, largely orphaned files | A full remove-orphans scan found 3 files, 243 KB. It was all live remote cache. There was never an orphan problem. |
| Age-based pruning would reclaim the disk | --days 7 and even --days 2 removed zero bytes. Nothing is stale — the cache is entirely recent, because inflow is ~5.6 GB/day. The disk is simply smaller than the firehose. |
| Emoji purging is worth scheduling | 1.82 GB reclaimed, fully regrown in 36 hours. Dropped from the schedule. |
| The emoji purge would take ~6 hours | 22 minutes. The first estimate extrapolated from 50 seconds that were mostly Rails booting. |
| Gate the migration on >25 GB free | That number was sized against the mistaken 24 GB estimate. Phase 2/3 needed ~6 GB. The gate was wrong, not the outcome. |
The account was adam | Adam, capitalised. Mastodon matches handles case-insensitively, so both resolve — but SQL does not. |
| The 22.04 → 24.04 upgrade would take 45–90 min | 12 minutes, and it never interrupted Mastodon. |
do-release-upgrade uninstalled ufw and left INPUT policy: ACCEPT where it had been DROP. Exposure did not change in practice — only 22/80/443 listen externally — but default-deny was gone, and nothing announced it. Reinstalled and re-enabled from the preserved /etc/ufw rules. Check your firewall after any release upgrade. It also disabled the Docker apt repo (renamed to docker.list.distUpgrade), which would have quietly stopped Docker security updates.
05 Adam’s account, and the truth about restore
The account is Adam — capitalised, id 110248000680880933, admin. Mastodon resolves handles case-insensitively, so [email protected] is this account; there is no separate lowercase one. It came through the migration at 941 statuses, 30 attachments (7.01 MB), 795 following, 360 followers, and has since posted — 946 at the time of writing, which is the clearest evidence the new stack is genuinely in use.
Mastodon has no import for the account archive. The .tar.gz from “Request your archive” is a portability and readability format — outbox.json, actor.json, media. No version, 4.7.1 included, can load it back. There is also no tootctl command that restores one account.
So the only thing that brings back Adam’s posts is the full database restore. Per-account exports are a real but partial second net.
| Artefact | Restorable? |
|---|---|
| mastodon_2026-09-05.dump | Yes — the real path Everything: all accounts, posts, follows, tokens. |
| local-media_2026-09-05.tar.gz | Yes Untar into public/system, then chown 991:991. |
| following / lists / blocks / mutes / bookmarks .csv | Yes Via Import in the web UI. |
| archive .tar.gz | No import exists Readable reference only. |
| account-adam/*.csv | Manual Reconciliation aid, not a restore path. |
| Followers | Not exportable Held on remote servers; they re-federate over time. |
Post text contains newlines, so one record often spans several physical lines — statuses.csv reads as 1,061 lines for 941 posts. Use a CSV reader, or count in SQL.
06 Still open
One verified backup exists and has been copied off-box. Nothing recreates it tomorrow. The instance has been running two days past that snapshot. This is now the largest single risk, and it is one cron line plus somewhere to put the file — see section 02.
Nightly pruning at --days 2 should hold the cache near ~11 GB against ~5.6 GB/day of inflow. It is a treadmill: whatever cadence you pick, you run it forever, and a busier feed moves the equilibrium.
The durable fix is S3-compatible object storage for media, which Mastodon supports natively and which OCI provides. That removes the disk ceiling instead of managing it. The first genuinely productive nightly run is the night of 7–8 Sep; watch free space then to confirm it stabilises rather than climbs.
Adam’s own web-UI exports (Settings → Import and export) were never taken — the one step needing his login. Not gating anything, since the database restore is the real path.
fail2ban. SSH is being brute-forced from the open internet; ordinary background noise, but it is cheap insurance and key-only auth is already on.
Old kernels. 5.15.0-1030 and 6.8.0-1060 are still installed; autoremove will not take them while the linux-image-oracle meta-package holds a reference. Keeping one fallback kernel is sensible anyway.
The clock that is no longer ticking
The original deadline behind all of this was Ubuntu 22.04 losing standard support in April 2027. That is gone: 24.04 LTS is supported to 2029, and because Mastodon no longer depends on what the OS ships, the next release upgrade is an ordinary do-release-upgrade rather than a rebuild.
