On September 28, 2026, we moved the primary StatusPage.me production infrastructure from Contabo in Munich to Netcup in Nuremberg.
The migration was successful.
PostgreSQL moved to the new primary, DNS switched cleanly, public and custom-domain status pages came back over TLS, monitoring resumed, and we did not need to roll back.
It was also not one of those suspiciously perfect infrastructure migrations where every diagram survives contact with reality.
A maintenance date picker corrupted UTC timestamps.
A backup script still contained the PostgreSQL port from the old server.
Our first backup-retention design turned out to be structurally incapable of providing the guarantee we wanted.
A service failed because it could read its configuration file but could not traverse the directory containing it.
And after everything was running, a real restore drill uncovered leftover migration plumbing that was still quietly sitting around.
None of these caused data loss.
All of them were useful.
This is what we changed, what worked, what broke, and what we learned.
What we were moving
Before the migration, StatusPage.me’s main application and PostgreSQL primary ran in Munich.
The new topology moved production to Nuremberg and changed Munich’s role completely.
After the migration:
Nuremberg
- Production application
- PostgreSQL 18 primary
- Redis
- Scheduler
- Monitoring and server telemetry
- Backup producer
- Two-generation local backup retention
Munich
- No production application writers
- No production scheduler
- No production backup job
- Off-site encrypted backup receiver
- Rollback asset during the migration hold period
The database migration also moved us from PostgreSQL 17 to PostgreSQL 18.
Rather than performing a large dump-and-restore during the maintenance window, we used one-way PostgreSQL logical replication from Munich to Nuremberg.
That let most of the data move while production was still running.
The preparation mattered more than the cutover
Most of the dangerous work happened before the maintenance window.
That was intentional.
We audited PostgreSQL compatibility, rehearsed restores, measured connection usage, built sequence-reconciliation tooling, prepared a static maintenance configuration for the old primary, reduced DNS TTLs, and repeatedly tested the parts of the migration that would be hardest to improvise under pressure.
One useful discovery came from connection-pool analysis.
Several very different application components had effectively been using the same generic PostgreSQL connection-pool defaults.
That makes very little sense once you look at actual workloads.
A scheduler and a stateless web frontend do not need the same pool behavior.
We moved to component-specific connection limits before the migration and performed capacity testing against the new PostgreSQL server.
During the final cutover, the scheduler was running with a configured pool of 40 connections and total application sessions settled around 69, comfortably below the investigation threshold we had established during rehearsal.
Logical replication does not move your sequences
PostgreSQL logical replication copies table data.
It does not keep sequence state synchronized.
That is a rather important detail when your new primary contains rows with IDs that the local sequence does not yet know about.
StatusPage.me had 95 application-owned sequences.
We were not interested in handling those manually during a production freeze.
So we built a dedicated sequence-reconciliation tool.
It validates sequence ownership and definitions, checks the expected sequence set, confirms that application writers are frozen, confirms that the migration subscription has been disabled, generates the required reconciliation SQL, and performs read-only validation afterward.
It also fingerprints the ownership mapping so unexpected schema drift causes a hard stop rather than silently producing questionable SQL.
At cutover, all 95 sequences reconciled successfully.
That is exactly the kind of task that is much nicer to automate before midnight than to debug after midnight.
Preventing Munich from becoming a second writable production
DNS changes are not instantaneous.
Even with a low TTL, some clients can continue reaching an old address for a while.
That meant simply switching DNS and stopping Munich was not enough.
We wanted Munich to remain available for stale-DNS clients, while being completely incapable of forwarding application traffic to the old database.
We built a dedicated static Caddy maintenance configuration for the old primary.
During cutover, Munich served static HTTP 503 maintenance responses and no longer proxied StatusPage.me application traffic.
The generated configuration explicitly excluded the application backend ports while preserving unrelated services such as mail.
The TLS allowlist was generated from real StatusPage.me slugs, deleted-slug grace entries, and verified customer custom domains.
The important part was simple:
Munich could still answer requests, but it could no longer accept application writes.
That gave us a much cleaner boundary during DNS propagation.
The maintenance window
The public maintenance window was scheduled for:
September 27, 23:45 UTC → September 28, 01:45 UTC
T0 was midnight UTC.
The cutover sequence was roughly:
- Put Munich behind the static maintenance fence.
- Stop application writers, scheduler, server telemetry, failover watchdog, and backup cron on Munich.
- Allow PostgreSQL logical replication to fully catch up.
- Disable and retain the Nuremberg subscription.
- Reconcile all 95 sequences.
- Start the Nuremberg application stack.
- Switch DNS.
- Verify public TLS, account login, admin login, the StatusPage.me status page, and a real customer custom domain.
- Start scheduler and server telemetry on Nuremberg.
DNS was verified through the authoritative servers as well as multiple public resolvers.
We deliberately did not add an AAAA record during the migration.
One infrastructure change at a time remains an underrated strategy.
The production role reversal completed well inside the scheduled maintenance window, and the remaining time was used for validation and cleanup.
The real point of no return
The point of no simple return was not the DNS change.
It was not the maintenance page.
It was not even disabling logical replication.
It was the first durable write accepted by Nuremberg after replication stopped.
Before that write, Munich still represented the last complete source of truth and the migration could be abandoned relatively cleanly.
After it, the two databases could diverge.
From that moment onward, rolling back would require reconciling writes that existed only on Nuremberg.
That distinction was written into the migration plan explicitly.
It made the cutover much easier to reason about.
What went well
Quite a lot.
PostgreSQL replication caught up cleanly.
The subscription was disabled only after the target was fully synchronized.
Sequence reconciliation completed for all 95 sequences.
The Munich maintenance fence behaved as intended.
DNS propagated cleanly.
Public TLS worked.
Customer custom-domain TLS worked.
The application stack came up normally on Nuremberg.
Database connection usage remained comfortably within the capacity we had tested.
No database rollback was required.
No data was lost.
Most importantly, the cutover itself was relatively uneventful.
The interesting problems happened around it.
Our maintenance scheduler almost scheduled the wrong maintenance
Before the migration even began, we discovered that our maintenance scheduling form had a particularly unpleasant interpretation of UTC.
The form said the selected times were UTC.
The JavaScript library did not agree.
Bootstrap Date Range Picker produced browser-local Moment objects, and the old code then called .utc() on them.
That converted the selected wall-clock value through the browser’s timezone rather than treating the selected calendar fields as UTC.
A cross-midnight maintenance window became corrupted before it reached the backend.
The backend and PostgreSQL timezone handling were fine.
The bug was entirely in the UI.
We first repaired the conversion logic, then decided that this entire class of problem did not deserve another opportunity.
We removed the old range picker and replaced it with Air Datepicker using separate start and end datetime fields.
The rule is now intentionally boring:
If the form says 23:45 UTC, we store 23:45 UTC.
The browser timezone does not get a vote.
We also added server-side validation for past times and invalid maintenance windows.
Infrastructure work has a remarkable ability to expose unrelated bugs five minutes before you need the feature.
Our first post-cutover service failure was a Unix permission bit
After the main application was live, the server telemetry agent failed.
The configuration file itself had the expected permissions.
The service account simply could not traverse the directory containing it.
/etc/statuspage was adjusted to the correct group ownership and mode, and the service immediately recovered.
Heartbeat and metric ingestion resumed normally.
No database incident.
No networking issue.
No complicated distributed-systems mystery.
Just Unix permissions doing what Unix permissions have been doing since long before most of us were born.
Then we discovered our backup script still lived in Munich
Moving production also meant reversing the backup direction.
Before the migration:
Munich → off-site
After the migration:
Nuremberg → Munich
During that work, we found that our vendored AutoPostgreSQLBackup script still contained Munich’s old PostgreSQL port:
19897
Nuremberg runs PostgreSQL locally on:
5432
The script was corrected, tested, and committed.
That last part matters.
A production fix that exists only on one server is not really fixed.
A future deploy would otherwise have silently restored the old port.
Our backup-retention design failed twice
This was probably the most interesting part of the migration.
The goal sounded simple:
- Nuremberg should keep the newest two complete local backup generations.
- Munich should keep encrypted off-site backups for 14 days.
Our first managed-retention implementation assumed that one backup generation consisted of:
statuspostgres_globals
Production had other ideas.
The real backup job uses DBNAMES=all, which meant the complete set contained:
statuspostgres_globalspostgrestemplate1
Each also had its metadata sidecar.
We generalized the retention logic to derive the actual member set dynamically.
Problem solved.
Except it still was not.
The vendor rotation model could not provide the guarantee we wanted
AutoPostgreSQLBackup organizes its daily tier around weekday slots.
If multiple backups are run on the same weekday, the previous file in that weekday slot gets removed while the next backup is being created.
That meant our original idea of using daily/, weekly/, and monthly/ as the source of truth could never reliably guarantee two complete generations.
The underlying storage model and our safety guarantee were simply incompatible.
So we stopped trying to make the vendor directories do something they were not designed to do.
A generation store built with hardlinks
The final design introduced a separate local generation store:
local-generations/<generation>/
After a backup generation has been:
- created,
- encrypted,
- locally validated,
- transferred to Munich,
- checksum-verified remotely,
- and atomically promoted on the receiver,
the complete generation is captured from latest/.
Each file is hardlinked into a temporary generation directory.
The complete set is validated again.
Then the directory is atomically renamed into its final generation name.
Retention operates only on that store.
The newest two complete generations are kept.
Older generations are removed as whole units.
Incomplete generations never count.
The vendor’s daily/weekly/monthly rotation is no longer part of the guarantee.
Because the generation store uses hardlinks on the same filesystem, capturing a generation initially consumes essentially no additional data space.
If the vendor later removes its own pathname, the retained inode stays alive through local-generations/.
We proved this using real production data.
One generation’s vendor-side path was removed by later rotation.
Its link count dropped from two to one.
The data remained fully present through the generation store.
That was exactly the failure mode the design was intended to survive.
Proving retention with real backups
We ran multiple real production backups on Nuremberg.
Once three local generations existed, retention in dry-run mode produced exactly what we expected:
RETAIN complete-set=2026-09-28_16h44m
RETAIN complete-set=2026-09-28_15h04m
PRUNE complete-set=2026-09-28_14h06m
We then enabled actual pruning.
It removed exactly the oldest generation.
The two newest remained complete.
No temporary directories remained.
No unrelated backup paths changed.
Munich still retained all of its off-site copies.
The local prune freed roughly 7 GB, which matched the oldest generation becoming uniquely owned by the generation store after the vendor paths had already disappeared.
That is considerably more reassuring than seeing a unit test print PASS.
The first unattended backup
After the manual tests passed, we enabled the daily backup schedule on Nuremberg.
The first unattended run happened at 00:00 UTC the following night.
At 00:30:42 UTC, the success notification arrived.
The automated pipeline had:
- created the backup,
- encrypted it,
- generated metadata,
- transferred eight artifacts to Munich,
- verified them,
- promoted them,
- run local retention,
- and sent success notifications through both Telegram and email.
At that point the new backup topology was no longer something we had only demonstrated manually.
It was operating unattended.
But a backup is not proven until you restore it
There was still one uncomfortable open question.
We had proven:
- backup creation,
- encryption,
- signatures,
- transfer,
- checksums,
- remote promotion,
- local retention,
- and unattended execution.
We had not yet proven that a real artifact produced by the new pipeline could actually be decrypted and restored.
So on September 29, we did exactly that.
The first real off-site restore drill
We used the 2026-09-29_00h00m generation.
That was intentionally the backup created by the first real unattended cron run.
The decryption private key is not stored on Nuremberg.
It is not stored on Munich.
It lives only in the offline recovery bundle.
That rule remained intact during the restore drill.
The encrypted artifacts were copied from Munich to an offline Mac environment on an external drive.
We first restored the tiny template1 member to prove the mechanics.
Then we pulled:
postgres_globals- the main
statusbackup, roughly 7.2 GB encrypted
Both were decrypted using the offline recovery key.
Their signatures were verified.
gzip -t passed.
We then restored the globals first and loaded the full plain-SQL status dump into an isolated PostgreSQL 18 test cluster.
The restored database was roughly 58 GB.
The restore log contained zero errors.
Comparing the restored snapshot with live production
We then compared several important tables.
| Table | Restored 00:00 UTC snapshot | Production at verification | Difference |
|---|---|---|---|
| users | 231 | 233 | +2 |
| status_pages | 180 | 181 | +1 |
| monitors | 224 | 225 | +1 |
| monitor_results | 25,706,675 | 25,803,682 | +97,007 |
Every difference went in the expected direction.
Production was ahead of the backup snapshot because roughly 19 hours of real activity had occurred since the backup was taken.
We found no missing data.
No corruption.
No restore errors.
That closed the most important remaining question around the new backup system.
It is now not merely backup-proven.
It is restore-proven.
The restore drill found one more migration leftover
Naturally, the restore also found something.
The restored database still contained the disabled logical-replication subscription from the Munich→Nuremberg migration:
status_munich_migration_sub
The corresponding inactive slot still exists on Munich, along with the old migration publication.
They are no longer used.
They are now scheduled for cleanup as part of the next replication project.
It was a useful reminder that restores do more than prove backups.
They also recreate the exact state you have accumulated, including things you thought you were finished with.
What we learned
Several lessons survived the migration particularly well.
Rehearse the dangerous parts
The database restore rehearsal, capacity testing, sequence tooling, and old-primary fencing were all prepared before T0.
That turned the actual cutover into a sequence of already-understood operations instead of an architecture workshop conducted during downtime.
A backup file is not the same thing as recoverability
We redesigned retention twice.
Then we ran multiple production backup generations.
Then we let the system prune one for real.
Then we waited for an unattended backup.
Then we decrypted and restored that exact class of artifact.
Only after all of that were we comfortable calling the backup pipeline proven.
Static old infrastructure is useful during DNS propagation
Leaving Munich alive but unable to accept application writes gave stale-DNS clients a controlled maintenance response.
That was much safer than either returning connection errors or accidentally letting the old application continue to write.
Hardcoded topology hides everywhere
We found assumptions encoded in:
- PostgreSQL ports
- filesystem paths
- service configuration
- backup direction
- host aliases
- safety-control paths
Most of those assumptions were harmless for years.
Moving the infrastructure turned them into bugs.
The real point of no return is about data
DNS can be changed back.
Caddy can be reloaded.
Processes can be restarted.
The moment that matters is when the new primary begins accepting writes that the old one will never receive.
That is the point at which rollback changes from switching traffic to reconciling data.
Where we ended up
After the migration, StatusPage.me now runs with:
Nuremberg
- Production application
- PostgreSQL 18 primary
- Redis
- Scheduler
- Monitoring and telemetry
- Encrypted backup production
- Two complete local retained backup generations
Munich
- No application writers
- No production scheduler
- No backup producer
- Encrypted off-site receiver
- 14-day off-site retention
The decryption private key remains offline.
The first unattended backup completed successfully.
A real off-site artifact from that pipeline has now been decrypted and restored into PostgreSQL 18 with zero restore errors.
That was the point at which we considered the migration and backup role reversal genuinely complete.
Not when DNS changed.
Not when the dashboard turned green.
When we proved that the data could come back.
Because backups are easy to celebrate.
Restores are what count.

