How a company leaves the cloud, one move at a time, by the makers of OneUptime, published three ways: a printable book, a reflowable EPUB, and a static website you can host anywhere.
Most writing about leaving the cloud is either an opinion piece or a programme plan for a company with a platform team. This handbook groups the work into stages, sized for a company whose platform is two or three engineers' work and whose cloud bill is somewhere around $100,000 a month.
π dist/Back-to-Metal.pdf Β Β·Β
π dist/Back-to-Metal.epub Β Β·Β
π site/ β open site/index.html
Decide, Buy, Build, Move and Run. The generated tables below show the current Move counts, stage boundaries and prerequisites.
Each Move is one job, with a stated cutover in minutes of user-visible downtime, a stated risk, and a rollback that names its point of no return. Rehearse the return path; some commitments and deletions cannot be undone.
The first edition of this book had 122 Moves. It was correct, and it was 975 person-days and four and a half years of elapsed time for three engineers, which is not a plan a company can act on. This edition is the same argument at a size somebody can finish: the whole thing is estimated from the Move files. Procurement and observation windows extend the calendar without adding engineer-days; they are accounted for separately from labour.
The exact figures are in the table at the foot of this file, and they are computed from the Move files rather than typed.
Move 04 concludes that you should keep paying somebody else for three things: a content delivery network, outbound email deliverability, and denial-of-service scrubbing at the edge. Each is a business other people run better than you will, and each is cheap next to what it replaces. Move 03 gives you permission to read three Moves, do the arithmetic and stop, which is a cheaper outcome than a programme abandoned in month five with two platforms running.
The comparison keeps the existing operations team, hours and salary in all three options, so the reference model adds no staffing cost. Cloud ops engineers transition to on-premises platform work. Colocation remote hands handles contracted cabling, drive swaps and other physical tasks, already priced in facility charges; rented-metal providers maintain their hardware under the support agreement.
On-premises operations roles can cost less than cloud specialist roles, offering further salary savings. The published figures keep the existing salary cost and count none of those potential savings. Price any recurring difference using your staffing plan, local pay rates, measured hours and support quotes. Migration effort is budgeted separately. Rent to preserve cash, or colocate owned machines for lower running costs and control over the fleet.
Every Move covers AWS, Google Cloud and Azure. The job is the same whichever you are leaving; only the extraction differs, so only the extraction is written three times. Each Move opens with three lines naming the real product on each provider and the one thing that is genuinely different there β a flag that needs a reboot, a tier that cannot do it at all, a resource that outlives its parent. The runbook itself is written once.
Existing cloud VMs move to KVM guests on Proxmox VE, with several guests sharing each physical host. Already-containerised services follow the Talos/Kubernetes path, directly on hardware or in guests when sharing a VM estate. Move 02 identifies which path each workload needs; Moves 11β14 cover provisioning, storage and the first migration, and Move 19 covers whole-VM restoration alongside Kubernetes and database recovery. Moving out of the cloud does not require converting a working VM application to containers. The published cost model is the container reference build; price your VM mix in Move 03.
Every Move's prerequisites are lower-numbered Moves. A reader who has reached Move 14 has, by construction, met every prerequisite of Move 14. That is not a stylistic choice β the build refuses to compile a book where it does not hold, and it is what lets the website's checklist read top to bottom rather than search a graph.
moves/ the Move files, numbered from 01 β the source of truth
book/ the typesetter and the site generator
parse.py markdown -> structured Move data
verify.py structure: the contract, the arithmetic, the house style
audit.py content: undeclared tools, unguarded destructive steps,
unpinned installs, drifted service names, repeated prose
build.py Move data -> a self-contained book.html
render.py html -> PDF via headless Chromium, vertically justified
site.py Move data -> a static website in site/
epub.py Move data -> a reflowable EPUB3
cover.py the paperback wrap, the Kindle cover, the hardback case
deps.py which Moves must be finished before which
kit.py the reference build, the laptop kit, the seven rules
costs.py the cost model, with the salary line in it
equivalents.py AWS to Google Cloud to Azure to what you run instead
rollback_data.py the safety page every Move's rollback derives from
symptoms.py the symptom index β "which Move do I need"
icons.py stage glyphs, the cutover gauge, the risk bars
style.css the print stylesheet
web/ the site's stylesheet and script
dist/ the built PDF, EPUB and covers
site/ the built website
build/ intermediate HTML (gitignored)
make deps # fonts, playwright, chromium
make verify # structure β must be clean
make audit # content β must be clean
make book # markdown -> HTML -> PDF -> KDP compliance check
make site # markdown -> static website
make # all of it, and the tables in this READMEPython 3.9 or newer with the packages in requirements.txt and Chromium; make deps installs
the packages and browser. The Makefile prefers
.venv/bin/python3 when it exists, so python3 -m venv .venv && make deps works without
activating anything.
verify.py and audit.py are the gate and both must come back clean. verify.py checks the
shape of every Move: the sections, the three-cloud block, the cost arithmetic, and that no Move
moving persistent state ships without saying how the state comes back. audit.py goes after the
things that are actually wrong in infrastructure writing β a runbook calling a tool the
prerequisites never mentioned, a destructive command with nothing standing behind it, software
installed without a version, a service spelled four ways, prose copy-pasted between Moves.
site/ is self-contained β the fonts are embedded in the stylesheet and nothing is fetched at
runtime β so it drops onto any static host unchanged. It is five pages and one page per Move:
the guide is the whole plan on one page, what it costs is the arithmetic, your
checklist is where you are and what is next, before you start is the safety page, and
about is the book. Every page carries one link to the next Move you have not ticked.
The most valuable contribution is a correction. If you ran a runbook and it did not work as written, that is worth an issue on its own. CONTRIBUTING.md covers the workflow; AGENTS.md documents the Move contract, the house style the build enforces, and what not to hand-edit.
Two licences, because this repository is two things:
- The software β the toolchain in
book/, the print stylesheet, the site's stylesheet and script β is MIT. - The content β the Moves, the written text of the book β is CC BY 4.0.
So you may share and adapt any of it, for any purpose including commercially, as long as you give credit.
| Moves | 20 |
| Stage 1 Β· Decide | 4, 01β04 |
| Stage 2 Β· Buy | 4, 05β08 |
| Stage 3 Β· Build | 4, 09β12 |
| Stage 4 Β· Move | 4, 13β16 |
| Stage 5 Β· Run | 4, 17β20 |
| Work | 107 person-days |
| End to end, two engineers | 41 weeks, including procurement and observation waits |
| At zero downtime | 18 of 20 |
| Whole book, end to end | 25 minutes of user-visible outage |
| Cannot be undone | 1 |
| Risk | 3 low, 7 medium, 10 high |
| Dependencies | 30, every one pointing backwards |
| Line savings across the Moves | $86,320 a month, after the facility line, before other platform costs |
| Saving, with everything counted | $80,885 a month β $970,620 a year, 71 per cent of a $114,644 bill with the salary on both sides |
Every Move names the real service on AWS, Google Cloud and Azure, and the one thing that differs on each.
| # | Move | Leaving | Risk | Cutover | Back out for |
|---|---|---|---|---|---|
| 01 | The bill, and the three lines that are most of it | β | Low | 0 min | Immediately |
| 02 | What you actually run | β | Low | 0 min | Immediately |
| 03 | The number that decides it | β | Medium | 0 min | Immediately |
| 04 | The three things you keep renting | Nothing β this Move decides what stays rented | Low | 0 min | Until a contract is signed |
| # | Move | Leaving | Risk | Cutover | Back out for |
|---|---|---|---|---|---|
| 05 | From rented vCPUs to cores you own | The vCPU as a unit of purchase | Medium | 0 min | Until the order is signed |
| 06 | Sixteen machines, and the two on the shelf | Elastic node capacity and cluster autoscaling | High | 0 min | Until the order is signed |
| 07 | A cage, not a data centre | The region-and-zone abstraction | High | 0 min | Until the contract is signed |
| 08 | The order, and the weeks you cannot compress | Provider-assigned addresses and managed transit | Medium | 0 min | Until the order is signed |
| # | Move | Leaving | Risk | Cutover | Back out for |
|---|---|---|---|---|---|
| 09 | Racking day | The provider's serial console and boot diagnostics | Medium | 0 min | Immediately |
| 10 | The network, and the way back in when it breaks | Cloud-managed private networking | High | 0 min | Immediately |
| 11 | The platform your workloads need | Managed VM and Kubernetes compute | High | 0 min | Immediately |
| 12 | Disks: what goes local, what goes on Ceph | Managed block and shared-file storage | High | 0 min | Immediately |
| # | Move | Leaving | Risk | Cutover | Back out for |
|---|---|---|---|---|---|
| 13 | Images, secrets and one-command deploys | Managed registries, secret stores and hosted CI runners | Medium | 0 min | Immediately |
| 14 | The first service, end to end | Cloud VM and container compute | Medium | 0 min | Immediately |
| 15 | Buckets, cache and queues | Managed object storage, Redis and queues | High | 0 min | 30 days |
| 16 | Postgres, the one that matters | Managed PostgreSQL | High | 15 min | 7 days |
| # | Move | Leaving | Risk | Cutover | Back out for |
|---|---|---|---|---|---|
| 17 | The front door | Managed layer-7 balancers and certificate services | Medium | 0 min | Immediately |
| 18 | Go-live, and how you abort | Weighted DNS and the parallel managed edge | High | 10 min | 7 days |
| 19 | Backups you have restored, and the pager | Managed backup, alarms and a rented paging service | High | 0 min | Immediately |
| 20 | Closing the account | The provider organisation, its audit trail and its support plan | High | 0 min | No |