Deploys, health & restarts
Everything on this page works without a daemon or orchestrator: state is files, coordination is a signal, enforcement is the kernel.
#Health gates
[run.health]
port = 5432 # TCP connect check against the instance's IP
grace = "30s" # budget for cold start (mounts + app init)
An instance is healthy once its declared port accepts a TCP connection.
grace is how long a fresh instance may take before the gate fails —
budget for cold-mount I/O and app initialization, not just process start.
#Rolling deploys
ply build . # new version → new image
ply deploy myapp-1.3.0-linux-x64.img
What happens:
- ply writes a deploy pointer file and sends the app's run parent
SIGHUP - The parent preps the new version while the old instances keep serving
- Instances roll one at a time: the old instance leaves the published pool (no new connection reaches it from this moment, in the kernel's DNAT or the relay), gets one second for what is in flight, then its stop signal; the new one starts and joins the pool only once it passes the health gate — before that, nothing is routed to it
- A failed gate aborts the roll and reverts that slot — untouched instances never left the old version
Zero downtime with --scale ≥ 2, and the failure mode is "some instances
still on the old version," never "everything down." --timeout <s> bounds
how long to watch before reporting partial progress.
The same rules carry the autoscaler: when a [scale]
policy adds or removes instances it uses this launch and stop path.
The same readiness rule holds for every launch, not only rolls: an instance takes traffic once its first published port accepts a connection, so a scale-up or a crash-restart never routes a request to a process that is still starting.
What the app owes. ply can stop feeding an instance new connections; it
cannot end the keep-alive connections clients already hold on it. On the
stop signal, a well-behaved server marks every further response
Connection: close — the client hangs up after reading it and reconnects,
to the other instances — waits until its open connections have drained,
and only then shuts down. What it must not do is close idle connections
itself: under load "idle" is the instant a client's next request is already
on the wire, and that request is lost. Go's SetKeepAlivesEnabled(false)
and Shutdown both close idle connections, so call Shutdown only after
the drain; bench/api/main.go has the pattern. Measured on the bench: a
rolling restart at scale 2 under 655k requests/s, 39 million requests,
zero lost — with the idle-close variant, one request per connection was.
There's no magic: you can watch the pointer file and per-instance state
files change under /run/ply/ while it happens.
#Deploying a sleeping app
An app asleep under [scale] min = 0 (see Autoscaling)
has no instance to roll. ply deploy still writes the pointer and signals
the parent; the parent swaps the image for the next wake and the deploy
reports complete at once:
ply: deploy complete — asleep, the next wake runs /var/lib/ply/apps/web/current.img
#When something restarts
ply why APP answers before you start guessing:
web — 2 instances, image /srv/web-1.2.0-linux-x64.img, published 10.77.0.1:8080
web.1 10.77.0.3 pid 4242 up 12m restarts 2
web.2 10.77.0.4 pid 4301 up 1h 4m restarts 0
last deploy: web 1.1.0 -> 1.2.0
restarts: 2 (newest first)
2026-09-05T17:16:40Z web.1 killed by signal 9 (SIGKILL), OOM-killed (oom_kill=1), up 1m 30s; policy on-failure -> restart in 2s
log: fatal: out of memory allocating 512 MiB
2026-09-05T17:15:10Z web.1 exit code 1, up 30s; policy on-failure -> restart in 1s
log: Error: connect ECONNREFUSED 10.77.0.1:5432
changes: 3 (newest first)
2026-09-05T17:14:00Z deploy web 1.1.0 -> 1.2.0
…
Exits, with their code or signal, OOM kills and uptime, come from the
journal; the log lines from the slot's ring; blocked traffic from the
egress log; changes from the journal. --json gives an agent the same.
#Rollback
Images are content-addressed files and registries are append-only — the old version still exists. Rolling back is deploying:
ply deploy myapp-1.2.9-linux-x64.img
#Restart policies
[run.restart]
policy = "on-failure" # or "always" / "never" (default)
backoff = "1s" # first respawn delay, doubles each failure
max_backoff = "60s"
The run parent respawns instances it started — same slot, same volume —
with exponential backoff that resets after healthy uptime. Ctrl-C and
ply rm are shutdown-aware (no respawn on intentional stops). ply ps
shows a RESTARTS column.
The parent stays boring by design: no socket, no API. It supervises only what it forked.
#Supervision across reboots
The restart policy covers crashes; host reboots are systemd's job:
ply systemd myapp.img --scale 4 --publish 80:3000 \
| sudo tee /etc/systemd/system/ply-myapp.service
sudo systemctl enable --now ply-myapp
The unit does not name the image file: ply systemd makes a current.img
link beside it and the unit runs that. Every later ply deploy re-points
the link after its roll succeeds, so a unit restart or a reboot comes back
on the version you deployed — not on the file the unit was written with.
ply deploy says when it re-pointed a link, and notes when an app was
started from a plain path that a restart would revert to.
#Continuous deployment
The page above is the push side: ply deploy rolls whatever you hand it,
so CI can scp an image and run one command over SSH. The published
actions (iluxav/ply@v1 to build, iluxav/ply/deploy@v1 to ship and
roll) do it in two workflow steps — the complete recipe is
Running on DigitalOcean.
The pull side is usually the better default: declare a deployment file naming a GitHub repo or release, and the host converges itself — on push, within a minute, with no server credentials in CI at all. That whole story, including building straight from a repo on a $4 droplet, is Deployments & CD.