docs /Deploys, health & restarts

Deploys, health & restarts

Everything on this page works without a daemon or orchestrator: state is files, coordination is a signal, enforcement is the kernel.

#Health gates

[run.health]
port = 5432        # TCP connect check against the instance's IP
grace = "30s"      # budget for cold start (mounts + app init)

An instance is healthy once its declared port accepts a TCP connection. grace is how long a fresh instance may take before the gate fails — budget for cold-mount I/O and app initialization, not just process start.

#Rolling deploys

ply build .                                    # new version → new image
ply deploy myapp-1.3.0-linux-x64.img

What happens:

  1. ply writes a deploy pointer file and sends the app's run parent SIGHUP
  2. The parent preps the new version while the old instances keep serving
  3. Instances roll one at a time: the old instance leaves the published pool (no new connection reaches it from this moment, in the kernel's DNAT or the relay), gets one second for what is in flight, then its stop signal; the new one starts and joins the pool only once it passes the health gate — before that, nothing is routed to it
  4. A failed gate aborts the roll and reverts that slot — untouched instances never left the old version

Zero downtime with --scale ≥ 2, and the failure mode is "some instances still on the old version," never "everything down." --timeout <s> bounds how long to watch before reporting partial progress.

The same rules carry the autoscaler: when a [scale] policy adds or removes instances it uses this launch and stop path.

The same readiness rule holds for every launch, not only rolls: an instance takes traffic once its first published port accepts a connection, so a scale-up or a crash-restart never routes a request to a process that is still starting.

What the app owes. ply can stop feeding an instance new connections; it cannot end the keep-alive connections clients already hold on it. On the stop signal, a well-behaved server marks every further response Connection: close — the client hangs up after reading it and reconnects, to the other instances — waits until its open connections have drained, and only then shuts down. What it must not do is close idle connections itself: under load "idle" is the instant a client's next request is already on the wire, and that request is lost. Go's SetKeepAlivesEnabled(false) and Shutdown both close idle connections, so call Shutdown only after the drain; bench/api/main.go has the pattern. Measured on the bench: a rolling restart at scale 2 under 655k requests/s, 39 million requests, zero lost — with the idle-close variant, one request per connection was.

There's no magic: you can watch the pointer file and per-instance state files change under /run/ply/ while it happens.

#Deploying a sleeping app

An app asleep under [scale] min = 0 (see Autoscaling) has no instance to roll. ply deploy still writes the pointer and signals the parent; the parent swaps the image for the next wake and the deploy reports complete at once:

ply: deploy complete — asleep, the next wake runs /var/lib/ply/apps/web/current.img

#When something restarts

ply why APP answers before you start guessing:

web — 2 instances, image /srv/web-1.2.0-linux-x64.img, published 10.77.0.1:8080
  web.1  10.77.0.3  pid 4242  up 12m  restarts 2
  web.2  10.77.0.4  pid 4301  up 1h 4m  restarts 0
  last deploy: web 1.1.0 -> 1.2.0

restarts: 2 (newest first)
  2026-09-05T17:16:40Z  web.1  killed by signal 9 (SIGKILL), OOM-killed (oom_kill=1), up 1m 30s; policy on-failure -> restart in 2s
    log: fatal: out of memory allocating 512 MiB
  2026-09-05T17:15:10Z  web.1  exit code 1, up 30s; policy on-failure -> restart in 1s
    log: Error: connect ECONNREFUSED 10.77.0.1:5432

changes: 3 (newest first)
  2026-09-05T17:14:00Z  deploy           web 1.1.0 -> 1.2.0
  …

Exits, with their code or signal, OOM kills and uptime, come from the journal; the log lines from the slot's ring; blocked traffic from the egress log; changes from the journal. --json gives an agent the same.

#Rollback

Images are content-addressed files and registries are append-only — the old version still exists. Rolling back is deploying:

ply deploy myapp-1.2.9-linux-x64.img

#Restart policies

[run.restart]
policy = "on-failure"     # or "always" / "never" (default)
backoff = "1s"            # first respawn delay, doubles each failure
max_backoff = "60s"

The run parent respawns instances it started — same slot, same volume — with exponential backoff that resets after healthy uptime. Ctrl-C and ply rm are shutdown-aware (no respawn on intentional stops). ply ps shows a RESTARTS column.

The parent stays boring by design: no socket, no API. It supervises only what it forked.

#Supervision across reboots

The restart policy covers crashes; host reboots are systemd's job:

ply systemd myapp.img --scale 4 --publish 80:3000 \
  | sudo tee /etc/systemd/system/ply-myapp.service
sudo systemctl enable --now ply-myapp

The unit does not name the image file: ply systemd makes a current.img link beside it and the unit runs that. Every later ply deploy re-points the link after its roll succeeds, so a unit restart or a reboot comes back on the version you deployed — not on the file the unit was written with. ply deploy says when it re-pointed a link, and notes when an app was started from a plain path that a restart would revert to.

#Continuous deployment

The page above is the push side: ply deploy rolls whatever you hand it, so CI can scp an image and run one command over SSH. The published actions (iluxav/ply@v1 to build, iluxav/ply/deploy@v1 to ship and roll) do it in two workflow steps — the complete recipe is Running on DigitalOcean.

The pull side is usually the better default: declare a deployment file naming a GitHub repo or release, and the host converges itself — on push, within a minute, with no server credentials in CI at all. That whole story, including building straight from a repo on a $4 droplet, is Deployments & CD.