Release candidate: this page describes
410fbfb, which is separate frommain. See version and availability.
Deploy a CPRa instance¶
Begin with a small, observed deployment using the monitor intervals and target types you actually need. CPRa has one active owner and one local Raft voter. Its deployment contracts preserve state across restart; there is no distributed failover.
Assign ownership and permissions¶
Run one process for a monitor configuration. Independent instances do not coordinate recovery or alert ownership. Give the process account only the access needed for its configured checks and actions.
Keep manifests containing credentials outside version control. A manifest authorizes outbound requests and local or remote recovery actions. -ssrf-protect restricts HTTP(S) destinations; it does not sandbox other protocols or local actions.
Protect remote access¶
Create a token file readable by the service account, then start the server:
./bin/cpra -yaml /etc/cpra/monitors.yaml \
-web.addr 0.0.0.0:8060 \
-web.auth-file /run/secrets/cpra-token
Put an HTTPS reverse proxy in front of this HTTP listener. Browser login is username cpra with the token as password.
./bin/cpractl --server https://monitor.example.com \
--token-file /run/secrets/cpra-token get overview
monitor.example.com is a placeholder for your protected endpoint. A token file overrides CPRA_AUTH_TOKEN. Liveness, readiness, and metrics requests also need the token when authentication is enabled.
Observe before enabling recovery¶
Confirm that checks classify your target correctly and notifications reach a test destination. Then add the intended recovery action and verify its effect.
Each incident admits one operation. If a response is lost, inspect the target's actual state before repeating the action. Restart restores committed incident state and holds interrupted started actions as unknown. Mount the persistent data directory and retain complete backups.
Native services, Compose and Helm¶
Use native installation for systemd, launchd and
Windows SCM, standard directories, installation ownership and upgrades.
System services always receive explicit absolute configuration and data paths.
Linux system state is /var/lib/cpra; user services use the XDG state directory.
The production image consumes the exact staged release executables and runs as UID/GID 1001. Compose uses a stable named volume, read-only configuration, bounded logs, a 60-second stop allowance and a loopback-published authenticated API. Required bind files must exist; file-backed Secret ownership follows the host.
The Helm chart uses one StatefulSet owner with retained storage. Prefer RWOP on verified CSI storage; RWO is an explicit compatibility option and does not fence multiple pods on one node. Large manifests use a separate read-only configuration volume. They do not belong in Helm values, ConfigMaps or Secrets. Use versioned immutable token Secrets and intentional rollouts for rotation.
Follow the container and Helm guide for exact commands, headless mode, probes, maintenance, backup and Helm 3/4 compatibility. Optional host controls and Kubernetes recovery privileges remain explicit.
Health and shutdown¶
/api/v1/healthzreports process liveness./api/v1/readyzrequires initialized admission, controller progress and available durable storage. Explicit empty configurations can be ready; projection freshness is separate./metricsexposes runtime and pipeline metrics.- SIGINT or SIGTERM marks readiness unavailable and stops admission; diagnostics remain available while accepted work drains or is cancelled.
The Unix/container application budget starts at 45 seconds within a 60-second supervisor allowance; Windows service handling uses a 15-second cap. Deadline expiry is visible and interrupted external outcomes remain unknown. Use an external supervisor to restart a failed process and an external observer for CPRa itself.
Current limits · Troubleshooting
Persistent volume and complete backup¶
Mount a private volume at /var/lib/cpra for the packaged container; the shipped
runtime configuration selects that directory. For a native process, set
storage.directory to an absolute persistent path. Avoid sharing the data
directory between processes. There is one voter and no automatic failover.
Stop CPRa before taking a file-level backup. Copy identity.json, raft.db,
snapshots/, and the entire retained history/ catalog and segment set together.
Use cpractl local backup to hold the exclusive state lock while copying and verifying all files. Restore all files with cpractl local restore into a new private directory and start with the compatible
binary and monitor configuration. Retain credentials separately. A storage
failure stops new admission and makes readiness unavailable; CPRa never silently
switches to memory. See complete recovery procedures.