Skip to content

Command Palette

Search for a command to run...

Operations

Resource limits

Cap per-instance CPU and memory with cgroups v2 on Linux, enforced before the process starts.

Configure limits

Phelix can apply per-instance CPU and memory limits on Linux via cgroups v2. Configure them in the phelix.yaml resources block - both fields are optional:

phelix.yamlyaml
resources:  cpu: "500m"  memory: "512Mi"
FieldAcceptsNotes
cpumillicores (250m, 500m, 1000m) or cores (0.5, 1, 2), max 3 decimalsA bandwidth ceiling, not CPU affinity or a reserved core. The cpu.max period is fixed at 100000µs, so 500m writes "50000 100000". The kernel minimum is 10m; smaller values fail validation rather than rounding up.
memorypositive integer with binary units Ki, Mi, Gi, Ti512Mi writes 536870912 to memory.max. Zero, negative, malformed, overflowing, empty, and unsupported-unit quantities are rejected.

Per-instance cgroups

Each running instance gets its own cgroup - including each replica and both the old and new instances during a deployment. Descendants inherit that instance's limits. Phelix configures the limits before atomically spawning the process into the cgroup (CLONE_INTO_CGROUP), so there is no unrestricted launch window. Setup or attachment failure aborts the launch; it never retries without limits.

Persistence and rollback

Build and rebuild save this policy as the app's current runtime configuration in apps.json, separate from versioned artifacts. Subsequent starts, restarts, and rollbacks use that current policy.

Host requirements

  • Linux cgroups v2, kernel 5.7+ with clone3 permitted.
  • A writable delegated parent cgroup with the requested cpu/memory controllers already enabled in cgroup.subtree_control. Set PHELIX_CGROUP_ROOT to its clean absolute path, or Phelix resolves its current cgroup via /proc/self/mountinfo and /proc/self/cgroup.
  • Phelix does not move the launcher or alter ancestor controller settings. Containers and systemd services must supply suitable delegation and syscall permissions. Non-Linux hosts reject configured limits; apps without limits keep the existing launch path.

Memory-limit violations (RESOURCE_OOM)

Memory-limit violations are detected through the instance cgroup's memory.events counters: an instance is classified as resource OOM when oom_kill increases during its lifetime - not from the exit status or SIGKILL alone. The classification unit is the cgroup, so a killed descendant counts even if the main process survives.

Resource OOM surfaces as the structured RESOURCE_OOM error - distinct from a generic process crash, a non-zero exit, and a failed health check - carrying the configured limit and PID. The launcher reads the final counters before cleanup releases its lease, so cleanup never erases the evidence. Exits for any other reason (including unreadable evidence) keep the existing behavior.

Deployment safety

  • A blue-green candidate that hits its memory limit fails its health check and never becomes the serving instance - the deployment fails with the resource-OOM reason while the old instance keeps serving.
  • A rolling replacement killed by its memory limit is rejected through the existing rolling failure/recovery path (the previous replica is restored and still serving), with the resource-OOM reason attached.
  • A detached cleanup helper waits for the instance cgroup to empty, then removes only that owned leaf. It covers normal exit, crash, and failed startup, is idempotent, and survives the CLI exiting.

Current limitations

Current memory usage, memory event counters, and CPU accounting are readable internally for lifecycle classification and future monitoring, but there is no dashboard, backend API, or live resizing in this phase - resource changes still apply only to newly launched instances. This applies to native app processes, not Docker/Kubernetes workloads or build commands.