Skip to content

Host preparation and recovery

New and existing appliances use one post-install workflow. The installer provides Ubuntu, K3s, the dashboard, diagnostic helpers and a root-owned local worker. It does not contain a second GPU package-selection or reboot path. The base kernel must already boot the computer and reach storage/network; post-install preparation cannot repair an installer that cannot boot. For an installer-created NVIDIA display host, the base host playbook separately installs a pinned display-owning driver and schedules one reboot after successful first convergence, so the physical setup/TUI console remains available. The same host role keeps nvidia-persistenced active once the driver is usable. The GPU Operator's CDI device specifications refer to its Unix socket, so an inactive service can prevent newly created NVIDIA model containers from starting even when the operator and nvidia-smi appear healthy. Check systemctl is-active nvidia-persistenced and test -S /run/nvidia-persistenced/socket before changing the GPU Operator. The media does not select the driver or schedule that reboot. CPU-only and AMD-only hosts skip it; optional AMD kernel/profile changes remain administrator-confirmed. New USB media use Ubuntu 26.04 LTS and its native generic kernel for both the installer and installed system. This base-install choice is separate from the reviewed GPU package profiles below; it does not upgrade an existing appliance or change its distribution repositories. Ubuntu 24.04 host preparation remains available for existing installations. See installer kernel selection.

Fresh K3s installations are pinned through the overridable Ansible default k3s_version: v1.36.4+k3s1, rather than resolving a moving install channel. This release includes the native containerd v3 configuration drop-in imports used by the NVIDIA toolkit. The installation task retains its creates: /usr/local/bin/k3s guard: ordinary host convergence does not upgrade an existing cluster. Existing clusters require a separately planned K3s upgrade and confirmation that the generated containerd configuration imports config-v3.toml.d/*.toml before adopting the new NVIDIA configuration. An override must preserve that prerequisite; choosing another K3s version is not a claim of GPU support. See the pinned K3s release.

User workflow

System → Model cache uses this same host-operation boundary for explicit cleanup of downloaded models. It never removes credentials or container images, and refuses cleanup while local model deployments can use the cache. See model cache management for scope and upgrade requirements.

Ethernet and Wi-Fi use the same bounded host-operation mechanism through System → Settings → Network, adding independent local timeout and boot recovery. See network management. Network, power and GPU changes never execute concurrently on the same host.

Open System → Hardware → GPU nodes → Host preparation. GPU operators are shown first. Each node contains its kernel plan, preparation acknowledgement, advanced profile override and collapsed GPU memory configuration. Help and diagnostic details are available through info icons, while action warnings and exact-host confirmations remain explicit. The local worker periodically inspects the OS, running kernel and all PCI display GPUs. It publishes a versioned, host-bound plan, including exact package changes and whether a reboot is necessary. Inspection itself never installs packages or restarts the host.

For a matching reviewed plan, acknowledge the experimental hardware profile, select Review hardware preparation, and type the computer's exact name in the confirmation dialog. The confirmation authorizes the displayed packages, one orderly restart if required, and AMD GPU profile activation. The local root worker runs the existing Ansible preparation role, persists its progress, and resumes verification after reboot. Closing the dashboard does not cancel an accepted operation. Existing working kernels are not removed. After host verification, Registering waits for fresh eligible hardware and an allocatable Kubernetes GPU. This completes preparation without an engine smoke test. Engine validation is an optional manual action in System → Hardware; no test-model or inference-image download is started by host preparation. The collapsed Advanced · AMD runtime profile section is a manual override, not another required setup step: preparation already activates the matching runtime profile. The override changes no host packages or kernel.

Package and power operations target only the selected computer. GPU runtime profile selection uses the existing AMD ModuleActivation, which applies to matching cluster Nodes; the confirmation explicitly includes that scope. Concurrent preparation requests must not overwrite a changed module decision: the worker detects configuration changes and stops its profile handoff. Runtime profile selection is still cluster-wide. The node's Verify Ollama and Verify vLLM buttons request only that engine on that node; per-model runtime placement remains a separate extension.

The shipped x86-64 Strix Halo (1002:1586) host package profiles are:

Host OS Profile / version Exact package Target kernel
Ubuntu 26.04 strix-halo-ubuntu-26.04 / 1 linux-generic=7.0.0-31.31 7.0.0-31-generic
Ubuntu 24.04 strix-halo-ubuntu-24.04 / 1 linux-generic-hwe-24.04=7.0.0-31.31~24.04.1 7.0.0-31-generic

The 26.04 package pin comes from the official resolute-updates package listing. It is a bounded additional preparation plan, not an OS upgrade or a hardware certification. The new Ubuntu 26.04 installation path still requires local hardware and inference acceptance; both profiles remain experimental. An already working native 7.0 kernel needs no package change or restart: explicit preparation only activates the matching GPU profile. A suitable kernel with a failing driver instead remains blocked for diagnosis. The worker does not reinstall the pinned package or downgrade a newer working kernel. NVIDIA, Intel and other AMD devices without this preparation requirement retain their existing operator path; no matching profile is not proof of GPU support. Profiles must be maintained and revalidated for future security updates; the worker never selects arbitrary latest releases. The Ansible package gate allows the linux-generic meta-package only on 26.04, and the 24.04 HWE meta-package only on 24.04. It never adds third-party package repositories.

Mixed systems and experiment mode

All GPUs inside one computer share its kernel. Different cluster Nodes have independent kernels. The normal path blocks a Strix Halo kernel change on an unreviewed mixed/multi-GPU computer, without disabling working vendor operators. A working Strix Halo host driver alongside another vendor can still undergo AMD engine validation without changing its kernel.

When a shipped host profile matches the OS/architecture and required Strix Halo device but the additional GPUs are unreviewed, administrators can enable Experiment mode. The UI shows the full detected GPU list and the proposed profile. A separate acknowledgement and exact-host confirmation are required. Other GPU drivers, graphics output or connectivity may break; have physical console access and the previous kernel available in the bootloader. There is no claim of automatic bootloader rollback or remote power recovery.

Experiment mode permits only the displayed, shipped package plan. It cannot accept an arbitrary kernel URL, shell command, package name, repository, ROCm installer or architecture override. Unknown OS/architecture and incomplete PCI inventory remain blocked; a new package profile needs code review first. It does not bypass inference eligibility. Multi-AMD layouts not supported by the current node-scoped plugin profile may finish as PreparedUnverified: host preparation is distinct from GPU availability. Successful AMD smoke tests do not certify other vendors or the entire mixed system. Tests remain local evidence, never automatic additions to the public compatibility catalog.

Fixed and dynamic GPU memory

System → Hardware → GPU nodes → GPUs → GPU Configuration AMD → Shared GPU memory provides two administrator controls on one supported Strix Halo GPU with compatible kernel/firmware evidence, including hosts with additional NVIDIA GPUs bound to the nvidia driver:

  • Fixed GPU reservation (firmware): a discrete slider containing only the options reported by that computer's uma/carveout_options. This RAM is carved out at boot and is unavailable to Linux.
  • Dynamic GPU memory limit: a slider in 1 GiB steps for the TTM ceiling on GPU use of shared Linux RAM. It is not a reservation, an additional memory pool, or guaranteed free memory. CPU workloads compete for that RAM.

The panel separates current values from the unsubmitted draft. Changing a slider sends no host operation. A non-step-aligned kernel default can have a rounded draft, but that alone cannot enable submission. If the active limit exceeds the allowed maximum for the current firmware reservation, the panel instead offers a limited corrective draft for review without another slider movement. It still requires exact-host confirmation and never applies itself. The preview estimates Linux RAM as current MemTotal plus the old fixed reservation minus the new fixed reservation. The dynamic ceiling must leave at least 16 GiB outside GPU dynamic allocations; this allowance is not a Kubernetes memory reservation. Actual post-boot MemTotal, not this projection, governs the second stage.

The worker selects the Strix Halo device by its verified PCI address; NVIDIA VRAM is neither reconfigured nor added to the shared-memory budget. The TTM limit is system-wide, so additional AMD GPUs, other GPU vendors, missing driver bindings and NVIDIA GPUs using nouveau or vfio-pci remain blocked. Complete PCI inventory and driver bindings are rechecked at approval and after each restart. After an approved memory reboot, an unchanged but not-yet-bound NVIDIA GPU gets up to ten minutes for its operator-managed driver to start; no further memory write or restart runs during that wait. The PCI hardware identity must remain unchanged; full driver identity is compared once the driver has bound. A changed companion GPU, wrong driver binding or timeout stops the operation. This memory workflow never changes the kernel, NVIDIA driver or operator configuration; ordinary preparation/experiment-mode safety rules remain separate.

Select Review memory configuration and enter the exact computer name in the final disruption confirmation to apply. A fixed-reservation change may require up to two restarts: first apply and verify the firmware choice, then apply the dynamic limit through the existing Ansible role and verify it after another boot. A dynamic-only change requires one restart. All workloads on that computer are interrupted; no migration or automatic firmware rollback is promised. Keep physical console/recovery access available. Closing the browser does not cancel an accepted operation.

The worker independently checks the reviewed hardware/configuration identity, actual firmware options, active memory values and safety allowance. Missing firmware controls, mixed/unknown GPU layouts, incomplete/stale evidence and competing boot/modprobe overrides disable this flow. Hardware experiment mode does not override those memory guards. A memory operation installs no kernel or packages, changes no inference profile and starts no GPU validation workload. Its success confirms the requested memory settings, not inference compatibility, performance, cgroup accounting or successful model execution.

The hardware fingerprint excludes only the configurable UMA VRAM capacity: changing that capacity is the intended operation, not evidence of a replaced GPU. PCI identities, architecture, driver, kernel, OS and firmware versions remain checked. The desired firmware reservation and dynamic limit are verified separately against their exact approved values after reboot. If a previous operation stopped after applying the firmware choice, its old dynamic limit may still be active. Diagnose the failure, refresh the current evidence and confirm a new plan; terminal failed requests are never replayed automatically.

The next-boot dynamic setting is owned in /etc/modprobe.d/90-magicstick-ttm.conf; the role refreshes initramfs. Neither a live TTM-only write nor the deprecated amdgpu.gttsize override is used. Firmware option indices are resolved locally; browser requests contain no device path or shell command. Linux UMA controls and AMD shared-memory guidance.

Restart and shut down

Administrators see a separate Computer power tab inside System, immediately after System Status (#/system/power), with Restart computer and Shut down computer. These controls are only mounted on this tab. Viewers and operators cannot access it, including through a direct link. If several managed computers are reported, select the target explicitly. Both buttons require typing its exact name and accepting interruption of all services/workloads on that computer. They are unavailable for stale/offline workers or while another operation is active.

The worker uses normal systemd shutdown -r +1 or shutdown -P +1: approximately one minute after local acceptance, not an immediate forced reset. Save work before confirming; there is no automatic Pod migration or zero-downtime promise. The dashboard distinguishes accepted, scheduled, and new boot observed. It cannot prove physical power-off while the machine is unreachable, and does not repeatedly send a power request. Power-on requires local action or separately configured remote power management. systemd shutdown manual

Execution and security boundary

  • magicstick-host-management.timer runs after boot and every 15 seconds after the previous tick completes. It waits for cloud-init's base installation to finish. No installation-time consent is inferred.
  • Root-owned implementation and profiles live in /usr/local/lib/magicstick/host-management/. Local progress is atomically persisted in /var/lib/magicstick/host-management/state.json (root only).
  • The worker publishes appliance.magicstick.dev/host-management on its local Node. API availability requires matching Node UID, boot ID and kernel plus evidence no more than three minutes old.
  • GET /api/host-management is viewer-readable. Administrator-only POST /api/host-management/operations requires existing authentication, same-origin/CSRF checks, exact-host confirmation, current Node/boot/plan IDs, explicit disruption consent, and a unique 32-hex request ID. No secrets or executable content are accepted. The actor subject is hashed for the request.
  • One immutable namespaced HostOperation per Node serializes browser requests. New requests expire after five minutes if not accepted locally. Active requests cannot be replaced; terminal replacement uses a resource UID precondition. The worker independently rechecks all inputs, and remembers processed IDs.
  • The dashboard may create/read/delete these requests in ai-system; it cannot write their status, patch Nodes or execute host commands. No privileged dashboard Pod, hostPath or network-facing root agent is introduced. The worker uses the existing root-only local K3s kubeconfig; cluster/root administrators remain trusted. Arbitrary external Kubernetes installs and agent-only Nodes without a local management credential do not silently gain host control.
  • Host convergence and host operations share /var/lib/magicstick/host-management/maintenance.lock inside a root-only directory. Convergence skips an active operation and a scheduled shutdown; it does not compete with APT preparation.
  • APT uses exact approved versions, signed configured repositories, no downgrades and no package removals. The prepare playbook is the same bounded Ansible implementation used for explicit local maintenance. Ansible APT contract

Progress and recovery

Typical preparation: Preparing → RebootScheduled → Verifying → Registering → Succeeded. A host that needs no restart enters registration directly. Succeeded means fresh host/driver evidence and an eligible registered GPU, not an engine smoke test, production model acceptance or KubeAI runtime-image adoption. Legacy in-progress Validating operations now complete from the same registration checks. Optional engine diagnostics remain visible under GPU compatibility.

Memory configuration uses Preparing → RebootScheduled → Verifying, with a second bounded cycle when both firmware reservation and TTM need changing. Root-owned state records the stage and reboot count before side effects. An interrupted/uncertain write is never automatically replayed. Changed hardware, unexpected actual memory or a conflicting local setting stops the operation; inspect the journal before issuing a new confirmed request.

Rejected, Interrupted, Failed and PreparedUnverified are terminal. An uncertain package execution, changed boot/profile, wrong kernel after reboot, expired request or failed engine test does not trigger an automatic retry, another kernel change, or a reboot loop. Inspect and submit a new confirmed request only when appropriate. GPU validation waits at most one hour, allowing for large image downloads; host driver verification waits five minutes.

sudo systemctl status magicstick-host-management.timer
sudo journalctl -u magicstick-host-management --since '-30 minutes'
sudo magicstick-gpu-preflight --json
kubectl -n ai-system get hostoperations

Do not delete an active request or its root-owned state to force a retry. If an administrator manually removes Kubernetes execution state, the local worker stops rather than guessing whether a power/package operation already happened. Retain the journal for audit and review local recovery before clearing state.

Local verification

PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover \
  -s magic-host/roles/host-management/tests
PYTHONDONTWRITEBYTECODE=1 python3 \
  magic-host/roles/host-management/tests/rancher_contract_test.py \
  --context rancher-desktop

The second test explicitly targets local Rancher, creates an isolated namespace and the CRD only if absent, and removes its own resources afterwards. It verifies real Kubernetes validation, immutability, RBAC and status persistence with a fake power executor. It neither installs a host worker nor patches/restarts any Node. Unit tests cover package/reboot decisions and post-boot continuation; physical power loss, bootloader fallback and actual mixed-GPU compatibility still require a separately approved hardware test.