GPU compatibility contracts¶
GPU discovery, Kubernetes allocation, and successful model inference are separate checks. A PCI ID matching an experimental profile does not mean that the AMD GPU Operator supports it, that ROCm kernels can run on it, or that both inference engines support the same stack. Never set a vendor support label just to make a Helm health check pass.
Engine validation is optional and manual. After the required profile consent, fresh host/driver checks and Kubernetes GPU registration, both catalogued GPU engines are selectable without running a test model. An unverified, running, failed or stale test does not disable a ready GPU. Actual runtime configuration, hardware eligibility and model-specific errors are still checked; availability does not claim successful inference or memory-accounting validation.
The first additional profile is strix-halo, matching AMD display devices with
PCI vendor 1002 and device 1586, with expected compute architecture gfx1151.
The expected architecture is compared with observed KFD/ROCm evidence; it is not
substituted for missing discovery data. Other AMD PCI IDs remain unknown unless
a separate, reviewed profile describes them. A successful HIP test alone is not
an Ollama or vLLM inference acceptance test.
Host evidence without changing the installation¶
Normal host convergence installs the diagnostic helpers:
sudo magicstick-gpu-preflight --json
Alternatively, run the source over SSH without copying or installing anything on the host. Verify its SSH host key first:
ssh ai@example.local 'python3 - --json' \
< magic-host/roles/gpu-compatibility/files/magicstick-gpu-preflight.py
The host helper only reads files and runs bounded rocminfo and dpkg-query
commands when present. It does not download packages, load modules, write
configuration, modify GPU permissions, label nodes, or start workloads. Missing
files, absent commands, execution failures and timeouts are reported as unknown
or unavailable, not as success. It needs no Python packages beyond the standard
library. A host-level rocminfo command is optional when runtime probing is done
in the actual container image.
The JSON includes OS/kernel versions, loaded amdgpu, firmware package evidence,
PCI identities, KFD/render character devices, per-device KFD architectures,
VRAM/GTT counters, TTM limits, and OS-visible RAM. Its stable hardware fingerprint
changes with kernel, driver, firmware package or device identity, not with free
memory. nodeAnnotation is a sanitized controller input proposal for a single
matching Strix Halo device; the diagnostic script does not publish it itself. It leaves
memoryAccountingVerified false because file inspection cannot prove how GPU
allocations are charged to cgroups.
A separate root-owned magicstick-gpu-preflight.timer publishes that evidence
from 20 seconds after boot and every minute once local K3s is available. Failed
publication retries after 10 seconds without weakening identity checks. Its publisher
checks the Node's kernel and boot identity, adds its actual UID, and patches only
appliance.magicstick.dev/gpu-host-preflight. It does not modify eligibility
labels or disclose kubeconfig contents. A missing profile removes only a stale
annotation owned by this publisher. Evidence has a UTC timestamp for freshness
checks. gpu_compatibility_node_name defaults to the local hostname and can be
overridden for a custom K3s node name; gpu_compatibility_publish_evidence=false
disables the publication timers without disabling the read-only diagnostic command.
A separate magicstick-memory-sample.timer runs every 30 seconds and uses
magicstick-gpu-preflight --memory-only through the same publisher. This cheap
path reads only /proc/meminfo and PCI-matched AMD VRAM/GTT sysfs counters; it
does not run ROCm, package queries, engine validation, or inference. It patches
only appliance.magicstick.dev/memory-sample, not compatibility evidence or
eligibility labels. The API checks Node UID, kernel, boot and a 90-second TTL.
CPU availability uses MemAvailable; shared GPU free additionally respects the
remaining GTT ceiling, without double-subtracting GPU allocations from RAM.
Stale or missing samples stay unknown on unified-memory hosts instead of falling
back to potentially inflated Kubelet working-set availability.
For tests, --root /absolute/snapshot/path uses a filesystem fixture and disables
all subprocesses. --require-profile strix-halo exits with status 2 if that exact
PCI profile is absent; this switch is not a compute test.
Kernel and shared-memory constraints¶
AMD documents required Strix Halo kernel fixes in upstream Linux 6.18.4+ and a specific Ubuntu 24.04 OEM backport at 6.14.0-1018+. ROCm userspace compatibility is version-specific even when those fixes are present. The helper reports this kernel evidence but does not certify an arbitrary ROCm/container combination. Older or unrecognized backports require further verification. AMD Strix Halo guidance
New installations use Ubuntu 26.04's native generic 7.0 kernel rather than
installing the Ubuntu 24.04 HWE stack. This satisfies the helper's kernel-version
check, but does not waive AMD driver evidence, experimental profile consent or
the mixed-GPU safeguards. A working host needs no additional package change;
optional engine validation remains separate. Existing 24.04 hosts retain their
own bounded preparation profile. See host package profiles.
Ubuntu 26.04 splits firmware into vendor packages: GPU evidence therefore tracks
linux-firmware-amd-graphics rather than only the linux-firmware metapackage,
so an AMD firmware update invalidates older hardware fingerprints. Ubuntu 24.04
keeps its existing monolithic-package evidence.
Ubuntu firmware package layout
Strix Halo has one GPU and unified physical RAM, not two GPU deployment targets.
GTT/TTM values are dynamic mapping ceilings inside Linux RAM, not an exclusive
reservation. The helper reports their conservative intersection as
gpuAccessibleMi; this is separate from the driver's ordinary device-allocation
capacity. Keep operating-system/Kubernetes headroom and validate actual
container accounting before promising memory protection.
AMD memory explanation
In the current JSON contract, physicalMemoryMi specifically means
OS-visible RAM from /proc/meminfo MemTotal, not total installed DIMM
capacity. It excludes the firmware GPU carve-out. gpuAccessibleMi remains a
mapping bound within Linux RAM, not total GPU memory. The PCI/render-node-matched
KFD local heap must agree with the firmware/VRAM and GTT/TTM counters before the
helper publishes gpuCapacityMi, gpuCapacitySource: kfd-topology and
gpuAllocationMode:
firmware-reserved: the KFD heap agrees with the active firmware carve-out and GTT is no larger. GPU budgets use this pool, outside LinuxMemTotal.shared-gtt: GTT is larger than VRAM and KFD agrees with TTM. Capacity is bounded by GTT, TTM and Linux RAM; model planning also intersects remaining host RAM budgets after safety headroom. The resulting Linux-capped capacity can be smaller than the firmware carve-out. Consumers preserve the host's corroborated domain instead of comparing that capped number with fixed VRAM and incorrectly treating it as contradictory evidence.unknown: missing, stale, ambiguous or contradictory evidence does not become an inferred capacity. GPU capacity is unknown and host requests stay conservative.
For example, a corroborated 64 GiB firmware heap with a 46 GiB dynamic ceiling reports one 64 GiB GPU capacity, not 46 or 110 GiB. A 512 MiB carve-out with a corroborated 109 GiB GTT heap reports 109 GiB, not 109.5 GiB. This matches the amdgpu APU allocation rule. Driver capacity is not proof that an engine can successfully allocate all of it; engine buffers, current usage and runtime validation remain separate.
The read-only inventory also publishes installedMemoryMi from populated
SMBIOS type-17 devices (dmidecode --type 17) and firmwareReservedMi when the
selected uma/carveout_options entry agrees with mem_info_vram_total. Missing
tools, unknown SMBIOS device sizes or conflicting firmware evidence remain
unknown; installed capacity is never inferred by adding Linux and GPU counters.
Only the aggregate capacities are published, not DIMM identifiers or raw DMI
output. These two inventory fields do not change scheduling, model estimates,
GPU eligibility, firmware settings or reboot plans.
The dashboard separates installed RAM, the fixed GPU carve-out, Linux-visible
RAM, the dynamic GPU ceiling and driver-reported model capacity/allocation
domain. The dynamic value is inside Linux RAM and
must not be added to it. Where the measured totals reconcile, the difference
installed - fixed - Linux-visible is labelled other firmware/platform memory,
not extra GPU capacity. Existing CPU/GPU gauges remain conservative model-budget
views, not hardware-inventory totals. Fixed GPU free memory is not inferred
from Linux MemAvailable; without vendor usage metrics it stays unknown.
Neither a GTT ceiling nor a Kubernetes memory request protects RAM for a future GPU process. Protecting a dynamic AI budget would require bounded non-AI/host workloads, system headroom and verified GPU/cgroup accounting under load. This implementation does not install that protection or claim a hard GPU limit. Firmware reservation choices remain exactly those advertised by the BIOS; there is no arbitrary larger carve-out or automatic firmware/TTM change.
Explicit Ansible preparation¶
Administrators can also configure the fixed firmware reservation and dynamic GPU memory ceiling using sliders in System → Hardware → GPU nodes → GPUs → GPU Configuration AMD → Shared GPU memory. Only discovered firmware options are offered, with a 16 GiB CPU/OS allowance for the dynamic ceiling and explicit restart confirmation. The host worker verifies actual RAM after changing the carve-out before applying TTM through the role below. This does not expand the scheduler's capacity using unverified projected RAM. See the memory workflow and recovery contract.
The normal user entrypoint is now System → Hardware → GPU nodes → Host preparation for both new and existing machines. It uses the same Ansible role below through a local root worker, with explicit package/reboot confirmation and resumable verification. System also provides administrator restart/shutdown controls. See host management, including the bounded experiment mode for unreviewed GPU combinations. The command below is a low-level local maintenance/check-mode entrypoint, not a second installer workflow.
Generic AMD detection never performs a kernel, firmware, driver or ROCm upgrade. The independent preparation entrypoint is disabled by default:
ANSIBLE_ROLES_PATH=magic-host/roles ansible-playbook \
-i magic-host/inventory/localhost.yml magic-host/playbooks/gpu-prepare.yml \
-e gpu_compatibility_prepare_host=true \
-e gpu_compatibility_profile=strix-halo \
-e @/path/to/reviewed-host-preparation.yml --check --diff
Create the private input file only after reviewing the installed hardware and
the exact Ubuntu/ROCm support combination. The role currently bounds preparation
to Ubuntu 24.04 and 26.04; this OS check is not a validated-stack claim. Review
check-mode output before explicitly rerunning without --check.
| Variable | Contract |
|---|---|
gpu_compatibility_prepare_host |
Defaults to false; all preparation is behind this opt-in. |
gpu_compatibility_profile |
Must be strix-halo, with matching hardware detected locally. |
gpu_compatibility_package_versions |
Mapping of allowed APT package names to exact versions. Empty by default; no automatic latest, repository additions, downgrades or vendor installer scripts. |
gpu_compatibility_ttm_limit_mib |
null leaves existing configuration unchanged; a positive integer writes the managed next-boot TTM mapping limit; 0 removes only that managed override. |
gpu_compatibility_system_reserve_mib |
At least 8192 MiB must remain outside an explicit TTM mapping limit. This is a safety floor, not a universal sizing recommendation. |
Allowed pinned package families are linux-firmware, rocminfo,
linux-firmware-amd-graphics on Ubuntu 26.04, the native
linux-generic meta-package on Ubuntu 26.04, the
linux-generic-hwe-24.04 meta-package on Ubuntu 24.04, and explicit
linux-image-*, linux-modules-*, linux-modules-extra-* and linux-headers-*
packages. All require exact package-version pins. Select versions from already
trusted Ubuntu repositories; the role
does not invent a certified kernel/firmware stack. An empty package map is a
valid read-only preparation check after helper installation.
TTM changes affect /etc/modprobe.d/90-magicstick-ttm.conf and refresh the
initramfs. They are not applied to a running GPU. Package changes can also need
a reboot, but the role never reboots, blacklists/loads amdgpu, enables KMM, or
changes Kubernetes GPU eligibility. Its evidence timer publishes non-secret
diagnostic metadata, not support claims. Schedule a controlled reboot separately and
rerun diagnostics afterwards. To remove a Magic Stick TTM override, explicitly
set its limit to 0, review the diff, apply and reboot. Keep a previously
working kernel available when planning a kernel upgrade.
Device-specific dashboard diagnostics¶
The read-only host preflight publishes a separate displayDevices PCI inventory
with name, driver, vendor/device IDs, architecture and memory evidence where
available. The AMD-only devices and Strix Halo memory-safety contract stay
unchanged. Appliance.status.hardwareOperators.<module>.devices exposes physical
cards and optional per-engine results, never time-slicing replicas.
The dashboard offers all GPUs on a node or one exact device. Each Ollama/vLLM
request requires confirmation and records identity-bound device-validation-*
annotations on the existing vendor ModuleActivation, not Git-owned appliance
spec or model settings. The API validates all selected targets before writing.
Cross-provider conflicts report which providers already accepted the request;
there is no automatic retry.
The controller serializes tests across devices and engines. Each Job requests a
real GPU resource and selects the current node/boot/hardware identity. AMD DRA
uses only the matching ResourceClaim; NVIDIA retains its existing nvidia
runtime and nvidia.com/gpu resource. No host-device mounts, CPU fallback,
automatic model removal or kernel/memory changes are introduced. Tests may wait
for active workloads. Failed or stale tests never disable a GPU or automatically
rerun. Missing runtime-image evidence is not a verified pass. Device diagnostics
do not repin model runtime images.
The current exact-device path supports one physical GPU per vendor on a node, including one AMD plus one NVIDIA and their sharing replicas. Same-vendor multi-card nodes remain visible but cannot be blindly tested through a node-level legacy allocation. MIG, missing physical inventory and providers without an approved diagnostic are explicitly unavailable; an all-GPU request never silently skips them. During rolling upgrades the old inventory remains visible, but exact-device diagnostics stay disabled until fresh host evidence arrives.
A small computation proof in the exact inference image¶
Run magicstick-hip-smoke.py in a disposable diagnostic container based on the
same pinned ROCm vLLM image used for models, with access to the selected GPU.
The helper requires the image's own ROCm-enabled PyTorch; it does not install
dependencies. Bind or stream the script and invoke:
python3 /path/to/magicstick-hip-smoke.py --expected-architecture gfx1151
It rejects CPU/CUDA-only builds, unavailable GPUs and unexpected architectures.
It performs a small FP32 GPU matrix multiplication, synchronizes the device,
checks that the result is still a GPU tensor, and compares finite output with
an exact CPU reference. JSON success still includes
modelInferenceValidated: false.
Follow that test with tiny Ollama and vLLM model requests separately, using the actual configured image digests. Verify GPU execution, correct responses, memory accounting, repeated requests and restart behavior. Record image, hardware fingerprint and outcomes; invalidate old evidence after relevant host or image changes. Do not silently enable CPU fallback, architecture overrides, Vulkan, or an alternative device plugin when a ROCm test fails.
A dashboard engine result of passed means that the bounded GPU smoke test
succeeded for that host and exact runtime image. It does not certify every
model, quantization, context length, answer quality, sustained performance or
restart scenario, and it does not establish GPU cgroup memory accounting. A
provider becomes Ready from hardware and registered-resource readiness,
independently of either engine test. Inspect the separate Ollama/vLLM diagnostic
results when investigating a model failure. The compatibility profile remains
experimental even after those small tests pass.
The operator publishes configured AMD images through
magicstick-gpu-runtime-images. KubeAI imports generic resource profiles first
and these image pins last, without duplicate inline AMD image tags. This keeps
Helm values merging from restoring a different runtime. Catalog release tags
work without validation; successful optional tests can pin their actual digest.
runtimeReady requires the exact configured image reference in
KubeAI's installed configuration and Ready controller Pods with matching
configuration checksums. See the Flux values-reference contract.
Use Verify Ollama or Verify vLLM under System → Hardware → GPU nodes
to request a bounded test only for that engine and node. Its scoped request does
not reset other engine results. hardware validate --yes and the TUI's
validation action still request both engines on matching nodes. Profile
saves and host preparation never start them. A request is bound to the host boot,
hardware and runtime images at its first test; changed evidence becomes stale
without automatically launching another test. Request a new run to refresh it.
Legacy automatic host-<request-id> requests are retired on upgrade.
The manually requested model fixture uses the raw completion 2 + 2 = with a two-token
output budget and checks the exact arithmetic answer. It deliberately avoids
model-specific chat and thinking templates: a tiny model's instruction-following
failure must not be confused with a broken GPU kernel. The same fixed fixture
can be compared on CPU and GPU; arbitrary or merely nonempty output never passes.
Local checks¶
python3 -m unittest discover -s magic-host/roles/gpu-compatibility/tests -v
ANSIBLE_ROLES_PATH=magic-host/roles ansible-playbook --syntax-check magic-host/playbooks/local.yml
ANSIBLE_ROLES_PATH=magic-host/roles ansible-playbook --syntax-check magic-host/playbooks/gpu-prepare.yml
Related operational contracts: host automation, GPU operations, modules, and operator orchestration.