Reproducible inference and training optimization for consumer AMD Radeon RDNA GPUs — what runs, how it compares to stock, and whether to use upstream, an image, or a patch.
RDNA4 (gfx1200/1201) primary · vLLM / llama.cpp / DiffSynth / ComfyUI · LLM·VLM·video · GEMM / attention / MoE kernels · RCCL / QuickReduce.
| Topic | Platform | Highlight | Detail |
|---|---|---|---|
| Qwen3.6-27B-FP8 · vLLM TP=2 | 2×R9700 · graph · MTP on | Live+ TG128 decode 74.1 vs Community 64.3 / Stock 46.1 · 256k @ c≈2 (FP8-QKV + UA) | PERFORMANCE |
More models will land as additional rows here; full tables stay in each topic’s PERFORMANCE.md.
- Docker is user space only; host GPU must pass
rocminfo/rocm-smi. - Results do not generalize across VRAM, GPU count, or PCIe/NUMA/P2P without revalidation.
working/is untracked; only reviewed artifacts enter the release tree (promotion rules).
| Path | Purpose |
|---|---|
docs/ |
Getting started, methodology, hardware, upstream |
frameworks/ · kernels/ · comm/ |
Released topics |
working/ |
Optimization workbench |
LICENSE. Upstream subtrees keep their own licenses.