Qwen 3.8 27B on One GX10: The Receipts
I cheered Qwen 3.8 27B before the weights existed. Then I loaded NVFP4 on the ASUS Ascent GX10 and timed it. Threads, comments, and the numbers from that week.
Open-weight models, local inference stacks, VRAM planning, and homelab setups for running AI on your own hardware.
View all local AI articlesFourteen pillar articles at /blog/<slug> — always on the blog URL, never moved under /local-ai/.
I cheered Qwen 3.8 27B before the weights existed. Then I loaded NVFP4 on the ASUS Ascent GX10 and timed it. Threads, comments, and the numbers from that week.
I posted Qwen 3.8 27B three times and meant it. Why a 27B open-weight model still matters on a 128GB Spark-class box, and how I would actually load it.
Hermes agent, DeepSeek V4 Flash 0731, ASUS Ascent GX10 — local only. How I shipped the Moon Trail browser game without cheating on a cloud model.
I unboxed an ASUS Ascent GX10 — NVIDIA DGX Spark in a 150mm cube — and started running local models on 128GB of unified memory. No brochure. First notes from the bench.
Cornerstone WikiWayne guide: Best GPU for Local AI (2026). Open-weight, practitioner-tested local AI.
Shopping the secondary market without overspending on VRAM you cannot use.
Load checkpoints, wire KSampler, export PNGs locally.
Cornerstone WikiWayne guide: ComfyUI Local Stable Diffusion Guide. Open-weight, practitioner-tested local AI.
What `-ngl` / GPU layer sliders actually do.
Compose services for local chat without cloud relay.
Quant-specific VRAM bands for Meta Llama 3 8B class models.