Qwen3.6-35B-A3B Runs on a 16 GB GPU for Local Coding Agents
Running a 35-billion-parameter model on a single consumer graphics card usually means compromises. A new open source project called FastLocalAI tries to keep those compromises small. It's a tuned setup for running Alibaba's Qwen3.6-35B-A3B on an NVIDIA GPU with just 16 GB of VRAM, paired with the OpenCode coding agent. The project is available on GitHub at https://github.com/24high/FastLocalAI. It uses llama.cpp's llama-server and moves part of the model into regular system memory so the rest fits on the card.
Why the model doesn't fit and how it works anyway
Qwen3.6-35B-A3B is a Mixture-of-Experts model. The quantized Q4_K_M version from Unsloth, which the project uses by default, is about 22 GB. That is too large for a 16 GB card. The launcher solves this by splitting the model. Attention and dense layers stay on the GPU, along with some of the expert layers. The remaining experts sit in system RAM, handled by llama.cpp's --n-cpu-moe option. By default, the experts of 28 of the model's 40 MoE layers stay in RAM. According to the project, that leaves about 2 GiB of VRAM free at a 128k context on a 16 GB card. Lowering the number speeds things up but uses more VRAM.
