A guide on what models you can run on all three configurations of the Framework Desktop.
Key specifications and details
The Framework Desktop comes with three memory configurations built around AMD's Ryzen AI Max APU, codenamed Strix Halo: 32GB, 64GB, and 128GB. When Framework Desktop launched in 2025, much of the local AI community naturally focused on the 128GB version. It's easy to see why: Ryzen AI Max brought 128GB of memory shared between the CPU and integrated GPU to a desktop system, making it possible to run local models that would not fit into the memory on most consumer graphics cards.
Fast forward to September 2026, and memory supply constraints across the industry have pushed up prices, including for the Framework Desktop. Fortunately, this has coincided with the release of new, smaller open-weight models that are well suited to the 32GB and 64GB configurations, making the less expensive versions more capable and helping to offset the impact of higher memory prices. The new Qwen3.8-27B can run within 32GB with a suitable quantization, while models such as Qwen Image and MiniMax-H3 bring image, video, and audio generation within reach on the same configuration.
The 128GB configuration still gives you the widest choice of models, quantizations, and context sizes, but it is no longer the only version worth considering. This guide looks at what you can run on all three configurations, with specific model and quantization recommendations, tested Linux performance, and complete image and video workflows. The video below goes deeper into inference engines, quantization, mixture-of-experts models, and speculative decoding.
Popular user interfaces like LM Studio and Ollama use llama.cpp as their backend. Lemonade uses it as its main LLM backend and also lets you choose others.
There are now several experimental llama.cpp forks made specifically for Strix Halo. Nathan Wilson, from the Strix Halo Homelab Discord, has done a lot of work on a Strix Halo fork that improves performance on this hardware. Some of that work is now being prepared for upstream llama.cpp. Other engines target specific model families. DwarfStar, for example, is dedicated to DeepSeek V4 Flash. AI Toolbox Cockpit provides pre-built Linux containers for the stable and experimental llama.cpp and DwarfStar configurations tested on the Framework Desktop.
Our Take
What stands out to me here is AMD. I think it has a slight, maybe even a fairly solid, lead over the other companies and its rivals when it comes to processors.
Just like NVIDIA leads in graphics cards, AMD is out in front on the CPU side, and both of them are aiming hard at AI.
Panos, GamingBroject editor
vLLM is another option here, but it is more geared towards discrete GPUs and data centers. It is designed to serve many users at once, using a different architecture to maximize throughput across concurrent requests. It can run on Strix Halo, but at the low concurrency of a personal AI server like the Framework Desktop, llama.cpp usually performs better. vLLM is also not supported on as many types of hardware as llama.cpp.
With 64GB, you can run Qwen3.8-27B or Qwen3.6-35B-A3B at higher precision and still leave room for context. You can also run a larger model such as Qwen3.5-122B-A10B. Its 52.5GB Q3 file fits, but leaves less memory for context than the smaller models. This is the main choice at 64GB: a larger model, higher precision, more context, or more than one model loaded at once.
Inkling-Small remains another option if you need multimodal input, because it accepts text, images, and audio. Its 82.3GB two-bit file leaves more memory for context, while the 107GB and 119GB three-bit files use that memory for higher precision instead.
Official video
Official video, verified official YouTube channel via frame.work.
Official source: Framework Official Blog
Follow GamingBroject on Google: add us as a preferred source




Community
0 comments