Running agents at home is something most people can afford to do, and something that will improve their lives, often on hardware they already own. A year ago this was not the case. Local models were the domain of large businesses, and highly technical tinkerers.

We’ve combed 184 websites and collected Hardware prices and sales data, model configs, and hardware registries to put together a high level overview of the industry, community, and technology

Why now

Key takeaway: a model that fits one 24 GB GPU now scores where the best model in the world was in February 2026. It came out about six months after that frontier model.

chart-qwen-same-size.png

Qwen's dense 27B models and 33B parameters (dots) against every new frontier record (lab logos), all scored on the same version of the Artificial Analysis index. The dashed line is the time from GPT-5.3 Codex to Qwen3.8 27B at the same score. Hollow dots are scores AA estimated. Source: Artificial Analysis, September 2026 snapshot.

software-before-after.png

1. Hardware

Key takeaway: memory size decides what you can run, memory bandwidth decides how fast. Big unified-memory boxes hold the most model per dollar; discrete GPUs are the most speed per dollar.

Two numbers decide almost everything