$npx skillfedfor your agent
REPO

A 125B model at 74 tokens per second on a gaming GPU is real, with caveats

on: Niko1221/Strata

Strata solves a specific, concrete problem: a 125-billion-parameter mixture-of-experts model normally demands server-grade hardware with hundreds of gigabytes of GPU memory. A gaming PC has 12-24 GB of VRAM. The gap looks unbridgeable until you notice that the model's 24,576 specialists — the "experts" in the MoE architecture — are not all needed at once. Each token requires only 10 of them. Strata exploits that sparsity by distributing the model across every tier of a consumer machine: the most-used experts stay on the GPU, all experts live in RAM, and a large lookup table sits on the SSD. The GPU and CPU work in parallel so neither tier sits idle waiting for the other.

The measured numbers on an RTX 5070 with 64 GB of RAM are worth taking seriously. The IQ2_XS quantization — the recommended starting point — produces 74 tokens per second on short chat and 60 tokens per second at 128K context. The IQ3_S quantization, which the README claims matches the full uncompressed model on published benchmarks, runs at 52 and 41 tokens per second respectively. A 32K-token prompt is ingested in roughly 17 seconds at Q2_0. These are not theoretical peaks; they come from a specific, named hardware configuration.

A speculative-decoding layer adds another layer of efficiency. A small internal helper model guesses several tokens ahead, and the large model verifies all guesses in a single forward pass, keeping only the ones it agrees with. The README puts the practical speedup at 1.6 to 1.8 times. Crucially, the big model makes every final decision, so quality is not traded away for throughput.

The Coder variant is worth noting separately. ISTA-DASLab pruned half the experts, keeping those most relevant to code and tool use. The result fits in 32 GB of RAM rather than 64 GB, runs a 262K context window on 64 GB, and scores 91% of the full model on SWE-bench Verified and 99% on LiveCodeBench according to its authors. For agent workloads that are primarily code generation or tool-calling, that tradeoff is probably worth it.

For agent builders specifically, the OpenAI-compatible endpoint at http://127.0.0.1:8080/v1 means dropping this into any existing agent framework is a configuration change, not a code change. Anthropic-API-compatible apps are also supported via a /v1/messages route. The model handles images, adjustable thinking depth, and up to 128K context depending on quantization and available RAM.

The honest constraints: this is a single-request-at-a-time server, so concurrent agent workloads are not the target. The first cold start locks 35-55 GB into RAM and can freeze the machine for several minutes. And the whole thing requires an NVIDIA RTX 30-series or newer card — AMD and Apple Silicon are not in scope here. Within those boundaries, running a 125B-parameter model locally at readable speed on hardware most developers already own is a genuine capability shift for anyone who wants inference without API costs or data leaving the machine.

A practical, well-documented system for running a 125B MoE model on a single gaming GPU by distributing experts across VRAM, RAM, and SSD in parallel.

Install it

Sources & links