AI assistants & agents
Strata - Local LLM
Run a 125B-parameter LLM on your own gaming PC, free and offline.
AI assistants & agents
Run a 125B-parameter LLM on your own gaming PC, free and offline.
Product insights
Server-class model quality on consumer hardware, with no per-token cost.
Models this large normally need a data-center GPU or a paid cloud API.
Gamers and developers with a 12 GB+ graphics card and 32 GB+ RAM who want a private local LLM.
Free and open source (MIT). You only need your own hardware and about 80 GB of disk space.
Strata is the project people usually mean by “Strata LLM”: a local LLM engine and app that runs the 125-billion-parameter Qwen3.8-Flash-Next on a Windows or Linux gaming PC. The most-used experts stay on the graphics card, the rest sit in RAM, and a lookup table lives on the SSD. On an RTX 5070 it writes around 50 to 95 tokens per second depending on the model size. It chats, writes code, reads pictures, and connects to coding agents, all offline.
Get Strata from GitHub, then double-click START-HERE.bat on Windows or run ./setup.sh on Linux.
The installer checks your card and RAM and recommends a size, then downloads about 70 GB once.
Open http://127.0.0.1:8080 to chat, or point any OpenAI-compatible app at http://127.0.0.1:8080/v1.
It usually refers to Strata, an open-source app on GitHub that runs the 125B Qwen3.8-Flash-Next model on a consumer gaming PC. The name is also used by unrelated research projects, such as a context-caching paper.
An NVIDIA RTX 20–50 series or a recent AMD Radeon card with 12 GB of VRAM or more, at least 32 GB of RAM (64 GB runs every size), and about 80 GB of free SSD space.
Yes. It is MIT-licensed open source and runs entirely on your PC, so there are no per-token fees.
LM Studio runs many open models that fit your machine. Strata is built around one very large model and squeezes it onto a gaming PC by splitting it across GPU, RAM, and SSD.
Official website snapshot
github.com
Strata (often searched as “Strata LLM”) is a free, open-source app that runs Qwen3.8-Flash-Next, a 125-billion-parameter model, on a normal gaming PC with a 12 GB NVIDIA or AMD card. It spreads the model across GPU, RAM, and SSD, then gives you a browser chat and an OpenAI-compatible local API.