Strata official homepage preview

AI assistants & agents

Strata - Local LLM

Run a 125B-parameter LLM on your own gaming PC, free and offline.

Category
AI assistants & agents
Maker
Niko1221 (open source)
Source
github.com
Updated
2026-10-01
Open on GitHub

Product insights

Strata at a glance

Value proposition

Server-class model quality on consumer hardware, with no per-token cost.

Problem solved

Models this large normally need a data-center GPU or a paid cloud API.

Audience

Gamers and developers with a 12 GB+ graphics card and 32 GB+ RAM who want a private local LLM.

Pricing and access

Free and open source (MIT). You only need your own hardware and about 80 GB of disk space.

Market

  • Local AI
  • Open-source LLM tools

Tech stack

  • Qwen3.8-Flash-Next (125B MoE)
  • GPU + RAM + SSD offloading
  • OpenAI- and Anthropic-compatible API

What is Strata?

Strata is the project people usually mean by “Strata LLM”: a local LLM engine and app that runs the 125-billion-parameter Qwen3.8-Flash-Next on a Windows or Linux gaming PC. The most-used experts stay on the graphics card, the rest sit in RAM, and a lookup table lives on the SSD. On an RTX 5070 it writes around 50 to 95 tokens per second depending on the model size. It chats, writes code, reads pictures, and connects to coding agents, all offline.

How to use Strata

  1. 01

    Download and run the installer

    Get Strata from GitHub, then double-click START-HERE.bat on Windows or run ./setup.sh on Linux.

  2. 02

    Pick a model size

    The installer checks your card and RAM and recommends a size, then downloads about 70 GB once.

  3. 03

    Chat or connect your apps

    Open http://127.0.0.1:8080 to chat, or point any OpenAI-compatible app at http://127.0.0.1:8080/v1.

Key features of Strata

  • 125B model on a 12 GB graphics card
  • Browser chat with a live hardware monitor
  • OpenAI- and Anthropic-compatible local API
  • NVIDIA and AMD, Windows and Linux

What is Strata best for?

  • Private chat and document work
  • A free backend for coding agents
  • Testing a frontier-size open model at home

FAQ about Strata

What is Strata LLM?

It usually refers to Strata, an open-source app on GitHub that runs the 125B Qwen3.8-Flash-Next model on a consumer gaming PC. The name is also used by unrelated research projects, such as a context-caching paper.

What hardware do I need?

An NVIDIA RTX 20–50 series or a recent AMD Radeon card with 12 GB of VRAM or more, at least 32 GB of RAM (64 GB runs every size), and about 80 GB of free SSD space.

Is Strata free?

Yes. It is MIT-licensed open source and runs entirely on your PC, so there are no per-token fees.

How is Strata different from LM Studio?

LM Studio runs many open models that fit your machine. Strata is built around one very large model and squeezes it onto a gaming PC by splitting it across GPU, RAM, and SSD.

Official website snapshot

github.com

Strata (often searched as “Strata LLM”) is a free, open-source app that runs Qwen3.8-Flash-Next, a 125-billion-parameter model, on a normal gaming PC with a 12 GB NVIDIA or AMD card. It spreads the model across GPU, RAM, and SSD, then gives you a browser chat and an OpenAI-compatible local API.