SOMA.

A tiny language model running entirely in this tab — weights, sampling, everything. No server. Nothing you type leaves your device.

Eight recurrent experts with a learned SetAttn-EMA combiner.

– params – of weights – nats/char –
fetching weights…
Opened directly from disk? Browsers block pages on file:// from fetching local files, so pick the weights manually — choose (or drop) the soma.bin sitting next to this page.
For the full experience serve the folder instead: python3 -m http.server → http://localhost:8000

temp 0.80 top-k 40 top-p 0.95 repeat 1.05 length 400
copied ✓
⚡ – tok/s ms/tok – ctx 0 paths entropy ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ – bits mem – backend …

What is this

SOMA is a character-level recurrent language model that runs entirely in your browser. The default model combines eight independently trained recurrent experts with SetAttn-EMA; the original tiny, backprop-free SOMA remains available from the model selector.

Both models read and write bytes directly. The champion is a one-billion-parameter research model compressed to four-bit weights; the tiny model is a 617K-parameter demonstration.

How it runs here

The champion runs through WebGPU using HQQ four-bit ONNX weights. SOMA Tiny uses hand-written WebAssembly SIMD (f32x4) in a Web Worker, with a typed-array fallback.

Weights are fetched once and inference remains on your device. Generation makes no server request. The only telemetry is an anonymous page-view count; the text you type and generate never leaves the tab.