# Models and routing in Jaisiu

Jaisiu runs local GGUF models on your machine by default. The first install picks a model that fits the RAM and CPU you have. Tasks that the local model can do (code review, ticket triage, draft replies) stay on your box. Tasks that need a bigger model route to the Pryzm broker, which is a token pool you already paid for with your Jaisiu license. Routing is automatic and per-task, not per-seat — agents decide per call.

## Routing rules

- **Local first.** If a task is within your model's context and capability, it runs locally. No token spend.
- **Broker fallback.** If a task is too large or too hard, the agent hands it to the broker. The broker uses a pooled token cap; once the pool is empty, agents pause.
- **No silent cloud.** A task never leaves your machine unless it goes through the broker and the broker log shows it.

## Bring your own model

Jaisiu reads GGUF files from `~/.jaisiu/models/`. Drop a file in, restart the gateway, and the next task can use it. Compatible with anything llama.cpp loads (Qwen, Gemma, DeepSeek, Mistral, Phi, Llama, GLM, and the small Mistral-derivative set we ship by default).

## How much does the broker cost?

The token pool is part of your Jaisiu license. Basic includes 200M tokens / month, Fleet includes 4B / month. Add-on packs (100M for $29, 400M for $69) extend the pool without a subscription change. Hard cap — agents pause at zero, never overage.

See the [model broker reference](https://pryzm.at/efss/viewer.html#docs%2Fgetting-started%2FPRYZM-MODEL-BROKER.md) for the full routing config and CLI flags.
