Models and routing in Jaisiu
Jaisiu runs local GGUF models on your machine by default. The first install picks a model that fits the RAM and CPU you have. Tasks that the local model can do (code review, ticket triage, draft replies) stay on your box. Tasks that need a bigger model route to the Pryzm broker, which is a token pool you already paid for with your Jaisiu license. Routing is automatic and per-task, not per-seat — agents decide per call.
Routing rules
- Local first. If a task is within your model's context and capability, it runs locally. No token spend.
- Broker fallback. If a task is too large or too hard, the agent hands it to the broker. The broker uses a pooled token cap; once the pool is empty, agents pause.
- No silent cloud. A task never leaves your machine unless it goes through the broker and the broker log shows it.
Bring your own model
Jaisiu reads GGUF files from ~/.jaisiu/models/. Drop a file in, restart the gateway, and the next task can use it. Compatible with anything llama.cpp loads (Qwen, Gemma, DeepSeek, Mistral, Phi, Llama, GLM, and the small Mistral-derivative set we ship by default).
How much does the broker cost?
The token pool is part of your Jaisiu license. Basic includes 200M tokens / month, Fleet includes 4B / month. Add-on packs (100M for $29, 400M for $69) extend the pool without a subscription change. Hard cap — agents pause at zero, never overage.
See the model broker reference for the full routing config and CLI flags.