<aside> ⚙️
Document purpose: Make ISAC local-first, model-agnostic, resource-aware and extensible without allowing any model to become ISAC's identity.
Document status: Architecture/design specification. Current implementation status is controlled by the roadmap and stage evidence.
</aside>
The default ISAC install should use approved free/open-weight local models appropriate to the device, such as suitable Qwen- or Kimi-family releases where licensing, runtime support and benchmarks are acceptable.
| Layout | Behaviour |
|---|---|
| Unified | One stronger local model serves Fast, Main and Deep roles through different GCF budgets. |
| Dual | Two medium/local models split roles according to measured performance. |
| Triad | Three smaller models provide dedicated Fast, Main and Deep/specialist roles. |
| Automatic | Recommended. GCF benchmarks viable layouts and selects the best smart-per-GB configuration. |
Fast/Main/Deep are cognitive roles, not mandatory separate downloads. Vision, embedding, reranking, STT, TTS and specialist models may also exist as utility roles.
On install, ISAC should verify and benchmark actual RAM/VRAM, load time, latency, throughput, tool reliability and task quality on the user's hardware, then store the device-specific model profile.
Optional provider support should include:
ISAC should provide a step-by-step setup flow inside the installer/terminal, test authentication, explain possible cost/privacy impact and let the user define when cloud escalation is allowed.
Support OpenAI-compatible endpoints, local network model servers, remote self-hosted models and advanced custom providers through adapters.
GCF should prefer the least expensive model that can reliably satisfy the task while respecting privacy, cost and resource policy.