<aside> 📊

Document purpose: Evaluate ISAC's claims using repeatable system-level evidence rather than model-only benchmark scores.

Document status: Architecture/design specification. Current implementation status is controlled by the roadmap and stage evidence.

</aside>

Benchmark Families

Resource, intelligence, offline, memory, personalisation, skills, initiative, latency, reliability, portability and efficiency.

Resource Metrics

Idle RAM, active RAM, peak RAM, VRAM, CPU, model load/unload time, retrieval memory, storage growth, battery impact and thermal behaviour where measurable.

Smart-per-GB

Compare the same task set across:

Measure task success, RAM, VRAM, latency and token cost.

Offline Evaluation

Test general questions, coding, files, system diagnosis, learned skills, memory retrieval, Knowledge Vault retrieval and conversation with internet physically unavailable.

Memory Evaluation

Correct retrieval, false retrieval, contradiction handling, user correction, long-term recall, consolidation, storage growth, RAM impact, encryption and recovery.

Personalisation Evaluation

Measure whether ISAC learns expertise, communication style and workflow preference, and whether explicit corrections reverse bad inferences.

Skill Evaluation

Success, failure detection, recovery, improvement, regression and permission compliance.

Initiative Evaluation