<aside> 🗣️
Document purpose: Make voice and accessibility core architectural interfaces while keeping them optional to install or enable.
Document status: Architecture/design specification. Current implementation status is controlled by the roadmap and stage evidence.
</aside>
flowchart LR
U["User speech"] --> STT["STT"]
STT --> C["Conversation Runtime"]
C --> I["ISAC"]
I --> TTS["TTS"]
TTS --> O["ISAC voice"]
Support local STT/TTS first where practical, plus optional BYOK providers such as ElevenLabs for voice output. Cloud voice must never be required for core operation.
ISAC may learn speaking pace, vocabulary, preferred spoken response length, terminology and whether the user prefers shorter voice responses. This adapts communication but does not imitate the user's identity.
Potential capabilities include TTS, STT, larger/readable terminal presentation, reduced clutter, simple explanations, repeat/rephrase, slower voice and keyboard-only navigation.
Voice can be enabled during first run or added later. A text-only ISAC remains a fully supported configuration.
Voice provider API keys remain in the credential vault, never in memory or conversation logs.