<aside> 🗣️

Document purpose: Make voice and accessibility core architectural interfaces while keeping them optional to install or enable.

Document status: Architecture/design specification. Current implementation status is controlled by the roadmap and stage evidence.

</aside>

Voice Pipeline

flowchart LR
    U["User speech"] --> STT["STT"]
    STT --> C["Conversation Runtime"]
    C --> I["ISAC"]
    I --> TTS["TTS"]
    TTS --> O["ISAC voice"]

Providers

Support local STT/TTS first where practical, plus optional BYOK providers such as ElevenLabs for voice output. Cloud voice must never be required for core operation.

Communication Learning

ISAC may learn speaking pace, vocabulary, preferred spoken response length, terminology and whether the user prefers shorter voice responses. This adapts communication but does not imitate the user's identity.

Accessibility

Potential capabilities include TTS, STT, larger/readable terminal presentation, reduced clutter, simple explanations, repeat/rephrase, slower voice and keyboard-only navigation.

Setup

Voice can be enabled during first run or added later. A text-only ISAC remains a fully supported configuration.

Security

Voice provider API keys remain in the credential vault, never in memory or conversation logs.