$ cd ~/products/iris/core
nf iris suite · 1/4
Iris Core
Offline Voice Assistant
You speak, it listens, thinks and talks back — with no network request ever leaving the machine. Wake word, transcription, reasoning and speech all run on-device.
- +Wake word — Zipformer keyword spotting via sherpa-onnx
- +Speech-to-text — Streaming Zipformer, no cloud round-trip
- +Local reasoning — llama.cpp accelerated on the Metal GPU
What it does
Core is the voice pipeline. It listens for the wake word, transcribes what follows as you speak, passes it to a local language model with the conversation so far, and speaks the answer back. Every stage runs on the machine; no audio or text is sent anywhere.
Speaking without gaps
Speech synthesis runs one sentence ahead of playback: while one sentence plays, the next is already being synthesised. There is no pause between sentences, and the first audio arrives as soon as the first sentence is complete, typically one to two seconds, instead of after the whole answer.
Interrupting
The stages talk over a typed event bus instead of calling each other. Speaking over Iris stops its speech immediately, and cancelling a request reaches every stage at once.
Models
Any GGUF model runs through llama.cpp, so changing the model is a configuration change, not a code change. Conversations keep multi-turn history, and the assistant's persona is a YAML profile.
Status
- [x]End-to-end voice conversation, fully on-device
- [x]Wake word, streaming transcription and pipelined speech
- [x]Multi-turn history and persona profiles
- [x]CI on macOS arm64
- [ ]Published latency and transcription-accuracy benchmarks
- [ ]Noise suppression and automatic gain control
- [ ]Voice selection