The on-device dramatic audiobook engine.
Most text-to-speech reads. Prosodia performs. It stages every book as a production — read ahead, interpreted, and voiced with emotional direction — entirely on your device. No cloud, no telemetry, no page of your book leaving your hands.
An on-device language model that reads ahead of the narration, annotating each passage with emotional direction — valence, arousal, tension — the way a director blocks a scene before the actors take it.
A neural voice trained by our Sonora project, performing the Director's notes through a Rust synthesis core — small enough to live on a phone, expressive enough to be worth listening to.
A coordinator that keeps the performance flowing — gap-free audio, look-ahead rendering, and graceful pacing across sentence and speaker boundaries.
Sonora’s de-risk milestone (July 2026): a multi-speaker, 24 kHz voice whose energy can be directed on a continuous dial — the first trained channel of the Director’s valence/arousal/tension vocabulary. Same text, same speaker, same seed; only the direction changes. Artifacts in the model registry.
The Director can also push past the trained range with classifier-free guidance — extrapolating the model’s own learned direction at render time.
A single unbroken 103-second render — no drift in pace, loudness, or clarity. Audiobooks need paragraphs, not sentences.