Artificial Humanity presents

Prosodia

The on-device dramatic audiobook engine.

Most text-to-speech reads. Prosodia performs. It stages every book as a production — read ahead, interpreted, and voiced with emotional direction — entirely on your device. No cloud, no telemetry, no page of your book leaving your hands.

The Company

The Director

An on-device language model that reads ahead of the narration, annotating each passage with emotional direction — valence, arousal, tension — the way a director blocks a scene before the actors take it.

The Actor

A neural voice trained by our Sonora project, performing the Director's notes through a Rust synthesis core — small enough to live on a phone, expressive enough to be worth listening to.

The Stage

A coordinator that keeps the performance flowing — gap-free audio, look-ahead rendering, and graceful pacing across sentence and speaker boundaries.


Hear the direction

Sonora’s de-risk milestone (July 2026): a multi-speaker, 24 kHz voice whose energy can be directed on a continuous dial — the first trained channel of the Director’s valence/arousal/tension vocabulary. Same text, same speaker, same seed; only the direction changes. Artifacts in the model registry.

Directed hushed (energy −1) — “The lighthouse keeper woke before dawn…”
Neutral (energy 0)
Directed emphatic (energy +1)

Amplified direction

The Director can also push past the trained range with classifier-free guidance — extrapolating the model’s own learned direction at render time.

Neutral — “We won! We actually won the championship!”
Energy +1
Energy +1, guidance ×3

Long-form stability

A single unbroken 103-second render — no drift in pace, loudness, or clarity. Audiobooks need paragraphs, not sentences.

~1¾ minutes of continuous narration, one pass
Status: the energy channel is trained and verified (controllability ρ ≈ 1.0, speaker identity preserved, intelligibility unchanged — the full evaluation ships with the model). Valence and tension training is underway on the same plumbing, alongside a 247-voice casting space.

Follow along