Shanks

Speech in and out natively through the Gemini Live API — no transcribe-then-synthesise chain. It reads hardware state and drives the desktop, so answering and acting are the same turn.

Shanksvoice session · gemini-live
idle
Transcriptwhat was said, and what it did
Runs automatically · the face is drawn live, dot by dot

Native duplex audio

Audio goes straight in and out of the Gemini Live API. Nothing is transcribed to text and re-synthesised, so the turn does not stack three models' worth of latency.

It can actually act

Reads hardware state, sets volume, toggles Wi-Fi, launches applications and looks things up — the answer and the action are one turn, not a suggestion to go do it yourself.

Remembers across sessions

What you tell it persists, so preferences and context survive a restart instead of resetting every time you open the mic.