Traditional voice terminals are half-duplex — you cannot cut in while the device speaks. EBOX achieves true interruptible full duplex with three layers: cut in anytime during playback and the device goes quiet within 100ms to listen.
Same voice interaction — the whole experience gap lives in one question: can you cut in?
Hold to talk → wait for it to finish → then speak. Want to correct? Let it read the whole list. Especially painful in a noisy plant.
Cut in, interrupt, correct — anytime. The device hears itself and hears you; the rhythm feels like talking to a colleague.
First the device "hears itself", then "hears you", and the scheduler decides who speaks first.
Neural VAD distinguishes speech / pause / noise in 75dB plant noise, sensing who speaks at millisecond level — the first reaction to an interruption comes from the edge, not the cloud.
Millisecond local response · network-free
Dual-filter AEC (linear + non-linear) cancels the device's own playback echo with 40dB suppression — the device "knows what it is saying" and never mistakes its own voice for a command.
Espressif 2026 new-gen full-duplex AEC
A 160ms acoustic window continuously judges dialogue state (listening / speaking / pausing / idle), with turn decisions in ~250ms — telling "still thinking" from "finished", at human pace.
Open-sourced by Soul AI × Shanghai Jiao Tong University × Northwestern Polytechnical University · Apache-2.0
A 5-minute live demo beats ten pages of specs.