EBOX × EAR CLOUD EAR CLOUD
Home / Core Technologies / Edge AI Voice Algorithms

Edge Voice AI: hear it clearly, answer it right

The first gate on an industrial site sits on the terminal chip: wake, denoise and detection all run locally — only valid speech goes upstream. Noise-resistant, bandwidth-saving, privacy-preserving.

WakeNet10 wake word2–3ms per frame on-deviceMultiNet offline fallback
PIPELINE · ON-DEVICE CHAIN

Five steps, all on the terminal chip

From capture to "should this be uploaded?" — the edge does it in one pass; the cloud only receives speech worth recognizing.

WakeNet10 wake word

Custom brand wake word with a TTS-synthesis training pipeline; accuracy 95–98% of real human recordings.

AFE audio front-end

AEC echo cancellation + noise suppression + array beamforming, locking onto the voice amid noise.

VADNet detection

Neural VAD separates speech from pauses at millisecond level.

MultiNet offline fallback

30 on-device commands work without network — the line never goes deaf.

Audio stream upstream

Encrypted WebSocket, cloud-grade LLM recognition.

ASR COMPARISON

ASR · Model comparison

EBOX uses a new-generation open-source LLM-architecture recognition model.

ModelLanguagesStrengths & limitsLatency
Qwen3-ASR-1.7BUsed by EBOX30 languages + 22 Chinese dialects
zh/en/ja/ko/vi/th/de/es/pt/id全覆盖
LLM-architecture recognition; 10k-token industry context biasing — the rarer the SKU or batch number, the better it works; real-time inference on pure CPU, no GPU neededStreaming recognition while speaking
0.1–0.3s per second of audio
full text when you finish
Legacy local ASR
Paraformer class
Mostly zh/enMature architecture, but limited language coverage and weak hot-word expansionFast, but languages limited
Whisper90+ languagesStrong general-purpose; no hot-word mechanism, hallucination risk, heavy computeHard to run real-time on CPU
Cloud APIBroad coverageDepends on external network, voice data leaves the plant, pay-per-use foreverAffected by network jitter
TTS COMPARISON

TTS · Engine comparison

Local synthesis engine — clean licensing, no cloud dependency.

ModelLanguagesStrengths & limitsLatency
Piper / VITS local synthesisUsed by EBOX60+ languages
Vietnamese with 3 voices at launch
Fully local CPU synthesis, MIT-family license safe for commercial use; one runtime, engines added per language; front number/bin-code reading-normalization layer for unambiguous announcementsFirst packet 100–300ms
streaming playback, speaks while generating
MeloTTS6 languagesCPU real-time, zh/en mixed reading — a multi-language backup engineReal-time
Cloud TTS140+ languagesGreat quality but needs internet — unavailable on industrial intranetsAffected by network RTT
95–98%
wake-model accuracy (trained against real human recordings)
2–3 ms
per-frame on-device processing
30
offline commands when the network drops
0
false-wake audio uploaded (silently filtered at the edge)
Next:Full-Duplex Voice

Want to see how these algorithms perform on your floor?

Book a demo — bring your own industry hot words to try.