I think you're identifying some of the right problems here. All voice assistants are based on turn-taking, and when the VoiceAI hits one of those failure points and just comes back with "I didn't get that" it leaves the user in a frustrating state trying to debug what's wrong.
I work at SoundHound where we've been worried about these issues. (I'm going to plug our recent work...) Our new approach is to do natural language understanding in real-time instead of at the utterance (turn) taking level. That way we can give the user constant feedback in real-time. In the case of a screen that means the user sees right away that they are understood, and if not, a better hint of what went wrong. For example a likely mistake is an ASR mistranscription for a word or two.
We still need to prove this is a better paradigm for VoiceAI in products that people can try for themselves, and are working towards that goal. I hope that voice interfaces that were clunky with turn-taking will finally be more naturally usable with real-time NLU.
Why not on the same device? Have a separate small simple SoC completely segregated from everything else, except shared battery, with 2 NICs and a physical switch to swap between using the firewall interface and the regular phone. Although this may make more sense for a regular computer plus router, with a cell phone there's multiple radios, not just a single simple IP connection...
Issue is that we would have to get device makers to buy into it, and also trust them that they show us everything. Also we wouldn't be able to retrofit existing devices. Most people dont like tinkering with things. A universal device small enough to fit in your pocket, with a nice little display or a usb connector to download data to a laptop and configure rules, is more desirable imo.
I work at SoundHound where we've been worried about these issues. (I'm going to plug our recent work...) Our new approach is to do natural language understanding in real-time instead of at the utterance (turn) taking level. That way we can give the user constant feedback in real-time. In the case of a screen that means the user sees right away that they are understood, and if not, a better hint of what went wrong. For example a likely mistake is an ASR mistranscription for a word or two.
We still need to prove this is a better paradigm for VoiceAI in products that people can try for themselves, and are working towards that goal. I hope that voice interfaces that were clunky with turn-taking will finally be more naturally usable with real-time NLU.
https://www.youtube.com/watch?v=5WLYH1qHfq8