Showing posts with label IxID. Show all posts
Showing posts with label IxID. Show all posts

24.2.15

Speech Synthesis



http://wolfpaulus.com/tag/speech-synthesis/
Yesterday was a big or maybe even a huge day, for voice user interfaces.
Microsoft introduced us to Cortana and Amazon introduced the Fire TV box, which includes a remote control, supporting Voice input. Considering that both are not 1st to market (SIRI, ComCast’s x1-xfinity), their entry further validates VUIs.
FireTV’s voice recognition seems so good that when Jon Fortt was demonstration it today on live television (CNBC), he dared to ask it for “Pawn Stars” http://video.cnbc.com/gallery/?video=3000264102
Wedbush analyst Shyam Patil wrote that Nuance Communications was likely powering Amazon.com’s Fire TV voice search, while reiterating a neutral rating and $15.00 price target on Nuance. (Nuance traded today for $17.59)
The consolidation regarding Speech Recognition and Synthesis seems to continue. Apple has acquired speech recognition pioneer Novauris last year, but this had not been announced until today. One of the biggest differentiators about Novauris in terms of the competitive landscape, is that they operated in both the embedded (i.e. on-device, like OpenEars, PocketSphinx) and server space (like LumenVox, Nuance), and they also owned the core engine.
“NovaSearch doesn’t carry out recognition at the word or sequence-of-words level, but rather identifies complete phrases from start to finish by matching them against a potentially huge inventory of possible utterances. This enables it to assemble information about what has been spoken over utterances of virtually any length and take near-optimal decisions.”

9.2.15


http://www.nime.org/proceedings/2007/nime2007_285.pdf
http://andrewcook.co.uk/index.html

http://www.trimpinmovie.com/ 


The performer’s specialized
experience of the physical world allows them to act on the
instrument in a way that will produce the sounds they desire, and
the audience’s general experience of the physical world, their
embodied knowledge and conceptual models [10] built over a
lifetime’s experience, means that they can draw a meaningful
correlation between what they see and the sounds that are
produced.