Wednesday, July 16, 2008

Google Video Search via Speech Recognition




Finally a hint on the expected Google move to the speech recognition arena.

Google announced at the Official Google Blog, the availability of a new video search capability based on speech recognition.

It was release as a gadget you can embed on your iGoogle homepage and is a good preview of things to come.
The gadget only searches videos uploaded to YouTube's Politicians channels, which include videos from Senator Obama's and Senator McCain's campaigns, as well as those from dozens of other candidates and politicians. It usually takes less than a few hours for a video to appear in the index after it has been published on YouTube.

So apart from congratulations to the google team who are exposed to the public for the first time, how are they compared to other speech recognition engines aimed for broadcast quality? The google team refer to their precision: "While some of the transcript snippets you see may not be 100% accurate, we hope that you'll find the product useful for most purposes." While I do not understand what are the purposes for just searching within the YouTube political channel, people should be aware of much more mature solutions developed in the past years. From the pioneering work of BBN and IBM to the existing online solutions like everyzing, tveyes, blinx, snipp.tv by NSC and more. Based on the perceived quality, the google team has a long way to go in order to get to the first league and to be able to analyze data which is not at broadcast quality. The good news as users is that the YouTube data is easier to process relative to telephony calls speech recognition performed widely today at contact centers by companies like Verint, Nice, Autonomy, Utopy, Nexidia, CallMiner and other players.

Friday, July 11, 2008

SuperHuman Speech Recognition

Last week the the Speech Technologies Group at the IBM Haifa Research Lab (HRL) coordinated a full-day seminar on Speech Technologies. The seminar was a great success with more than 100 participants.

The Keynote presentation at the IBM Speech Technologies Seminar 2008 was: "Superhuman Speech Recognition: Technology Challenges and Market Adoption" by Dr. David Nahamoo, IBM Fellow, Speech CTO and Business Strategist, IBM Watson Research Center. You can view the presentation below. More presentations will be posted soon.




Read this document on Scribd: SuperHuman Speech Recognition Jul 2 2008

Saturday, July 5, 2008

11 Indian languages available from Nuanace


Nuance just extended their Indian languages support. In addition to Hindi and Indian English, they support also: Marathi, Malayalam, Tamil, Kannada, Telegu, Bengali, Gujrati, Oriya, and Punjabi.
I wonder what will be an automatic language identification results when trying to discriminate between these languages automaticallyl.




Nuance Communications Launches 9 New Indian Languages for Speech Recognition



By VARIndia Correspondent

Nuance Communications has released 9 new Indian languages for speech recognition in the contact centre...

Tuesday, June 17, 2008

Nuance-Vlingo: If you are not sued, you do not exi(s)t

Nuance is not only active to push speech based search to the iphone competing with the vlingo/Yahoo offering. Apparently, Nuance just filed a lawsuit against Vlingo for infringing one of their 1000 patents. Nuance and its predecessor (ScanSoft) has a long history of lawsuits which were in sync with their M&A and business strategy. In some cases, when competing on a large account or negotiating a good M&A price, Nuance used the lawsuit mechanism to get a better deal.

Just few examples from the past:


ScanSoft, ART could settle lawsuit with acquisition

"Peabody-based ScanSoft Inc. may settle a lawsuit by acquiring defendant ART Advanced Recognition Technologies Inc., according to an Israeli newspaper report.

The deal could amount to tens of millions of dollars, according to Globes Online, which attributed the report to a Hebrew newspaper called Yediot Ahronot."


ScanSoft files suit against Voice Signal

"Voice Signal Technologies Inc., a Woburn start-up that sells speech-recognition systems used in wireless phones, has been hit with a patent-infringement lawsuit by ScanSoft Inc., after Voice Signal refused to accept what executives yesterday called a lowball takeover bid from ScanSoft."

There is also the famous TellMe case.

I think it is a great feedback to Vlingo whose management includes are ex-Nuance employees. If you are not sued in this industry, you do not exi(s)t.

Sunday, June 15, 2008

iPhone speech recognition

The iPhone is attracting many developers wishing to add its next cool application. Nuance recently introduced vsearch - a voice search application. Similar to Vlingo's recent application on the blackberry.


The official video from Nuance







I believe the following video demonstrate better the hands free aspect of search.




Is it a gimmick or will people actually use it? What happens when there are speech recognition mistakes? Are there such mistakes? I like much more the unofficial video demonstrating the yahoo/vlingo voice search.



If someone has some statistics on voice base search, I will be glad to receive it.

Tuesday, May 27, 2008

Speech analytics in the contact center - what's driving adoption rates?

Ok, so many analysts have been talking about the growth rate of speech analytics continuing to accelerate through 2008 and beyond. While I will concede this is a reasonable prediction for this technology, the reality is that, while more and more companies are budgeting for, evaluating and even purchasing these technologies, the adoption of these solutions into core customer business practices, and more importantly the quantifiable business benefits delivered tell a story of fuzzy results and an ever-denser fog through which to see potential measurable results.

Over a series of posts on this topic, its my goal to offer insight from experiences working with over 100 clients, my view of effective and not so effective approaches and hopefully some practical remedies that can be applied to your business today.

I’m not going to spend time here revisiting the history of the world of speech analysis or take some dive into the weeds on the technology. If the title of this article resonated with you, I assume you’ve done your homework and have probably struggled with some of these same issues during a recent project. Or, you’ve postponed such a project because cracking the code on this technology has been too elusive to provide a level of confidence in moving forward. No. My mostly benevolent, and maybe a tiny bit self-serving, objective here is to share my experiences and opinions I’ve refined over the past five years working with these solutions.

If you want more details on the technology, there are plenty of sources out in cyberspace for all the facts and figurers, bits and bites you want; if you’re into that sort of thing. Do you’re research and then pick this back up and read through it before you make your next move.

Some of the headings under which the adoption rate of these solutions fall include:

-Value realized by early adopters
-Once bitten, twice shy
-Overwhelmed quality management functions
-The hype cycle - oversold rudimentary capabilities
-The “Toy in the Happy Meal® Syndrome”
-Fuzzy ROI
-Best Practices
-Managing organizational change

As this series progresses, we'll tackle each of these, individually and as they potentially influence eachother, in combination. I hope this series of posts stimulates others to contribute their experiences. I look forward to our journey.

Friday, May 23, 2008

Speech Technologies Seminar 2008

The Speech Technologies Group at the IBM Haifa Research Lab (HRL) invites speech professionals to a full-day seminar on Speech Technologies, to be held on Wednesday, July 2, 2008.

This full-day seminar provides a forum for the research and development communities from both academia and industry to share their work, exchange ideas, and discuss issues, problems, and work-in-progress, as well as future research directions and trends. The seminar agenda will be posted at a later date. It will include frontal presentations and a poster session.

The seminar will take place at the HRL site on the Haifa University campus, in the auditorium (room L100). Lunch and light refreshments will be served. Participation is free.

See http://www.haifa.il.ibm.com/Workshops/speech2008/index.shtml for detail.

Program
09:00 Registration

09:30 Opening Remarks
Oded Cohn, Director, IBM Haifa Research Lab

09:45 Challenges of Speech Solutions in Call Centers
Nava Shaked, Manager, CRM & Call Center, IBM Israel

10:15 Actionable Intelligence via Speech Analytics
Ofer Shochet, Senior VP, VERINT

10:45 Discriminative Keyword Spotting
Joseph Keshet, IDIAP

11:15 Break

11:30 Recent Advances in Speech Dereverberation
Emanuel Habets, Bar-Ilan University & Technion

12:00 On Improving the Quality of Small Footprint Concatenated Text-to-Speech Synthesis Systems
David Malah, Head of Signal Processing Lab, Technion
12:30 Keynote. Superhuman Speech Recognition: Technology Challenges and Market Adoption
David Nahamoo, Speech CTO and Business Strategist, IBM Watson Research Center

13:30 Lunch

14:30 Using Speech Processing Technologies in Audio Search Applications
Ido Itzhaki, Director, Business Development, NSC

15:00 Intra-class Variability Modeling for Speech Processing
Hagai Aronowitz, IBM Haifa Research Lab

15:30 Retrieving Spoken Information by Combining Multiple Speech Transcription Methods
Jonathan Mamou, IBM Haifa Research Lab

16:00 Poster Session & Refreshments
Bookmark and Share