Wireless Speech Recognition ..

Speech recognition is now primarily wireless; We've migrated fast, to universal wireless access-communcation devices.

Often, the speech recognition is remote based - And the better signal we send it, the better it performs.

Here, we hope you'll find ideas, technology or projects using hands free and/or mobile devices to make wireless speech recognition a rewarding and useful universal tool!

Tuesday, June 10, 2008

The most remote speech-driven system yet?

 
 On Bantayan Island in the Philippines, to prove its new speech-driven services work, the Government Service Insurance System (GSIS) unveiled a remote voice biometrics (a/k/a "voice recognition") system for its members on a picturesque island located in the northernmost tip of Cebu.

 "If it will work in Bantayan Island, it will work anywhere," said GSIS President and General Manager Winston Garcia, in a briefing here. Garcia said the service is being launched as part of its modernization program. "This system removes the need for them to go to a GSIS office to renew their pension status. They can now do so remotely, via a phone call," noted Garcia, explaining how the system is intended to reduce large inbound call queues into GSIS offices.

 The service was beta-tested in the United States prior to the local (Cebu) launch; E.G. Washington, New York, Chicago, Los Angeles, San Francisco, and Hawaii.

 The speech driven "GSIS Voice Activated Processing System" (G-VAPS) enables its 1.2 million members to transact with the GSIS using their unique voice as their "electronic signature,". The system is secure, and is be able to detect "tape-recorded" as opposed to live speech. Active members and pensioners call a US-based GSIS Teleservice toll free number to use the service.

 Currently, it allows members to apply for and process loans. It can also be initially used by pensioners to renew their status to active members; Active members can then go to GSIS office to record their voice after presenting proper identities.

 "This is voluntary but it provides them convenience," said Garcia when asked if he expects all members to avail of this new service. He also noted GSIS Wireless Automated Processing System kiosks remain another option to apply for loans.

 Garcia further noted the system was customized by its own information technology department.

 

Labels: , , ,

Saturday, June 07, 2008

 
 Cited primarily from an article at http://computer.getmash.net/;

 Speech recognition has long languished in the no-man’s land between sci-fi fantasy (”Computer, engage warp drive!”) and everyday usage reality.

 But that’s changing fast, as advances in computing power, artificial intelligence, powerful API's & newly available WSR Macros, make speech recognition the next powerful step for everyday use by "non-geek" users, user-interface design and now electronic voice-based security.

 As to voice-based security: A whole host of highly advanced speech technologies, including emotion and lie detection, are moving from the lab to the marketplace.

 This not a new technology,” says Daniel Hong, an analyst at Datamonitor who specializes in speech technology. “But it took a long time for Moore’s Law to make it viable.”

 Mr. Hong estimates at the speech technology market is worth more than $2 billion, with plenty of growth in embedded and network apps.

 And it’s about time. Speech recognition's technology has been around since the 1950s, but only recently have computer processors and accompanying artificial intelligence become powerful enough to handle the complex algorithms required to recognize our speech, both local & remote, improve our lives & productivity, and open our eyes to the long tail of speech recognition fields.

 A few examples:
 There are already several capable voice-controlled technologies on the market. You can issue spoken commands to devices like Motorola’s Mobile TV DH01n, a mobile TV with navigation capabilities, and a host of telematics GPS devices. Microsoft recently announced a deal to slip voice-activation software into cars manufactured by Hyundai and Kia, and its TellMe division is investigating voice-recognition applications for the iPhone. And Indesit, Europe’s second-largest home appliances manufacturer, just introduced the world’s first voice-controlled oven.

 Yet as promising as this year’s crop of specch-controlled devices are, they’re just the beginning.

 Speech technology comes in several flavors, including the speech recognition that drives voice-activated mobile devices; network systems that power IVR's using automated speech recognition, the unequalled desktop Vista Speech recognition, now with available macros {which we use to post & write articles) and the long-standing the standard in the Healthcare industry, the highly impressive network-based Philips SpeechMagic systems.

 Voice biometrics (the true technical description of the often mis-used phrase "voice recognition") is a particularly hot area. Every individual has a unique voice print that is determined by the physical characteristics of his or her vocal tract. By analyzing speech samples for telltale acoustic features, voice biometrics can verify a speaker’s identity either in person or over the phone, without the specialized hardware required for fingerprint or retinal scanning.

 The technology can also have unanticipated consequences. When the Australian social services agency Centrelink began using voice biometrics to authenticate users of its automated phone system, the software started to identify welfare fraudsters who were claiming multiple benefits — something a simple password system could never do.

 The Federal Financial Institutions Examination Council has issued guidance requiring stronger security than simple ID and password combinations, which is expected to drive widespread adoption of voice verification by U.S. financial institutions in coming years. Ameritrade, Volkswagen and European banking giant ABN AMRO all employ voice-authentication systems already.

 Advanced voice-recognition systems that can tell if a speaker is agitated, anxious or lying are also in the pipeline.

 Computer scientists (e.g. at Carnegie Mellon) have already developed software that can identify emotional states and even truthfulness by analyzing acoustic features like pitch and intensity, and lexical ones like the use of contractions and particular parts of speech. And they are honing their algorithms using the massive amounts of real-world speech data collected by call centers and free 411 speech-driven services such as the extremely popular Goog411.

 A reliable, speech-based lie detector would be a boon to law enforcement and the military. But broader emotion detection could be useful as well. Our host company which developed the now-standard Law Enforcement "Mobile Prosecutor" is presently experimenting with embedding it with voice-stress analysis.

 In another example, a virtual call center agent that could sense a customer’s mounting frustration and route her to a live agent would save time, money and customer loyalty.

 “It’s not quite ready, but it’s coming pretty soon,” says James Larson, an independent speech application consultant who co-chairs the W3C Voice Browser Working Group.

 Companies like Autonomy eTalk claim to have functioning anger and frustration detection systems already, but experts are skeptical. According to Julia Hirschberg, a computer scientist at Columbia University, “The systems in place are typically not ones that have been scientifically tested.”

 According to Hirschberg, lab-grade systems are currently able to detect anger with accuracy rates in “the mid-70s to the low 80s.”

 They are even better at detecting uncertainty, which could be helpful in automated training contexts. (Imagine a computer-based tutorial that was sufficiently savvy to drill you in areas you seemed unsure of.)

 Lie detection via voice stress & syntax-pattern deviation analysis is a tougher nut to crack, but progress is being made.

 In a study funded by the National Science Foundation and the Department of Homeland Security, Hirschberg and several colleagues used software tools developed by SRI to scan statements that were known to be either true or false. Scanning for 250 different acoustic and lexical cues, “We were getting accuracy maybe around the mid- to upper-60s,” she says.

 That may not sound so hot, but it’s a lot better than the commercial speech-based lie detection systems currently on the market. According to independent researchers, such “voice stress analysis” systems are no more reliable than a coin-toss.

 It may be awhile before industrial-strength emotion and lie detection come to a call center near you. But make no mistake: They are just around the proverbial corner. And they will be preceded by a mounting tide of gadgets that you can talk to, argue with and intelligently discuss topics with.

 Don’t be surprised if, some day soon, your Bluetooth headset tells you to calm down. Or informs you that your last caller was lying through his teeth.

 Now that Windows Speech Recognition Macros for Windows Vista™ are in feverish development, both in-house (Microsoft Speech Components Group [ listen_+at+_microsoft.com ], and the beta group inside the Microsoft Speech Yahoo Technical Group) - desktop speech recognition is advancing daily by leaps and bounds, literally.

 Powerful WSR macros that can, for example:

  • Open e-mail messages from a specific (non-Inbox) account with TO: / CC: / BCC: and Subject: fields already completed;

  • Macros that can move large blocks of extant text in and out of specific locations inside different applications;

  • Navigate & move items in and out of various folders inside Vista Explorer;

  • Spoken database lookups

  • are already evolving and being used & improved daily. It will not be long before speech recognition becomes "what we just use" for most of our daily work & living activities..

    (A detailed post on the powerful new WSR Macro Tool & evolving macros is coming soon; We're gathering data, useful macros and research to be sure it is both interesting & useful to all types of speech recognition users)

     

    Labels: , , , , , , , , ,

    Wednesday, May 07, 2008

    Tweet by phone.. TwitterFone!

     
     Yes, it's true.. Now you can Tweet by phone.
     

     · Click to view TwitterFone's website · 
     We didn't get an invite yet, but we'll be sure to post more when we do..

     

    Labels: , ,

    Tuesday, May 06, 2008

    SPEAK: Speech-enabled auto attendant for SMB's!

     
     Active Voice has rolled out a turnkey speech enabled auto-attendant that's priced and aimed at small to mid-sized business markets, named "SPEAK".
     

     · Click to visit the website · 
     Via the press release & web page:
    "The engineering framework of SPEAK is designed to be a sophisticated turnkey solution that is considerably more affordable, while being easy to implement, easy to deploy and easy to use. Thus, SPEAK provides SMBs a practical solution for speech attendant, corporate directory as well as mobility access."

     Some productivity features the maker notes:

  • Upgrades customer service while reducing “zero-out” calls

  • Eliminates the push button frustration as well as the need to search for numbers

  • Provides total 24/7/365 self-service information such as employee and department directories, company information, driving directions, business hours, etc.

  • Performs dynamic call routing to ensure all incoming callers are treated in a personalized and professional manner

  • Frees front desk support to deal with important face-to-face matters

  • Keeps mobile workforces connected efficiently and provides hands-free access for safety

  • Reduces the need for maintaining and printing company directories

  • Secures access to your telephony resources across all networks


  •  Per Eyal Inbar, General Manager of Marketing and Product Development:
     "Active Voice SPEAK leverages LumenVox's cutting-edge Speech Engine and Digium's open standard hardware, enabling SPEAK to be offered at a price point that is unbeatable in the market and at a level of sophistication that is yet to be experienced by SMBs. The bottom line of SPEAK is simplicity, affordability and ease of use."

     

    Labels: , , , , , ,

    Monday, April 14, 2008

    DoCoMo's new handset; built-in speech recognition

     
     Japan's NTT DoCoMo Inc (NYSE: DCM) is releasing the FOMA "Raku-Raku Phone Premium" F884i mobile phone today, with built-in speech recognition for remote transcription of email text.
     
     · Click to read about DoCoMo's technolgies · 
     The handset, made by Fujitsu Ltd, contains new & proprietary technology to enter e-mail text using remote speech recognition. If the "voice input" button in the e-mail editing display is pressed, software that extracts the characteristics of the user's speech (and performs Analog-to-Digital conversion) will start, and access the DoCoMo's i-mode site.

     When users say what they want transcribed into an email, the in-box software sends the dictation to the i-mode server. There, speech recognition software manufactured by Advanced Media Inc outputs transcribed text. The F884i receives the transcribed text, displays it in the email's text display interface.

     As best we can tell, without a direct response from DoCoMo or Advanced Media - This appears to be DSR (Distributed Speech Recognition) which we've blogged about in the past; and we are tremendous fans of DSR as a global answer to near-perfect mobile speech recognition.

     We've also emailed David Pearce, founder and Chief Developer of this emerging technology to see if he has any information on whether this may, in fact be DSR..
      Check back later for details!

     

    Labels: , , , ,

    Monday, April 07, 2008

    SendChat - new universal speech-to-text messaging!

     
     TMCnet reports today that SR Virtual has developed SendChat, a "state-of-the-art, voice-to-text SMS" that mobile users can download to any mobile phone, regardless of which wireless carrier they are using. SR Virtual notes that SendChat will install as easily as a new ringtone, (welcome news to those of us who don't fancy tiresome fiddling with mobile phones..)

     TMC further reports that SendChat uses smart technology that continues to learn the user’s vernacular and diction; thereby increasing its transcription accuracy with every use.
      Very Cool.

     SendChat is server-based which is why it will work with phones on any wireless network, and also spares users the tedious chore of updating new versions as well.

     We don't have a download address, yet, but we will contact them shortly and see if we can glean further info; this promises to be a godsend to those of us who do text, but hate the "thumb typing" that goes along with it!

     

    Labels: , , ,

    Wednesday, April 02, 2008

    Speech recognition gets increased public confidence

     

    Callcentres.net has released a study in Australia with some quite welcome revelations..
    The high points --


  • Overall, customers were significantly more satisfied with their speech recognition experience in 2007 than they were in 2005.


  • The research also shows that speech recognition is the preferred self-service interface; 66 percent of survey respondents preferred speech across the internet, & 59 percent preferred speech recognition over touch-tone (DTMF-driven) IVRs, when using the telephone.


  •  Dr. Catriona Wallace, a director of callcentres.net said: "Confidence has emerged as a key factor influencing satisfaction with speech recognition. The research showed that frequent users of speech technology have a statistically significantly higher level of satisfaction with the experience than new or inexperienced users.

    "It's also interesting to note that while men are more willing to try speech recognition, it's women who are more likely to become real advocates of the experience once they've used it," Dr. Wallace explained.

     

    Labels: , , , ,

    Speech-enabled Yahoo oneSearch is released

     
    Today at CITA Wireless 2008, Yahoo announced version 2.0 of its oneSearch mobile search application now includng Voice-Enabled Search.

    A combination of predictive query entry and speech-recognition provided by vlingo, oneSearch is only available now for select Blackberry devices including the 8800 series, Curve, and Pearl, but Yahoo stressed that other handset support would follow shortly.


    "Consumers can search for anything, including flight numbers, locations, Web site names, local restaurants, and more, by simply speaking," a release from Yahoo detailed. The voice-activation software is now available for download on a number of RIM's BlackBerry devices, and Yahoo has said that over the next few months it will be compatible with more handsets.

    We're thrilled to see speech recognition emerging as a driving force for mobile search. We'd hope that Yahoo! and many others will begin to use vlingo's (the only close 2nd to Microsoft's voice searching on Windows Mobile/Smartphones) technology to expand and tweak mobile-centric searches.

    Even though hurdles still exist for mobile-centric speech searching (noise canceling, poor voice signal quality) it’s begun to receive some serious integration as of late. Other outfits like Free411, Goog411 and Ask.com are using speech recognition technology; and the new ChaCha has a pretty robust speech recognition built into its new mobile-centric searching.


    Labels: , ,

    Thursday, March 13, 2008

    The new Mobivox "Say it and Save it" Feature

     
     The calling service Mobivox that offers free international calling began offering a new feature title "Create a Contact" today.

     · Click here to view the Mobivox features · 
     This feature enables MOBIVOX users to add new contacts, This feature enables MOBIVOX users to add new contacts, by simply speaking a name to the service's recognition application, named "Voxgirl" !

    Via the Mobivox website:
    "Dial your local access number, and when VoxGirl answers, simply say 'Create a Contact', or press number 6 from a touchtone phone. You will then be asked to record the name of the contact you would like to add, and enter the telephone number using your telephone's keypad."

     Cool!

     Mobivox's Voxgirl incorporates advanced speech and noise profiling, trained noise model, "voice tag" and vocabulary size technology, and there is almost unlimited storage space for voice dial contacts.

     Nitzan Shaer, the company's COO comments:
     “This new feature will dramatically improve the speech recognition accuracy, and is especially useful for those with multi-cultural backgrounds and non-English contact names. These are precisely the people who use MOBIVOX the most to keep in touch with their home country and loved ones far away.”

     MOBIVOX also uses artificial intelligence to identify a frequently dialed number that doesn't appear in the user's the contact list, and accoordingly then prompt the user to create a contact for that number; A handy feature for users that don’t have their address books readily accessible.

     Also very cool.

     That's not all Mobivox does that's impressive.. The service will sync your Skype contacts and make them available to call using the remote voice access system. You can also query the service to see which of your Skype buddies are online.

     MOBIVOX also offers instant conferencing (allowing users to spontaneously add callers to an existing conversation), group calling, and the ability to transfer a call from one phone to another, as well.

     Mobivox's business model primarily hinges on revenue from international calling; Users purchase chunks of up to $100 international mobile-to-landline credits at a time; and Skype users can call their Skype contacts without having to buy credit from Skype, either!

     There aren't "extra" charges for Mobivox's service other than minutes you use up on your mobile (or domestic-calling) plan, and since Mobivox gives you a local number, you can avoid legacy long-distance per-call charges on landlines, where applicable.

     

    Labels: , , , , , , , ,

    Wednesday, March 05, 2008

    Over-the-phone note & task transcription

     
     Via Angel.com:
    "SalesByFone from Angel.com makes it possible to access, update, and manage accounts, contacts and leads directly in salesforce.com through voice commands over the phone. With a simple phone call, you can record your impressions about a just-completed meeting, set a follow-up task, or connect directly to a contact".

     Per PRWEB - March 5, 2008:
    Angel.com, the leading provider of hosted, on-demand call center applications, has partnered with SimulScribe, the largest provider of voicemail-to-text services and visual voicemail applications, to integrate speech-to-text functionality with Angel.com products and services. The first offering using speech-to-text functionality is Angel.com's new Salesbyfone application.

     SimulScribe's technology allows Salesbyfone users to transcribe meeting notes and other details over the phone and see notes appear, within seconds, in Salesforce.com contact records. Users can also automatically dial and send an e-mail to a contact simply by speaking it over the phone. These functions occur in near-real time, allowing users to quickly act on or respond to critical business situations as they happen.

     Salesbyfone is the latest in Angel.com's suite of IVR (Interactive Voice Response) integration applications for Salesforce.com. Salesbyfone provides phone-based access to Salesforce.com accounts, empowering sales executives and other users to access, update, and manage key prospect information directly in Salesforce.com through voice commands.

     

    Labels: , , , , ,

    Sunday, February 10, 2008

     
     The Hague’s Bronovo hospital in The Netherlands has converted to 100% speech recognition for all departments - it has improved service quality, and profits.

     Says Dr. Pieter Lambregts, who oversaw the speech recognition for the Neurology group in the hospital.."The overall efficiency gain and time saving allowed us to increase our department’s activities by 20% without having to hire additional secretaries.”

     Speech recognition is widely used in the Health Care industries in the Netherlands; SpeechMagic from Philips is the platform for over 80% of the speech recognition systems.

     Wilfred Reinhard, Bronovo’s IT Manager noted “We wanted to realize the country’s first all-speech recognition hospital because we knew how critical the availability of clinical information is for the delivery of care.”.

    The entire case study from G2 Speech is available here.

     

    Labels: , , , ,

    Wednesday, January 09, 2008

    Navigation, solely w/ remote speech recognition!

     
    NavStar Technologies, in partnership with Navteq, a provider of digital maps for vehicle navigation, has released "Voice Navigator" - an entirely speech recognition driven navigation device accessing only a remote server!

    True, wireless remote speech recognition... via NavStar:
    "No PC's, CD's, DVD's, tiny screens, difficult interfaces, or piles of 'stuff' to learn. Simply plug it in and start using it."


    NavStar Navigation Device


    According to NavStar, there is even "planning intelligence" built in:
    Users can plan their trip ahead of time, store it and then access the planned route with the push of a button.

    The device interfaces with user's mobile devices, using their wireless broadband service. Navteq's "Points of Interest" database includes over six million listings including airports, restaurants, hotels, banks, gas stations and other popular destinations, or, subscribers can state the exact address desired.

    Our techie hats are off to this wonderful advance in true wireless remote speech recognition!
     

    Labels: , ,

    Thursday, December 13, 2007

    Speech recognition's accuracy better than human transcription..

     
     Welcome news for speech recognition proponents!

     HealthImaging.com (a site for Healthcare IT professionals) posted a news article yesterday, December 13, about a presentation at the Radiological Society of North America (RSNA)'s annual meeting last month, documenting that the reports which were manually transcribed by humans, showed higher error rates than the reports that were transcribed through speech recognition!

     John Floyd, MD a partner in the 24-member Radiology Consultants of Iowa (RCI), reported “The rate for significant errors, requiring the preparation of an addendum, was 0.6 percent for speech recognition and 2 percent for traditional transcription.”

     Floyd also noted speech recognition significantly increased his firm's efficiency: "Separate data for this practice indicated that average turn-around time for traditional transcription was greater than 24 hours while that for speech recognition was less than one hour.."

     Dr. Floyd further confirmed that the accuracy rate for speech recognition reported by his group was independently verified by 3rd party analysis, conducted at one of the hospitals his partnership services.

     In an On10Net blog post, Bill Crounse MD, Healthcare Industry Director for Microsoft Corporation predicted earlier this year that speech recognition would open up new vistas in the healthcare industry..

     We're pleased to see his predictions coming true!
     

    Labels: , , ,

    Thursday, November 15, 2007

     
    An excellent speech recognition blog, (appropriately titled "Speech Recognition" :-) has posted a really interesting article about an intriguing remote-based wireless speech recognition system for Emergency Rooms, and also links to a terrific case study on the Crescendo site.

    The case study's diagram outlines a system that uses a remote central voice & data server_+_Speech Recognition (SpeechMagic) Server that is accessed via Pocket PC's, and allows ER physicians to "walk and talk" while maintaining access and the physicians are provided with up-to-date patient data on their PDAs - any place, any time.. driven via speech commands.

    It's fascinating reading, really worth checking out.
    .

    Labels: , , ,

    Wednesday, July 05, 2006

     
    We have found, and have been using a remarkable advance in true remote "Call your own PC" speech recognition found at Adondo's PAL website. This rather innovative and solidly performing application not only promises to revolutionize remote speech recognition, but the new "Personal Audio Link" allows you to you call your own PC or workstation (yes, it safely works through most corporate firewalls) and manage e-mail, including creating and sending them, entirely by speech alone and/or uses voice commands to get real-time traffic, weather, and stock quotes. It also reads you your favorite news sites, jumps back and forth between topics by speech alone, or reads you blogs and plays podcasts & audio from websites like your local radio stations.

    Sitting in front of your laptop or PC and using PAl to familiarize yourself with what it'll do (and it's functionality takes some getting used to; it does a lot more than just what meets the eye at first) is quite the experience. As you say voice commands, the application whisks you from window to window, each containing the website (where applicable) you're accessing... cool.

    It's use of the local Windows Media Player ensures you'll handle every type if popular music/audio format, quickly and cleanly with superb audio clarity. What's to us, always a pleasant surprise is hear some of the clever quips it speaks at opportune times, the program is extremely innovative and you get a real sense of "sensible" Artificial Intelligence as you use it more and more.

    Our Techie hats are off to the authors for this highly advanced mode of useful, hands free speech recognition that one quickly feels "naked" without . I for one quickly became used to driving down the road chatting with PAL getting real-time stock quotes and listening to podcasts without *touching anything. No more "thumb driving" to manage e-mails.

    The author of this particular post was stuck in a grocery store checkout line, and our Personal Audio Link beta called him and reminded him of an appointment. When he realized he had forgotten to bring this little-used contact's phone number, he then voice-dialed his PAL, asked it for the contact's work number, and never took his hands off what he was doing. Try that little trick with a Blackberry or a PDA.

    This application promises to be most certainly the next "crackberry" in the wireless speech recognition world. We could go on and on about how cool it is, but go to the website, try it and post a comment..!!

    Labels: ,

    eMicrophones

    Promote Your Page Too