Showing posts with label speech recognition. Show all posts
Showing posts with label speech recognition. Show all posts

Tuesday, May 27, 2014

Apple, Microsoft, Google -- how about a Kindle Killer?

I have a Kindle and I like it, but I don't love it. I like the form factor and the e-paper display and the touch interface is OK, but I really wish it used voice input in the user interface and was well integrated with my laptop.

Reading is an active process -- kind of a dialog with the author -- and I create a lot of associations and marginal notes. I want an e-reader that lets me record those in a database while reading. Database entries would include a note, timestamp, tags (global to me, local to the document -- perhaps even global to the world), the document it was associated with, the selected passage or figure it was associated with, etc. The links would be hot.

I'd like to be able to search this database in a variety of ways, for example, finding entries with a given tag or combination of tags, free-text search or those associated with a given document and display the results in a variety of ways, for example in a table showing only the first N characters of a marginal note and the passage it was associated with.

Voice recognition -- for commands and text entry -- should be a first-class input modality. If I have selected a passage for annotation, a single voice command should initiate speech-to-text mode so I could dictate my note and tags.

The e-reader and my laptop should be integrated. I should be able to copy the annotation database to my laptop with a single command and search and display reports in the same way as on the e-reader. I should also have the option of updating the e-reader database from my laptop, uploading it to the Internet or exporting it in various formats, for example, docx or xlsx.

I've outlined a few features I'd like in an e-reader -- what would you like to see?

The business case -- low-hanging fruit

When the Kindle came out, Steve Jobs dismissed it saying “It doesn’t matter how good or bad the product is, the fact is that people don’t read anymore -- forty percent of the people in the U.S. read one book or less last year. The whole conception is flawed at the top because people don’t read anymore.”

Amazon does not release official sales figures, but Forbes estimated that between 2007, when the Kindle was released, and the end of 2013, roughly 43.7 million Kindles were sold. Assuming a 3-year replacement cycle, about 30 million Kindle e-readers are currently in use -- for reading and purchasing content from Amazon. Regardless of Jobs' opinion, there is a sizeable market.

Apple is a device manufacturer par excellence, Microsoft has elevated "devices" to its tag line and Google sells some devices, so they all have varying degrees of manufacturing experience.

Speech recognition is the key technology that I want added to an e-reader and Apple, Microsoft and Google have world-class speech recognition products and research projects.

Amazon may be taking a loss on each Kindle and making it up on content sales, but Apple, Microsoft and Google have online stores too and could be selling the same content as Amazon.

The Kindle reminds me of the Apple Newton -- a very early "personal digital assistant" with a stylus and handwritting recognition as its primary input modality. The Newton failed because the handwriting recognition was slow and inaccurate -- like typing on a Kindle's virtual keyboard -- and it was a stand-alone device, barely integrated with Apple or Windows computers. (I've still got a Newton -- in mint condition because I only used it to look cool). Apple learned from the failure of the Newton -- the iPod was one component in a system along with the iTunes store and software.

It seems to me that a Kindle-killing e-reader would be low-hanging fruit for Apple, Google or Microsoft. If any one of them builds it, I'll buy it. (Think of the competition if they each built one)!

-----
Update 6/5/2014

There is a long (320 comments) discussion of this post on Slashdot.

Sunday, July 15, 2012

Dragon Systems founders suing Goldman Sachs $1 billion

The Bakers, 1990, NYTimes
Dragon Systems was the voice recognition company in the 1980s and 1990s -- they brought speech recognition from the research lab to the PC. (Founders James and Janet Baker, shown here, came from MIT). I recall using and reviewing DragonDictate back in the day.

The Bakers sold Dragon to Lernout & Hauspie for $580 million in stock, which sounds good at first, but L & H turned out to be a fraud and collapsed, leaving the Bakers with nothing.

The Bakers are now suing Goldman Sachs for one billion dollars.

The sale to L&H was brokered by Goldman Sachs, which collected millions of dollars in fees, and, it turns out that Goldman Sachs had previously considered investing in L & H, but had walked away after some digging into the company.

Does that sound familiar? A Wall Street firm making a commission by selling something that they themselves would not buy? This story grosses me out.

Thursday, December 01, 2011

Is speech recognition finally going to catch on with Siri?

Everyone agrees that some day we will be talking to our computers -- dictating memos, asking questions, giving commands, etc. The catch is that "some day" seems to be always five or ten years in the future.

People have been working on speech recogntion for a long time. My first exposure was a demonstration of Shoebox, a calculator with speech input, at the IBM pavillion at the 1964 World Fair. In spite of years of research and hacking, speech recognition has remained niche technology.

Have we finally seen the start of practical, ubiquitous speech recognition with Apple's Siri? Maybe.

Siri has a lot of infrastructure support that earlier speech recognition systems lacked. It sends the speech back to a server for recognition and that server has assimilated clues from massive amounts of data on speech patterns. Once recognized, it relies on other services for search and to look for answers to questions. If you ask how far it is from Los Angeles to New York, it will go to WolframAlpha for the answer. Ask it where to get Indian food in your neighborhood and it will go to Yelp. (What will Apple do if you ask where to find bomb-making instructions or dirty pictures)?

Google seems to have the recognition part down, but may be playing catch up with input parsing and answer retrieval.

In spite of Apple's secrecy, Siri has attracted a hobbyist following. Check out this video of a hobbyist using Siri to control lights and other things in a room.



The developer of that app had to jump through hoops using SiriProxy to get it to work. Here's hoping Apple provides tools to encourage this sort of thing -- that might be what it takes to finally get speech recognition off the ground.