Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Baen also has a problem with OCR'd text. I downloaded a Greg Bear book from them a few weeks ago and it was clearly OCR'd without a second look towards quality. It had so many typos and spelling errors that it was practically unreadable. I returned it for a refund.

To be fair Amazon also has this problem. It's hard being an ebook consumer when 90% of books either have DRM or are unreadably OCR'd.



I've seen that with a couple books from Smashwords, but I'd expect better from Baen. I think I've only bought books by their regular authors (Weber, Flint, Drake) and I've not had that problem.


I've never experienced this with Amazon. Presumably it is a publisher problem, not a retailer problem. What publishers have you experienced this with?


It's endemic in the industry. With a few exceptions, try downloading any book printed before, say, the year 2000. A good example is the Kindle copy of Dune that Amazon sells. Spend $15 ($3 more than having the print version delivered to your doorstep!) on one of the most famous science fiction books of all time only to find that it's riddled with spelling errors, grammar errors, and word omissions. (Or just read the Amazon reviews that tell you as much.)

Another good example is the older Dark Tower books. I bought book 3 a year ago and it was in the same unreadable state. I also returned it for a refund.

New books aren't a problem, they're designed with ebooks in mind. It's just older books that the publisher quickly OCR'd and spammed Amazon with to make a quick buck.


I've seen this problem too with Amazon, though only when I first got the Kindle a couple of years ago. I suspect those books date back to when publishers didn't take ebooks seriously. Now that they do, the formatting / spelling errors have mostly vanished.


I tried doing this with some case studies in graduate school - any open sourced OCR is very difficult to deal with.

Of course, this was a couple years ago. Can anyone recommend a library or API that offers a decent OCR?


I just started a project of scanning in and OCR`ing old school news-papers. Tesseract [1] works very well. The result is almost always completely read-able, with a few obvious mistakes (that seem like could be reduced to almost none with a fairly simple post-processor). In terms of usability, it is a terminal program that is run as `tesseract srcFile.jpg dstFile`. It also has a list of gui front-ends on the site (none of which I have looked at).

[1] http://code.google.com/p/tesseract-ocr/


Cool - Tesseract was what I came across last week while working on one of my own projects. It's the only one I know of right now, and it's nice to hear someone confirm it.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: