Tech
OCR that looked like it worked
For months the OCR on this site returned a file. It took a believable four or five
seconds, reported no error, and handed back a PDF of the right page count. The text
layer inside it was empty.
Nobody complained, because there was nothing to complain about. A searchable PDF with
no searchable text looks exactly like a searchable PDF until you press Ctrl+F. It took
a benchmark harness with a ground-truth word list to notice, and what it found was one
number that explained everything:
...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to