This is a problem domain where software hasn't caught up with what is possible, so people do in hardware what could be done in software.
With two or more photos or a stereo image (new iPhone?) one could triangulate to infer a flattened page, and produce images that look like they came from cut pages in a flatbed scanner. Now just pay someone well in Ethiopia to carefully turn pages without damage.
As any researcher can attest, our digital libraries now hold a century of scanned work of questionable quality. AI could infer scans indistinguishable from an outline font format original on an 8K monitor.
I once helped consult on the 1980's font wars, turning old formats and digital scans into Postscript and TrueType fonts. This was hard then, but will soon be understood as the "correct" way to scan text, when software catches up.
For the scientific literature, we need a ChatGPT equivalent to reconstruct LaTeX source that can reproduce each page. (We really need a successor to LaTeX that isn't such an arcane language, and can author fixed and flowable text with equal ease.)
This is very much a physical problem and there aren't too many shortcuts you can take.
I've been part of a preservation project and scanned a LOT of magazines. Generally, it doesn't matter much if you place it on a flat bed or not. Flipping the page manually takes a while either way.
There's already a known way for scanning magazines and books very fast: cut the spine and feed the pages to an automatic scanner. This is of course not applicable to anything you'd like to keep around after scanning, because your copy is destroyed.
All in all, the best way to automate scanning without destroying the item, will have to combine a top level camera with a machine to turn pages. I believe this is what was going on in Google's massive scanning project.
Maybe using x-ray could work for "scanning" some books without having to turn the pages. But I suppose there'll be a new set of problems to solve there.
I recently saw some work from the University of Kentucky on reading the Herculaneum scrolls. These scrolls were carbonized by volcanic activity (Pompeii?) and obviously can't be unrolled without disintegrating. They used some interesting CT (xray) scanning plus machine learning to distinguish the carbon-based ink from the mostly carbon substrate and retrieve legible text.
Of course, that only gets you the printed text. You might lose notes and doodles in the margins, or other physical evidence. But, it's certainly promising for works that are too delicate to physically open and inspect
RE "....There's already a known way for scanning magazines and books very fast: cut the spine and feed the pages to an automatic scanner. This is of course not applicable to anything you'd like to keep around after scanning, because your copy is destroyed......" I've always thought the pages could be rebound. Not perfect but a halfway solution.
Most signatures (https://en.wikipedia.org/wiki/Section_(bookbinding)) are glue bounded, not just sewn. Some use cases prefer/require the pages to be unbound because the printing goes all the way into the gutter and cutting the spine can also leave out some data. It's highly inneficient as you have to heat the spine carefully and then remove the glue residues. A tiny glue leftover can smear your autofeed scanner if not completely jam and tear the page. For a unique item, makes sense using a non-destructive scanning method, but for anything else, a carefully cut spine (or better yet, a bookbinding plow https://duckduckgo.com/?t=palemoon&q=bookbinding+plow&iax=im... ) leave a perfect cut and the loose pages can be kept in a ziplog bag for any future reference.
I guess you could, but everything where you would even consider cutting the spine is probably not worth fixing afterwards. E.g. it's a contemporary magazine, where you could just buy 2 if you really need to keep a physical copy.
How is setting up infrastructure to exploit third world labor a "software problem" exactly?
I think the problem isn't that software can't do de-warping well, it's that by the time you set up everything for book scanning you might as well use a setup that doesn't need it.
"Exploitation" is an arguable word, when without outside intervention there would be no employment above a few cents a day. Any employment by those who can pay even a dollar a day could instantly ratchet an entire family out of abject poverty.
You've also got to wonder about idea behind it that it's supposed to be trivial to send a large number of books to Ethiopia and back without damaging them.
There's also a feature where it tells you to turn the page, detects that it has been turned, takes a photo, etc. And in the background it flattens, splits into pages and OCRs the photos. With a little practice you can scan and OCR a whole book at 1-5 seconds per page.
Aren‘t you losing information in the parts that aren‘t perfectly straight? Yes, you can stretch those to recreate the original layout, but that would come at the cost of resolution in the interpolated sections of the page. Granted, not a problem for most books, but probably a reason prople are still looking for mechanical solutions to the problem.
I think the suggestion is that with AI you can interpolate to the actual letterforms, not to pixels.
Working a typical volume the letter “e” will appear hundreds of times and be identical, so there should be lots of data to help resolve ambiguities in the poorer parts of images.
Not to mention data that can be used across volumes.
If the goal is to ultimately OCR then it's moot. But yes, of course information is lost.
That being said, modern phone cameras are going to produce "scans" above 300 DPI, and while 600 DPI or higher might be tricky they're stills possible if you take partial shots of a document, assuming you can focus that close.
What you lose in quality you make up with convenience, I suppose.
> For the scientific literature, we need a ChatGPT equivalent to reconstruct LaTeX source that can reproduce each page. (We really need a successor to LaTeX that isn't such an arcane language, and can author fixed and flowable text with equal ease.)
Check out Nougat: OCRing scientific papers with a deep net trained end to end. It was released by Meta a few days ago.
“PDF format leads to a loss of semantic information, particularly for mathematical expressions. We propose Nougat (Neural Optical Understanding for Academic Documents), a Visual Transformer model that performs an Optical Character Recognition (OCR) task for processing scientific documents into a markup language, and demonstrate the effectiveness of our model on a new dataset of scientific documents.”
What about very old books, or handwritten books, or books with images in it?
The only valid archival approach has to be taking a good photo / scan while the page is as flat as possible.
Every further processing can be done later based on that, as a separate step... as technology advances.
Also carefully FLIPPING THE PAGES is literally the whole problem. Everything else can be solved by lowering a glass pane on the book and automatically taking a photo from above.
> As any researcher can attest, our digital libraries now hold a century of scanned work of questionable quality
I keep thinking so much collaborative potential is not being utilized. Imagine each (unique) Google Book is basically an editable wiki where people can directly correct OCR errors as they come across them (with an associated Talk page where they can give explanations, etc)
The insurmountable problem is that Google doesn't have, and can't get the rights to distribute all these books. Even regular Google employees can't see them (and I tried when in Legal there).
So this is neither a hardware problem nor a software problem. Unfortunately.
You are right! I was really thinking more about books published before the 20th century (which constitute the majority of my own usage of Google Books), and I had in mind something like a collaborative WikiSource integrated into Google Books.
With two or more photos or a stereo image (new iPhone?) one could triangulate to infer a flattened page, and produce images that look like they came from cut pages in a flatbed scanner. Now just pay someone well in Ethiopia to carefully turn pages without damage.
As any researcher can attest, our digital libraries now hold a century of scanned work of questionable quality. AI could infer scans indistinguishable from an outline font format original on an 8K monitor.
I once helped consult on the 1980's font wars, turning old formats and digital scans into Postscript and TrueType fonts. This was hard then, but will soon be understood as the "correct" way to scan text, when software catches up.
For the scientific literature, we need a ChatGPT equivalent to reconstruct LaTeX source that can reproduce each page. (We really need a successor to LaTeX that isn't such an arcane language, and can author fixed and flowable text with equal ease.)