Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The other day I was thinking about what would be required to preserve all scientific knowledge. There are a great deal of papers that have been made available by hackers, but the filesize is quite large. Many of them are scanned documents.

Raw text itself is pretty compressible, but pdfs record everything about the layout and typesetting, font choice, small smudges on the page, etc. You could maybe run OCR on them and get the text (if the OCR is reliable enough.) But then you lose equations, figures, unusual symbols, and other important info.



There has been some effort in Project Gutenburg to typeset public domain books (and at least one journal issue) in mathematics: http://www.gutenberg.org/wiki/Mathematics_%28Bookshelf%29 Doing this for all of the scientific literature is rather daunting, however.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: