Old OCR text cripples language model training, and FineBooks wants to fix that at scale | AIChainDay