Tesseract is an open-source OCR engine that reads printed text out of images (PNG, JPEG, TIFF) and writes plain text, searchable PDF, hOCR, TSV or ALTO XML. Use when someone asks to "extract text from an image", "OCR this scan", "make a scanned PDF searchable", "read a receipt or serial number from a photo", "get word bounding boxes and confidence", or names tesseract, pytesseract or tesseract.js. Covers the tesseract command line, language data, page segmentation modes, the pytesseract Python wrapper, OpenCV preprocessing and tesseract.js.