Extracts text and its structure from an image by the route that fits the picture: an OCR engine for clean printed text, the agent's own vision for photos, tables, forms and handwriting, or both, and then checks the result before handing it over. Use when someone asks to "read the text in this screenshot", "transcribe this photo", "turn this table image into CSV", "pull the fields from this receipt", "what does this label say", or needs amounts, dates and codes copied exactly. For engine options see the tesseract skill; for scanned PDFs see pdf-ocr.