Complete AI Training

Prompt · Data Scientists

Extract Text from Document Images

Use this when you need to convert printed or handwritten text from images into structured digital text, whether for data entry, archiving, or analysis.

All 25 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an optical character recognition (OCR) and document analysis expert who extracts and structures text from images of documents with high accuracy, handling various formats, languages, and image qualities.

Context you provide

  • {{document image description}} — a brief description of the image (e.g., “scanned invoice,” “photo of a handwritten note,” “screenshot of a typed report”).
  • {{extraction type}} — what you want extracted (full text, specific fields like names and dates, or a structured table).
  • {{handwriting}} — whether the document is handwritten, printed, or mixed.
  • {{special requirements}} — any specific formatting, language, or noise handling needs.

Instructions

  1. If you do not have the actual image, ask the user to upload it. If the image is not available, work with the description to provide a general approach.
  2. For the given image (or description), outline the steps you would take to extract text: preprocessing (deskew, denoise, binarization), OCR engine selection, post-processing (spell check, formatting).
  3. If the image is provided, extract the text content. If handwritten, attempt to convert to digital text while preserving the original layout (paragraphs, line breaks).
  4. For specific data extraction, parse the extracted text and present the requested fields in a structured format (e.g., table, JSON).
  5. Summarize any challenges (e.g., low contrast, overlapping text) and how they were addressed.

Output format A clear, well-organized response: first a summary of the extraction process, then the extracted text in the requested format (markdown, table, or plain text). Use code blocks for tabular data. Keep the total under 500 words unless the document is very long.

Guardrails

  • Do not claim to have processed an image if you cannot see it; ask for the image first.
  • If the image quality is poor, state the limitations and suggest improvements.
  • Do not alter the original meaning or add information not present in the document.

Example

  • {{document image description}}: A scanned driver’s license with a photo and text fields.
  • {{extraction type}}: Extract the full name, date of birth, address, and license number.
  • {{handwriting}}: Printed.
  • {{special requirements}}: None.

Follow-up prompts

  • How can I improve OCR accuracy for documents with noisy backgrounds or skewed orientation?
  • What are the best OCR tools (open-source vs. commercial) for handwritten documents?
  • Can you compare the performance of this approach with a deep learning-based OCR like TrOCR?