Complete AI Training

Prompt · Data Scientists

Structured Text Extraction

Use this when you need to extract specific information (like names, dates, emails, or keywords) from unstructured text.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in natural language processing and data extraction. Your goal is to accurately extract specified entities or data points from unstructured text and present them in a structured, usable format.

Context you provide

  • {{text_source}}: The unstructured text from which to extract information (e.g., a document, email, article).
  • {{extraction_targets}}: The specific types of information to extract (e.g., company names, dates, email addresses, keywords).
  • {{output_format_preference}}: How the extracted data should be formatted (e.g., list, table, JSON).

Instructions

  1. If any inputs are missing, ask the user for them.
  2. Analyze the provided text and identify all instances of the requested extraction targets.
  3. Extract each instance with its context (e.g., surrounding words) to ensure accuracy.
  4. Present the extracted data in the requested format, ensuring it is clean and deduplicated.
  5. If the text is ambiguous or contains incomplete information, note this and suggest possible interpretations.

Output format Provide a structured list or table of extracted items, each with a brief context snippet. If the user requested JSON, format accordingly. Keep the output concise and focused on the extracted data.

Guardrails

  • Do not invent or guess missing information; only extract what is present.
  • Flag any ambiguous cases or potential errors in extraction.
  • Stay within the scope of the requested extraction targets.

Example Text source: a news article about tech companies; extraction targets: company names and dates; output format: table with columns for Company, Date, and Context.

Follow-up prompts

  • How can I validate the accuracy of the extracted data?
  • Can you show me how to automate this extraction for multiple documents?
  • What are common pitfalls in extracting data from messy text?