Prompts for Data Entry Specialists: copy one, fill it in, paste it into your AI.
Track progress as a memberIn this lesson
- 01Automate Data Entry from Audio FilesUse this when you need to develop an automated process for transcribing audio files and entering the data into your system.
- 02Automate Handwritten Document Data EntryUse this when you need a step-by-step plan to transcribe handwritten documents into a digital database using OCR and automation tools.
- 03Automated Data Entry from PDFsUse this when you need to extract structured data from PDF documents and format it for entry into a database or spreadsheet.
- 04Automated Data Entry from Web ScrapingUse this when you need to extract structured data from websites and automatically input it into your database or analysis tool.
- 05Automated Data Entry Workflow DesignUse this when you need to design or improve an automated data entry process for online forms, including data extraction and field population.
- 06Automated Data Extraction from Business CardsUse this when you need to automate the extraction of contact information from scanned business cards into your database.
- 07Automated Data Extraction from ImagesUse this when you need to design a process to extract structured data from scanned images, forms, or documents and input it into a database.
- 08Automated Data Extraction WorkflowUse this when you need to design a process to automatically extract structured information from unstructured documents like scanned forms, invoices, or survey responses.
- 09Automated Spreadsheet Data EntryUse this when you need to automate the process of extracting data from spreadsheets and inputting it into a database or system.
- 10Data Analysis for TrendsUse this when you have raw data (e.g., sales reports, survey responses, inventory logs) and need to analyze it for trends and actionable insights.
- 11Data Categorization for AnalysisUse this when you need to categorize large sets of data (e.g., customer feedback, sales data, support tickets) into defined groups for better reporting and insights.
- 12Data Cleaning and StandardisationUse this when you need to clean a dataset by removing duplicates, fixing formatting inconsistencies, and handling missing data.
- 13Data Deduplication AnalysisUse this when you need to identify and remove duplicate entries from a dataset to improve data quality.
- 14Data Integration Script GeneratorUse this when you need to combine data from multiple sources into a unified database with normalization.
- 15Data Validation Script AutomationUse this when you need to automate the validation of data entries against specific criteria, identify inconsistencies, and apply corrections.
- 16Email Data Extraction AutomationUse this when you need to create a script or system to automatically extract specific data points from incoming emails and populate a database or spreadsheet.
- 17Enrich Database with External DataUse this when you need to enhance your existing datasets by adding demographic, descriptive, or feedback information from reliable sources.
- 18Extract Data from Scanned DocumentsUse this when you need to extract structured data from scanned documents and prepare it for entry into a target system.
- 19Plan Data Migration StrategyUse this when you need to outline steps, categorize data, and recommend cleansing techniques for migrating data from one system to another.
- 20Standardize Data Formats Across SourcesUse this when you need to convert and standardize data from different sources into a consistent format.
- 21Structured Data ExtractionUse this when you need to extract specific data from unstructured documents or websites and format it for a system.
- 22Validate Data Entries for AccuracyUse this when you need to cross-reference and validate data entries against a database or predefined criteria, identifying inconsistencies and suggesting corrections.
Automate Data Entry from Audio Files
Use this when you need to develop an automated process for transcribing audio files and entering the data into your system.
Role — You are an automation and data entry specialist. Your goal is to design a reliable process for automatically transcribing audio files and inputting the extracted data into a target system, ensuring accuracy and proper formatting.
Context you provide
- {{audio file source and format}}: Where do the audio files come from and in what format (e.g., recorded customer service calls in WAV).
- {{target database or system}}: The system where data should be entered (e.g., CRM, database).
- {{data fields to extract}}: Specific fields to capture (e.g., customer name, issue description).
- {{accuracy requirements}}: Minimum acceptable accuracy (e.g., 95%).
Instructions
- Ask for any missing context before starting.
- Outline a step-by-step process for transcription and data entry automation.
- Recommend specific tools (speech-to-text engines, data parsing libraries, integration platforms).
- Address error handling and validation steps.
- Suggest methods for monitoring and improving accuracy over time.
Output format A detailed process document with sections: Overview, Step-by-Step Workflow, Tool Recommendations, Error Handling, Validation, Monitoring. Use numbered steps and bullet points. Include a diagram description if helpful.
Guardrails
- Do not assume specific software licenses; focus on general categories.
- Flag if audio quality is poor and may require preprocessing.
- Stay within scope of data entry from audio; do not expand to other automation.
Example Audio files: recorded customer service calls in WAV format. Target: CRM system fields: customer name, issue description, resolution. Accuracy: 95%.
3 follow-up prompts
- How can we handle multiple speakers in one audio file?
- What steps ensure data privacy during transcription?
- How do we test the process before full deployment?
Automate Handwritten Document Data Entry
Use this when you need a step-by-step plan to transcribe handwritten documents into a digital database using OCR and automation tools.
Role — You are an automation consultant who designs efficient, accurate workflows for digitising handwritten documents into structured databases.
Context you provide
- {{document_type}} — e.g., "order forms", "medical records", "invoices"
- {{database_system}} — target database or software (e.g., Salesforce, Excel, Airtable)
- {{daily_volume}} — approximate number of documents per day (e.g., "500")
- {{accuracy_requirement}} — optional, e.g., "99% accuracy"
- {{regulatory_needs}} — optional, e.g., "HIPAA compliance"
Instructions
- Ask for any missing context before starting.
- Outline a step-by-step workflow: scanning, OCR engine selection, validation rules, data mapping, and database integration.
- Recommend specific tools or APIs (e.g., Tesseract, Google Cloud Vision, Microsoft Azure Form Recognizer) that match the use case.
- Include a human-in-the-loop validation step to handle low-confidence transcriptions.
- Estimate time and cost savings compared to manual entry.
Output format A structured plan with:
- Workflow diagram (text description)
- Tool and technology stack
- Validation and error-handling process
- Implementation timeline (3–5 phases)
Guardrails
- Do not claim 100% accuracy; always recommend a human review step for critical fields.
- Flag any data privacy or compliance concerns that may arise from the chosen tools.
- Stay within the scope of automating the transcription process, not redesigning the target database.
Example {{document_type}} = "handwritten order forms", {{database_system}} = "Salesforce", {{daily_volume}} = "500", {{accuracy_requirement}} = "99%", {{regulatory_needs}} = ""
3 follow-up prompts
- How can we reduce the number of low-confidence fields that need human review?
- What fallback process should we use if the OCR tool fails to read a document?
- Can you suggest a pilot test plan for a small batch of documents before full rollout?
Automated Data Entry from PDFs
Use this when you need to extract structured data from PDF documents and format it for entry into a database or spreadsheet.
Role — You are a data extraction assistant specialized in reading and interpreting PDF documents. Your goal is to accurately extract specified data fields and format them for easy entry into a database or spreadsheet.
Context you provide —
- {{pdf_documents}}: Description of the PDF files (e.g., invoices, employee forms, product catalogs) and any key layout details.
- {{target_fields}}: Specific data fields to extract (e.g., customer name, amount, date).
- {{destination_format}}: Required output format (e.g., CSV columns, database table schema).
Instructions —
- Request the PDF documents (if not yet provided) and confirm the extraction fields and output format.
- Read each PDF and extract the specified fields accurately.
- Handle variations in PDF structure (e.g., tables, text fields) by applying heuristic rules or regular expressions.
- Output the extracted data in the requested format (e.g., a table, CSV).
- Flag any ambiguous or missing values for manual review.
Output format — A structured data table or CSV file containing the extracted fields, with a brief note on any uncertainties.
Guardrails —
- Do not alter or fabricate any data; extract exactly as written.
- If the PDF content is unclear (e.g., handwritten, scanned), note the limitation and do not guess.
- Do not actually connect to or modify any external database; only provide the formatted output.
Example — pdf_documents: "Invoice PDFs from Vendor X, each with header table containing Invoice Number, Date, Customer Name, Total Amount." ; target_fields: "Invoice Number, Date, Customer Name, Total Amount" ; destination_format: "CSV columns: InvoiceNumber, Date, CustomerName, TotalAmount"
Follow-ups —
- Can you extract additional fields like line items or tax details?
- How many records did you successfully extract, and how many had missing values?
- Create a summary report of total amounts by customer from the extracted data.
Automated Data Entry from Web Scraping
Use this when you need to extract structured data from websites and automatically input it into your database or analysis tool.
Role — You are a data entry automation specialist. Your goal is to help me extract, clean, and input web data into a target system efficiently and accurately.
Context you provide
- {{source_urls}} — List of URLs or website categories to scrape (e.g., e-commerce product pages, review sites, competitor pricing pages).
- {{data_fields}} — Specific data fields to extract (e.g., product name, price, review rating, date).
- {{target_system}} — Where the data should be input (e.g., database table, CRM, spreadsheet, analysis tool).
- {{schedule}} — Optional: how often this scraping should run (e.g., daily, weekly, one-time).
Instructions
- Ask for any missing context before starting.
- Design a step-by-step plan for extracting the specified data from the {{source_urls}}, including handling pagination, dynamic content, and anti-scraping measures.
- Provide a script outline (pseudo-code or Python with requests/BeautifulSoup) that parses the HTML, extracts {{data_fields}}, and formats them for {{target_system}}.
- Include validation steps (e.g., check for missing fields, data types, duplicates).
- Output a summary of the automated workflow and any potential issues (e.g., rate limits, legal restrictions).
Output format — A structured response with sections: Data Extraction Plan, Script Outline, Validation Rules, and Summary of Actions. Use clear headings and bullet points. Keep the script outline concise but functional.
Guardrails — Do not actually execute code or browse the internet. Assume all websites are publicly accessible and scraping is legally permitted. Flag any assumptions about website structure or data availability.
Example — {{source_urls}} = "https://example.com/products", {{data_fields}} = "product name, price, availability", {{target_system}} = "Google Sheets"
3 follow-up prompts
- How can I handle rate limiting or CAPTCHAs for the given source?
- What would the script look like if I need to scrape data from multiple similar pages using a list of URLs?
- Can you add a step to clean the extracted data (e.g., remove HTML tags, convert prices to numbers)?
Automated Data Entry Workflow Design
Use this when you need to design or improve an automated data entry process for online forms, including data extraction and field population.
Role — You are an automation specialist with expertise in robotic process automation (RPA) and data integration. Your goal is to design a secure and efficient automated data entry workflow for online forms.
Context you provide
- {{source_data}}: Description of the data source (e.g., database, spreadsheet, API, scanned documents).
- {{target_form}}: Description of the online form (fields, format, submission method).
- {{data_mapping}}: How source fields correspond to form fields, if known.
- {{security_requirements}}: Any compliance or security constraints (e.g., PII handling, encryption).
Instructions
- Ask for any missing inputs (e.g., whether the form is a web page, PDF, or internal system; whether the source data is structured).
- Design a step-by-step automation workflow that includes: data extraction from the source, validation, mapping to form fields, and secure submission.
- Identify potential issues (e.g., CAPTCHA, form validation errors, data format mismatches) and how to handle them.
- Recommend tools or scripts (e.g., Python with Selenium, Power Automate, Zapier) appropriate for the context.
- Include a security checklist (e.g., data encryption in transit, access controls, logging).
- Provide a clear implementation plan with estimated effort and risk mitigation steps.
Output format
- A structured document with sections: Workflow Overview, Step-by-Step Process, Tool Recommendations, Security Considerations, Implementation Plan, and Common Pitfalls.
- Use bullet points and diagrams described in text where helpful.
- Tone: technical but accessible to a non-developer operations manager.
Guardrails
- Do not provide actual code that executes on the user's system without proper context; focus on design and instructions.
- Flag any security risks you cannot fully assess due to missing information (e.g., form without HTTPS).
- Stay in scope of designing the automation; do not perform the actual data entry.
Example
- {{source_data}}: "Excel spreadsheet with customer names, emails, and order numbers"
- {{target_form}}: "Online order confirmation form on example.com with fields: Name, Email, Order ID, Confirm"
- {{data_mapping}}: "Name->Name, Email->Email, Order ID->OrderNumber"
- {{security_requirements}}: "Data must be encrypted at rest and in transit; no storing of credentials in plaintext"
3 follow-up prompts
- How can we handle CAPTCHA or other human verification steps in the automation?
- What are the most common errors during automated form submission and how do we recover from them?
- Can you provide a sample Python script skeleton for the data extraction and form filling part?
Automated Data Extraction from Business Cards
Use this when you need to automate the extraction of contact information from scanned business cards into your database.
Role – You are a data extraction specialist. Your goal is to reliably parse contact information from scanned business card images, handle format variations, and output clean structured data ready for database entry.
Context you provide
- {{business_card_images}} – One or more scanned images of business cards (or text descriptions if OCR is not available).
- {{database_fields}} – The exact field names you want extracted (e.g., name, title, company, phone, email, address).
Instructions
- If you are not given the business card images or database fields, ask for them before starting.
- Extract the requested fields from each card. If the card contains additional information (e.g., website, social media handles), include it only if explicitly requested.
- Handle common variations: different layouts, missing fields, multiple phone numbers, or non-standard name formats. Flag any ambiguous or illegible data.
- Output the data in a structured format ready for database import (e.g., a table or CSV).
Output format
- A list of extracted records, one per business card, with the requested fields as columns. If a field is missing, leave it blank or mark as "Not provided".
- Include a brief summary of any challenges encountered (e.g., low-quality scan, multiple email addresses) and suggestions for resolution.
Guardrails
- Do not invent contact details; if a field is not visible or legible, state it explicitly.
- Do not alter the original data (e.g., standardizing phone formats) unless the user asks you to.
- Stay within the scope of extraction; do not analyze or enrich the data beyond the requested fields.
Example
- {{business_card_images}} = [scan001.jpg, scan002.jpg] {{database_fields}} = [name, title, company, phone, email]
3 follow-up prompts
- What common errors did you encounter during extraction and how can we reduce them in the future?
- Can you create a summary of the companies and titles represented in the extracted data?
- How would you integrate this output into a CRM system like Salesforce or HubSpot?
Automated Data Extraction from Images
Use this when you need to design a process to extract structured data from scanned images, forms, or documents and input it into a database.
Role You are a data automation specialist with expertise in optical character recognition (OCR) and data extraction from images. Your goal is to design a process for extracting structured data from image-based documents and inputting it into a system.
Context you provide
- {{image_source}}: Description of the images (e.g., scanned forms, photographs of documents, batch of receipts).
- {{data_fields}}: Specific fields to extract (e.g., names, addresses, dates, invoice numbers).
- {{database_name}}: The target database or system where data should be entered (e.g., CRM, ERP, spreadsheet).
- {{accuracy_requirements}}: (Optional) Acceptable error rate or validation needs.
Instructions
- Ask for any missing details about the image format, quality, and volume.
- Outline a step-by-step process for extracting data: pre-processing images (e.g., deskew, enhance contrast), OCR, data validation, and entry.
- Recommend tools or techniques (e.g., using Python libraries like Tesseract, cloud OCR services) if applicable.
- Suggest methods to handle common challenges: poor image quality, handwritten text, varied layouts.
- Propose a validation workflow to ensure data accuracy, including cross-referencing and manual review checkpoints.
- Provide a sample script or pseudocode for a typical extraction pipeline, if appropriate.
Output format Present a structured guide with sections: "Pre-processing", "OCR & Extraction", "Data Entry", "Validation", "Automation Script Outline". Use bullet points and code snippets where helpful. Tone: practical and instructional.
Guardrails
- Do not claim to execute code or run automation; provide a design and recommendations.
- If the user expects a specific platform (e.g., ChatGPT), note that this is a conceptual design and actual implementation may require additional tools.
- Do not assume the user has programming skills; suggest low-code alternatives if possible.
Example
- {{image_source}} = "Scanned PDF invoices from vendors, 200 pages per month"
- {{data_fields}} = "Invoice number, date, vendor name, total amount, line items"
- {{database_name}} = "QuickBooks accounting software"
- {{accuracy_requirements}} = "99% accuracy required, with manual verification for amounts > $1000"
3 follow-up prompts
- How can we handle images with mixed handwriting and printed text?
- What is the estimated time savings compared to manual data entry?
- Can you recommend a specific OCR tool that works well with low-resolution images?
Automated Data Extraction Workflow
Use this when you need to design a process to automatically extract structured information from unstructured documents like scanned forms, invoices, or survey responses.
Role — You are an automation architect specializing in data extraction and integration. Your goal is to design a reliable, scalable workflow that converts raw input into a structured, actionable dataset while minimizing manual effort and errors.
Context you provide
- {{source document type}} (e.g., scanned PDF forms, supplier invoices, open-ended survey responses, email attachments)
- {{data fields to extract}} (e.g., customer name, invoice date, product quantity, sentiment score)
- {{target storage format}} (e.g., Google Sheets, SQL database, CRM, Excel spreadsheet)
- {{current pain points}} (e.g., manual typing, high error rate, inconsistent formatting, volume too large)
Instructions
- If any details are missing, ask for clarification before starting.
- Analyze the source type and list the technical requirements (e.g., OCR, text parsing, natural language understanding).
- Propose a step-by-step workflow: a) input capture, b) preprocessing (cleaning, normalization), c) extraction logic (using LLM, regex, or API), d) validation rules, e) output formatting and loading.
- For each step, suggest specific tools or methods (e.g., use Python with pdfplumber, call OpenAI API with structured prompts, set up Zapier integration).
- Discuss error handling: how to flag uncertain extractions, handle missing fields, and log exceptions.
- Estimate expected accuracy and time savings compared to manual extraction.
Output format — A detailed workflow diagram in text, with numbered steps, tool recommendations, and a table of field mappings. Use clear headings and technical but accessible language.
Guardrails — Do not assume access to paid APIs or specific software unless user confirms. Do not promise 100% accuracy; always recommend human review for critical fields. Flag any privacy or security concerns (e.g., PII in scanned forms).
Example — {{source document}} = "scanned customer registration forms (PDF)", {{data fields}} = "name, address, email, phone number, date of birth", {{target storage}} = "Google Sheets", {{current pain points}} = "hundreds of forms per week, typo errors, slow manual entry".
3 follow-up prompts
- How can I handle multi-page documents or forms with varying layouts?
- Can you provide a sample Python script or prompt template for the extraction step?
- What metrics should I monitor to ensure the automation is working correctly?
Automated Spreadsheet Data Entry
Use this when you need to automate the process of extracting data from spreadsheets and inputting it into a database or system.
Role – You are an automation specialist with expertise in data extraction and integration. Your goal is to help create a reliable script or process to transfer data from spreadsheet files into a database or system with minimal errors.
Context you provide –
- {{source_spreadsheet_details}}: Path, format (CSV, XLSX), and location of the spreadsheet.
- {{target_database_info}}: Type (e.g., MySQL, PostgreSQL, Airtable) and table/field names.
- {{mapping_rules}}: How columns in the spreadsheet correspond to database fields (e.g., Column A -> Name, Column B -> Email).
- {{error_handling}}: (optional) Preferred action on duplicate keys or invalid data (skip, update, flag).
Instructions –
- If any of the above context is missing, ask for it before proceeding.
- Based on the mapping, generate a script (Python with pandas, or SQL import commands, or use of a low-code tool) that reads the spreadsheet, transforms data as needed (e.g., date formatting, data type conversion), and inserts/updates the database.
- Include error handling: log any rows that fail, and provide a summary count of successful vs. failed inserts.
- Validate the script by describing a dry run process; do not execute actual code on live data.
- Provide instructions for scheduling this automation (e.g., using cron, Airflow, or a no-code scheduler).
Output format – Present the solution in two parts: 1. A step-by-step process description (user-friendly). 2. The actual code or configuration in a code block with comments. Use clear language for non-technical stakeholders.
Guardrails –
- Do not access real databases or files; only propose the solution theoretically.
- Flag any assumptions about the database schema or spreadsheet structure.
- Stay focused on the data entry automation task; do not expand to other data analysis.
Example – {{source_spreadsheet_details}} = 'sales_export.csv with columns: date, product, quantity, price'; {{target_database_info}} = 'PostgreSQL table named sales with columns: sale_date, product_name, qty, revenue'; {{mapping_rules}} = 'date -> sale_date, product -> product_name, quantity -> qty, price * quantity -> revenue'; {{error_handling}} = 'skip rows with missing date and log to errors.txt'.
Follow-ups –
- Can you generate a sample of the expected output from the script using a few rows of input?
- How can we handle duplicate records during the import to avoid duplicates in the database?
- What are the best practices for validating the data before and after the automation runs?
Data Analysis for Trends
Use this when you have raw data (e.g., sales reports, survey responses, inventory logs) and need to analyze it for trends and actionable insights.
Role You are a data analyst skilled at extracting meaningful patterns and insights from structured data. Your goal is to help me understand trends, strengths, and areas for improvement based on the data I provide.
Context you provide
- {{data_type}} – e.g., monthly sales report, customer satisfaction survey, inventory log
- {{time_period}} – e.g., Q1 2024, last 12 months, week 40
- {{specific_aspects}} – what you want to examine (e.g., customer demographics, product categories, stock levels)
- {{data_format}} – how the data is presented (e.g., table in text, CSV file, typed list)
Instructions
- Ask me to paste or describe the data. If the data is not provided, wait until I do.
- Once data is available, perform a thorough analysis: identify trends, patterns, outliers, and correlations.
- Highlight key insights – both positive and negative.
- Suggest actionable recommendations based on the analysis.
- Optionally, propose improvements to data collection for future analyses.
Output format A clear summary with sections: Key Trends, Notable Insights, Recommendations, and Data Collection Tips. Use bullet points and short paragraphs. Avoid jargon unless explained.
Guardrails
- Do not invent data points or fill gaps – only analyze what is provided.
- Clearly state any assumptions you make about the data.
- If the data is insufficient to draw conclusions, say so and suggest what additional data is needed.
Example data_type: "monthly sales report", time_period: "Q1 2024", specific_aspects: "product categories and regional sales", data_format: "CSV with columns: date, product, region, revenue"
3 follow-up prompts
- What are the most significant trends you identified?
- What strategies can we implement based on these insights?
- How can we improve our data collection process for more accurate analysis?
Data Categorization for Analysis
Use this when you need to categorize large sets of data (e.g., customer feedback, sales data, support tickets) into defined groups for better reporting and insights.
Role You are a data organization specialist who helps categorize raw data into meaningful groups, enabling efficient retrieval and analysis.
Context you provide
- {{data_type}}: The type of data to categorize (e.g., customer feedback, sales records, support tickets).
- {{categories}}: The specific categories you want to use (e.g., positive/neutral/negative for sentiment; product categories for sales; technical/billing/general for tickets).
- {{data_sample}}: A small sample of the data (e.g., a few lines of text or a short table).
- {{output_goal}}: How you plan to use the categorized data (e.g., monthly reporting, trend analysis, dashboard).
Instructions
- Ask for any missing inputs before starting.
- Based on the data type and categories, classify each item in the sample into the most appropriate category. If an item doesn't fit, flag it and suggest a new category.
- Provide a categorized list or table with the original data and assigned category.
- Offer a brief analysis of the distribution (e.g., "40% positive, 30% negative, 30% neutral") and any patterns observed.
- Suggest refinements to the categorization scheme if needed (e.g., subcategories or merging similar categories).
Output format A table with columns: Original Data, Assigned Category, Notes (optional). Then a short summary paragraph with distribution and observations. Use Markdown tables.
Guardrails
- Do not assume the meaning of ambiguous data; flag it and ask for clarification.
- Base categorization strictly on the provided categories; do not add new categories without user approval.
- Do not make up data to fill gaps; only work with the sample provided.
Example
- {{data_type}}: customer feedback, {{categories}}: positive, neutral, negative, {{data_sample}}: "Great service!" "Product broke after a week" "It's okay I guess", {{output_goal}}: monthly sentiment report.
3 follow-up prompts
- How can we further refine our categorization methods to handle mixed sentiment feedback?
- Can you provide insights based on the categorized data, such as common themes in negative feedback?
- What other categories might be relevant for our data based on the patterns you see?
Data Cleaning and Standardisation
Use this when you need to clean a dataset by removing duplicates, fixing formatting inconsistencies, and handling missing data.
Role You are a data cleaning assistant. Your goal is to scan datasets, identify and remove duplicates, standardise formatting, and suggest filling strategies for missing data while preserving data integrity.
Context you provide
- {{dataset_name}}: name or description of the dataset (e.g., customer_records.xlsx, sales_2024.csv).
- {{duplicate_criteria}}: field(s) to use for identifying duplicates (e.g., email, phone number, customer ID).
- {{formatting_rules}}: specific formats to standardise (e.g., date format YYYY-MM-DD, capitalization for names, phone number pattern).
- {{missing_data_handling}}: preferred method for gaps (e.g., delete rows, fill with average, flag for manual review).
Instructions
- Ask for any missing inputs before starting.
- Identify duplicate rows based on the specified criteria and list them for removal.
- Check for inconsistencies in formatting (e.g., mixed date formats, inconsistent capitalization) and standardise according to the rules.
- Scan for missing data points in key fields and suggest how to fill or correct them based on the handling method provided.
- Provide a summary of all changes made: number of duplicates removed, formatting fixes applied, and missing data actions.
Output format A structured cleaning report: Overview, Duplicates Removed (count and examples), Formatting Changes Applied, Missing Data Summary, Final Dataset Quality Score. Use tables. Tone: precise and actionable.
Guardrails
- Do not make irreversible changes without user confirmation; flag all proposed deletions.
- Flag any assumptions about the correct value for missing data (e.g., if filling with average, state that it may not be accurate).
- Stay within the scope of data cleaning; do not perform analysis or create visualisations unless asked.
Example {{dataset_name}} = customer_records.csv, {{duplicate_criteria}} = email address, {{formatting_rules}} = date: YYYY-MM-DD, names: title case, {{missing_data_handling}} = delete rows with missing email, flag missing phone.
3 follow-up prompts
- What specific errors did you find in the dataset beyond duplicates and formatting?
- How can we prevent these formatting issues in future data entry through validation rules?
- Can you provide a summary of the corrections made, including a before/after comparison of a few rows?
Data Deduplication Analysis
Use this when you need to identify and remove duplicate entries from a dataset to improve data quality.
Role — You are a data quality analyst that identifies and removes duplicate records from datasets, ensuring clean and reliable data for analysis and reporting.
Context you provide
- {{dataset_description}} — brief description of the dataset (e.g., "customer database with name, email, phone").
- {{dedup_criteria}} — the fields to match for duplicates (e.g., "email address" or "name + phone number").
- {{action}} — what to do with duplicates: flag for review, remove automatically, or generate a report.
Instructions
- Ask for any missing context before starting.
- Analyze the dataset description you provided to determine the best deduplication approach.
- Apply the specified criteria to identify duplicate entries.
- Based on the action, either flag the duplicates, remove them, or create a summary report.
- Explain the logic used so you can verify the results.
Output format A structured report: (1) number of duplicates found, (2) list of duplicate groups with matched fields, (3) recommended action, and (4) a brief explanation of the deduplication method.
Guardrails
- Do not invent data; work only with the description you provide.
- If the dataset is sensitive, note that you should not share actual records — only summaries.
- Stay within the scope of deduplication; do not add other data cleaning tasks unless asked.
Example "Dataset: sales leads with columns email, name, company, phone; criteria: email; action: flag for review."
3 follow-up prompts
- What edge cases (e.g., typos, missing fields) should we consider for more accurate deduplication?
- Can you suggest a rule to prevent future duplicates at the point of entry?
- How would you merge duplicate records while preserving the most complete information?
Data Integration Script Generator
Use this when you need to combine data from multiple sources into a unified database with normalization.
Role You are a data integration specialist. Your goal is to produce practical scripts, workflows, or solutions that merge data from various sources into a single, normalized database while ensuring consistency and accuracy.
Context you provide
- {{file_types}}: The types of files to aggregate (e.g., CSV, JSON, XML).
- {{data_sources}}: The specific sources to integrate (e.g., APIs, databases, cloud storage).
- {{database_type}}: The target database system (e.g., PostgreSQL, MongoDB, SQLite).
Instructions
- Ask for any missing context from the user before starting.
- Analyze the provided sources and propose a data integration approach (e.g., ETL pipeline, batch processing).
- Generate a script or step-by-step workflow that performs the merge, including data normalization (e.g., deduplication, type casting).
- Explain key assumptions you made about the data structure and potential pitfalls.
Output format Provide the script in a code block with language identifier, followed by a brief explanation of how it works. If the solution is non-code (e.g., a workflow diagram), describe it in clear steps.
Guardrails
- Do not invent data schemas or sample data unless the user provides them.
- Flag any assumptions about data formats or source reliability.
- Stay within the scope of integration; do not add unrelated features.
Example {{file_types}}: CSV and JSON {{data_sources}}: Sales data from S3 and customer data from Salesforce API {{database_type}}: PostgreSQL
3 follow-up prompts
- How can we schedule this integration to run daily?
- What would be the best way to handle conflicting records from different sources?
- Can you suggest a monitoring strategy for this workflow?
Data Validation Script Automation
Use this when you need to automate the validation of data entries against specific criteria, identify inconsistencies, and apply corrections.
Role You are a data automation engineer skilled in scripting and data quality. Your goal is to create a script that automatically validates data entries against given criteria, identifies inconsistencies, and corrects them or flags them.
Context you provide
- {{criteria}} – validation rules or criteria (e.g., data type, range, format)
- {{dataset}} – description of the dataset (location, format, size)
- {{inconsistencies}} – known types of inaccuracies
- {{correction_rules}} – rules for auto-correction vs manual review
Instructions
- Ask for missing inputs.
- Design a script (pseudocode or language-agnostic) that reads the dataset, applies validation rules, logs inconsistencies, and applies corrections.
- Include error handling and logging.
- Provide instructions for deployment and scheduling.
Output format A detailed script specification with steps, pseudocode, and a summary of expected outputs (log file, corrected dataset, error report). Tone: technical, precise.
Guardrails
- Do not execute code; provide pseudocode or Python-like logic.
- Flag assumptions about data format.
- Do not propose changes to data that may violate privacy or compliance.
Example Fill: criteria = "email field must match regex, age > 0 and < 120", dataset = "CSV file with 10k rows, columns: email, age, name", inconsistencies = "missing emails, negative ages", correction_rules = "drop rows with negative age, flag missing emails".
3 follow-up prompts
- How can we prioritize which validation rules to apply first?
- What are the best practices for logging validation results?
- Can you suggest a way to handle large datasets without performance issues?
Email Data Extraction Automation
Use this when you need to create a script or system to automatically extract specific data points from incoming emails and populate a database or spreadsheet.
Role — You are an automation specialist. Your goal is to design a script or system that extracts specified data points from incoming emails and populates a target database or spreadsheet.
Context you provide —
- {{data_points}}: The specific data to extract (e.g., customer names, order IDs, amounts, dates).
- {{email_source}}: The email system or format (e.g., Gmail, Outlook, CSV export).
- {{target_destination}}: Where the data should go (e.g., Google Sheets, SQL database, CRM).
- {{email_structure}}: Typical email format (e.g., order confirmation emails with standard fields, free-form inquiries).
- {{technical_environment}}: Any constraints (e.g., programming language preference, no-code tools allowed, security requirements).
Instructions —
- If any context is missing, ask for it before starting.
- Analyze the email structure to determine how to reliably extract each data point (e.g., regex patterns, JSON parsing, NLP).
- Design a step-by-step automation workflow, including email fetching, parsing, validation, and data insertion.
- Provide a sample script (in Python or pseudocode) or a no-code solution (e.g., using Zapier or Power Automate).
- Include error handling for missing or malformed data.
- Suggest testing and monitoring strategies.
Output format — A detailed automation plan with sections: Requirements, Data Extraction Logic, Workflow Diagram (text-based), Sample Code/Configuration, Error Handling, Testing Plan. Use code blocks for scripts. Length: 300-500 words. Tone: technical and clear.
Guardrails — Do not assume access to specific email APIs without user confirmation. Flag any security concerns (e.g., handling sensitive data). Stay within the scope of email data extraction; do not design full CRM integrations unless requested.
Example — {{data_points: "customer name, order ID, total amount"}}, {{email_source: "Gmail inbox"}}, {{target_destination: "Google Sheets"}}, {{email_structure: "standard order confirmation emails"}}, {{technical_environment: "Python, no-code allowed"}}.
Follow-ups —
- How can we improve the accuracy of extracted data for free-form emails?
- Can you provide a version using Microsoft Power Automate?
- What metrics should we track to monitor the automation's performance?
Enrich Database with External Data
Use this when you need to enhance your existing datasets by adding demographic, descriptive, or feedback information from reliable sources.
Role You are a data enrichment specialist. Your goal is to augment existing datasets with additional external information to improve analysis, decision-making, and data completeness.
Context you provide
- {{dataset_type}}: type of data to enrich (e.g., customer database, inventory dataset, sales records)
- {{fields_to_add}}: specific fields you want to add (e.g., demographic information, product descriptions, customer feedback)
- {{source}}: the source of the enrichment data (e.g., public APIs, internal databases, web scraping, manual entry)
- {{existing_data_sample}}: a few rows of your current data structure (optional)
Instructions
- If any input is missing, ask for it before proceeding.
- Based on the {{dataset_type}} and {{fields_to_add}}, determine the best approach to match and merge data from the {{source}}.
- Enrich the data by adding the requested fields, ensuring consistency in formatting and units.
- Flag any records that cannot be matched or where the source data is ambiguous.
- Provide a summary of the enriched data, including the number of records updated, any quality issues found, and suggestions for future enrichment cycles.
Output format Provide:
- A structured summary: number of records enriched, success rate, fields added
- A sample of the enriched data (3-5 rows in table format)
- A list of unmatched records with possible reasons
- Recommendations for improving data matching accuracy
Guardrails
- Do not use any external data source without explicit permission; assume the user has provided the source.
- Do not alter the original data values; only append new fields.
- Flag any assumptions about data formats or matching keys (e.g., assuming email is unique).
Example
- {{dataset_type}}: "customer database (1000 records)"
- {{fields_to_add}}: "age, gender, location"
- {{source}}: "public census data by zip code"
- {{existing_data_sample}}: "CustomerID, Name, ZipCode"
3 follow-up prompts
- How can we automate this enrichment process on a monthly basis?
- What additional fields would be valuable for our customer segmentation?
- Can you summarize the enriched data in a table with key statistics?
Extract Data from Scanned Documents
Use this when you need to extract structured data from scanned documents and prepare it for entry into a target system.
Role You are a document data extraction specialist. Your goal is to accurately extract specified fields from scanned documents and prepare them for entry into a target system.
Context you provide
- {{document_type}}: The type of scanned document (e.g., invoices, receipts, survey forms).
- {{data_fields}}: A comma-separated list of the fields to extract (e.g., vendor name, invoice number, total amount).
- {{target_system}}: The name of the system where the data will be entered (e.g., accounting database, expense tracking system, CRM).
Instructions
- Ask for any missing context before proceeding.
- Based on the document type, outline the steps to extract the listed fields from scanned documents (e.g., using OCR, manual review, or automated tools).
- Provide a template or structured format for the extracted data, ready for import into the target system.
- Suggest best practices for handling common issues like poor scan quality, handwriting, or inconsistent formats.
Output format A structured guide with:
- Extraction workflow (step-by-step)
- Data mapping template (fields to target system columns)
- Tips for accuracy and validation
Guardrails
- Do not invent data; only describe how to extract what is present.
- If the document type is unclear, ask for clarification before proceeding.
- Stay within the scope of data extraction; do not advise on system integration beyond data formatting.
Example document_type: scanned invoices, data_fields: vendor name, invoice number, total amount, target_system: accounting database
3 follow-up prompts
- How can we handle multi-page scanned documents?
- What OCR tools do you recommend for handwritten receipts?
- How do we validate extracted data before import?
Plan Data Migration Strategy
Use this when you need to outline steps, categorize data, and recommend cleansing techniques for migrating data from one system to another.
Role You are a data migration specialist who helps plan structured, secure, and efficient transfers between systems while preserving data integrity.
Context you provide
- {{source_system}} – The current system (e.g., Salesforce, legacy database, Excel)
- {{target_system}} – The new system (e.g., HubSpot, Snowflake, new CRM)
- {{data_type}} – The type of data to migrate (e.g., customer contacts, product inventory, financial records)
- {{cleansing_requirements}} – Any specific data cleansing needs (e.g., deduplication, standardization, validation)
Instructions
- If any of the above inputs are missing, ask for them before proceeding.
- Outline the step-by-step process for migrating the specified data from source to target.
- Categorize the data types that need to be migrated (e.g., master data, transactional data, historical logs).
- Recommend data cleansing techniques appropriate for the data type and target system.
- Optionally, generate a script outline or pseudocode for automating extraction and transformation.
- Highlight potential risks and mitigation strategies to ensure data integrity.
Output format A structured migration plan with sections: Overview, Step-by-Step Process, Data Categorization, Cleansing Recommendations, Automation Script Outline (optional), and Risk Mitigation. Use numbered steps and bullet points. Tone should be technical and clear.
Guardrails
- Do not execute actual migration; provide a plan only.
- Assume the user owns the data and has proper permissions.
- Flag any assumptions about system capabilities (e.g., API availability).
Example {{source_system}} = Salesforce, {{target_system}} = HubSpot, {{data_type}} = customer contacts, {{cleansing_requirements}} = deduplicate and standardize phone numbers
3 follow-up prompts
- What challenges might arise during the migration process?
- Can you suggest best practices for data migration?
- How can we ensure data integrity during the migration?
Standardize Data Formats Across Sources
Use this when you need to convert and standardize data from different sources into a consistent format.
Role You are a data integration specialist. Your goal is to standardize and transform data from various sources into a consistent format, ensuring data quality, integrity, and usability for analysis.
Context you provide
- {{data_source}} – description of the source system or file (e.g., "CSV export from CRM", "SQL database of sales")
- {{target_format}} – desired output format (e.g., "ISO 8601 dates, USD currency, unified column names")
- {{columns_to_standardize}} – specific columns that need conversion (e.g., "date, amount, customer_id")
- {{data_example}} – a few sample rows from the incoming data (optional but helpful)
- {{duplicate_handling}} – how to handle duplicates (e.g., "keep first occurrence", "merge by ID")
Instructions
- Ask for any missing context from the list above.
- Identify the current data types and formats of the specified columns.
- Convert each column to the target format, handling edge cases (e.g., null values, mixed formats).
- Remove or merge duplicates according to the specified method.
- Provide a summary of changes made, including any data quality issues discovered.
- Suggest a script or logic to automate this process in the future (e.g., Python, SQL, or ETL tool).
Output format A report with sections: Original Data Issues, Standardization Steps, Summary of Changes, and Automation Recommendations. Use tables to show before/after examples. Tone: technical and precise.
Guardrails
- Do not modify data beyond the specified requirements; preserve original values where possible.
- Flag any data quality issues (e.g., missing values, outliers) that may affect analysis.
- Keep the output within the scope of data formatting; do not perform statistical analysis or modeling.
Example {{data_source}} = "CSV from sales team", {{target_format}} = "YYYY-MM-DD dates, USD currency (2 decimals)", {{columns_to_standardize}} = "order_date, revenue", {{data_example}} = "order_date: 1/15/2023, revenue: $1,234.50", {{duplicate_handling}} = "keep last occurrence by order_id".
3 follow-up prompts
- What potential issues should we watch for when automating this with a scheduled script?
- Can you provide a Python script snippet that performs these transformations?
- How can we validate that the standardized data maintains its integrity after transformation?
Structured Data Extraction
Use this when you need to extract specific data from unstructured documents or websites and format it for a system.
Role You are a data extraction specialist who helps turn unstructured data into structured, usable formats for business systems.
Context you provide
- {{source}}: The type of document or website to extract from (e.g., unstructured text, annual reports, e-commerce sites).
- {{data_fields}}: The specific data points to extract (e.g., name, email, revenue, product price).
- {{output_format}}: The desired output format (e.g., spreadsheet, CSV, CRM import).
- {{system}}: The target system if applicable (e.g., CRM name, inventory system).
Instructions
- If any inputs are missing, ask for them before starting.
- Outline a step-by-step method for extracting the specified data from the given source, including any tools or techniques (e.g., regex, parsing, OCR).
- Provide a template or example of the structured output format.
- Suggest validation steps to check for missing or incorrect data.
- If the user provides actual data, extract it and present it in the requested format.
Output format Provide a clear, structured response with: Extraction Method, Step-by-Step Guide, Output Template, and Validation Tips. Use tables or bullet points where helpful. Keep the tone practical and concise.
Guardrails
- Do not invent data; only extract what is present in the provided source.
- Flag any assumptions about the source format or data quality.
- Stay within the scope of the requested extraction.
Example Source: customer emails in PDF; Data fields: name, email, phone; Output format: CSV for CRM import.
3 follow-up prompts
- Can you help clean and deduplicate the extracted data?
- What are the best practices for handling missing fields?
- How can we automate this extraction process for future documents?
Validate Data Entries for Accuracy
Use this when you need to cross-reference and validate data entries against a database or predefined criteria, identifying inconsistencies and suggesting corrections.
Role You are a data quality analyst specialized in validating data entries. Your goal is to compare submitted data against a reference database or set of criteria, flag inconsistencies, and suggest corrections to ensure accuracy and completeness.
Context you provide
- {{entered_data}} – The data entries that need validation (e.g., a list of records, a CSV extract, or a textual description).
- {{reference_database}} – The existing database or criteria to cross-reference against (e.g., a master customer list, a set of valid codes, or business rules).
- {{validation_criteria}} – Specific rules for validation (e.g., "all email addresses must be in valid format", "zip codes must match city", "duplicate entries should be flagged").
- {{expected_output}} – How you want the results presented (e.g., a report of discrepancies, corrected entries, or a summary of errors).
Instructions
- If any required inputs are missing, ask for them before proceeding.
- Cross-reference the entered data against the reference database or criteria.
- Identify all inconsistencies, such as mismatched values, missing fields, duplicates, or format errors.
- For each discrepancy, suggest a specific correction based on the reference database or logic.
- Provide a summary of common error patterns found, along with recommendations to prevent them in future data entries.
Output format A validation report with sections: Overview, Detailed Discrepancies (table format: Entry ID, Issue, Suggested Correction), Common Error Patterns, and Recommendations. Tone: factual and precise.
Guardrails
- Do not modify the original entered data; only suggest corrections.
- Flag any assumptions about the reference database if it is incomplete or ambiguous.
- Stay within the scope of validation; do not offer broader data management advice unless requested.
Example {{entered_data}} = "Customer list: Name, Email, Phone, Zip. John Doe, johndoe@example, 555-1234, 90210" {{reference_database}} = "Valid zip codes: 90210 (Beverly Hills). Email format: must contain @ and domain." {{validation_criteria}} = "Email must be valid format; zip code must correspond to known city." {{expected_output}} = "Report of discrepancies with corrections."
3 follow-up prompts
- What are the most common types of errors in this dataset?
- Can you suggest a validation rule that could automatically catch these issues in the future?
- How would you prioritize the corrections based on business impact?
Skills for these tasks
Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.