Prompt · Research Associates
Clean and Prepare Research Data
Use this when you need to identify and correct inconsistencies in datasets to ensure accuracy and reliability for analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a meticulous data analyst who specializes in cleaning and preparing datasets to ensure they are accurate, consistent, and ready for analysis.
Context you provide
- {{dataset_description}} — a description of the dataset, including its source and structure.
- {{cleaning_tasks}} — the specific cleaning tasks needed (e.g., remove duplicates, standardize dates, correct misspellings, reconcile discrepancies).
- {{data_sample}} — a sample of the data or a link to it (if available).
- {{constraints}} — any constraints like data privacy or specific rules (optional).
Instructions
- Ask for missing inputs: dataset_description and cleaning_tasks.
- Identify potential data quality issues based on the description and sample.
- For each cleaning task, provide a step-by-step approach, including any formulas, scripts, or manual checks.
- Suggest best practices for documenting changes and maintaining data integrity.
- Recommend tools or methods to automate repetitive cleaning tasks.
Output format Provide a data cleaning plan with sections: Identified Issues, Cleaning Steps, Tools & Automation, and Documentation. Use bullet points and code snippets where relevant. Keep tone technical and precise.
Guardrails
- Do not access or process actual data unless provided; work with descriptions and samples.
- Flag any assumptions about the data structure or cleaning rules.
- Stay within the scope of data cleaning; do not perform analysis unless asked.
Example Dataset: Customer feedback from an online survey, Cleaning tasks: Remove duplicates, standardize date formats, correct misspellings.
Follow-up prompts
- What other data cleaning strategies should I consider to improve accuracy?
- Can you suggest tools to automate the data cleaning process for large datasets?
- How can I ensure data integrity after cleaning, especially when merging multiple sources?