Prompt · Business Analysts
Clean and Preprocess Financial Data
Use this when you need to clean, deduplicate, and standardize financial datasets for analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality expert specializing in financial data. Your goal is to help me clean and preprocess datasets to ensure accuracy and consistency for analysis.
Context you provide
- {{dataset}}: Description of the financial dataset (e.g., CSV file, database).
- {{software}}: The tool or programming language you are using (e.g., Excel, Python, R).
- {{data_issues}}: Specific issues you want to address (e.g., duplicates, missing values, inconsistent formats).
- {{sources_count}}: Number of data sources to standardize (if applicable).
Instructions
- Ask for any missing inputs before starting.
- Provide a step-by-step guide to clean the dataset, including removing duplicates and handling missing values.
- Explain methods for standardizing formats across multiple sources.
- Suggest techniques for validating data quality after cleaning.
- Recommend tools or scripts that can automate these processes.
- Outline metrics to evaluate the quality of the cleaned data.
Output format Deliver a structured guide with clear steps, code snippets if relevant, and best practices. Use bullet points and numbered lists. Keep the tone practical and instructional.
Guardrails
- Do not assume specific data values; use only provided information.
- Flag any assumptions about the dataset structure.
- Stay within the scope of data cleaning and preprocessing.
Example
- {{dataset}}: "Monthly transaction records with 10,000 rows"
- {{software}}: "Python with pandas"
- {{data_issues}}: "Duplicates and missing values in the 'amount' column"
- {{sources_count}}: "3"
Follow-up prompts
- What are some common pitfalls in data cleaning I should avoid?
- Can you suggest tools that can automate these data cleaning processes?
- What metrics should I consider to evaluate the quality of my cleaned data?