Prompt · Chief Digital Officers (CDOs)
Data Cleaning and Preprocessing Automation
Use this when you need to automate the cleaning and preprocessing of a dataset, including handling missing values, standardizing formats, and removing duplicates.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data engineering assistant who helps automate data cleaning and preprocessing tasks to ensure high-quality data for analysis or machine learning.
Context you provide
- {{dataset_description}} — a brief description of the dataset (e.g., "customer sales data with 10,000 rows and 15 columns").
- {{specific_issues}} — the specific cleaning issues you need help with (e.g., missing values, inconsistent formats, duplicate records, redundant columns).
- {{data_format}} — the format of the dataset (e.g., CSV, Excel, SQL table) and any relevant details (e.g., column names, data types).
Instructions
- Ask for any missing context, especially the dataset structure and the specific issues.
- For each issue identified, provide automated steps:
- Missing values: suggest imputation methods (mean, median, mode, drop) and provide code snippets (Python/pandas) to implement them.
- Inconsistent formats: propose standardization rules (e.g., date formats, string capitalization) and code to apply them.
- Duplicate records: recommend deduplication logic (e.g., based on key columns, fuzzy matching) and code.
- Redundant columns: outline criteria for identifying and removing columns (low variance, high correlation, missing data threshold) and automation code.
- Provide a complete workflow combining all steps, including error handling.
- Suggest validation checks to ensure data quality after cleaning.
Output format — A step-by-step guide with code snippets (preferably Python/pandas) for each issue, followed by a consolidated workflow. Include explanations of each step. Tone: instructional and practical.
Guardrails — Do not access or process actual data; provide code that the user can run locally. Flag any assumptions about column names or data types. Recommend testing on a sample before full automation.
Example — {{dataset_description}} = "sales transactions with missing customer IDs and inconsistent date formats", {{specific_issues}} = "missing values in 'customer_id', dates in 'MM/DD/YYYY' and 'YYYY-MM-DD'", {{data_format}} = "CSV with columns: transaction_id, customer_id, date, amount"
Follow-up prompts
- What common pitfalls should we avoid during data cleaning, especially with large datasets?
- How can we validate the quality of our cleaned data to ensure it meets analysis standards?
- What tools (beyond pandas) can streamline the data cleaning process further, especially for real-time or streaming data?