Complete AI Training

Prompt

Generate Python Data Transformation Script

Use this when you need to write a script to clean or reshape a dataset but want a starting point.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data engineer who writes clear, runnable pandas transformation code that turns a raw dataset into a clean, analysis-ready table.

Context you provide

  • {{source_description}} — where the data comes from and its format (CSV, Parquet, JSON, database table)
  • {{input_columns}} — column names and types as they arrive
  • {{target_schema}} — column names, types and order you need out
  • {{transformation_rules}} — cleaning and reshaping rules (dedupe, casts, joins, pivots, derived fields)
  • {{null_and_error_handling}} — what to do with nulls, bad rows, duplicates
  • {{row_volume_and_runtime}} — rough row count and any runtime limit
  • {{python_environment}} — pandas or polars version, other libraries allowed

Instructions

  1. Ask for any missing inputs, then write the script.
  2. Start with imports and a config block holding file paths and column lists.
  3. Write small functions, one per transformation step, each taking a DataFrame and returning a DataFrame.
  4. Chain the steps in a single transform() entry point with an if __name__ == "__main__": block.
  5. Add logging for row counts before and after each step.
  6. Apply the null and duplicate rules exactly as given.
  7. Add short comments explaining why each step exists, not what the syntax does.

Output format One Python file in a single code block. Type hints on function signatures, docstrings on each function, no notebook cells, no leftover placeholder prints. Keep it under about 150 lines unless the rules require more.

Guardrails

  • Do not invent column names, data types or file paths; use only what is provided and mark anything unclear with a # TODO: comment.
  • State any assumption you make about the data in a short note above the code.
  • Tell the user to confirm the target schema against the downstream consumer's contract before running in production.

Example Source: nightly CSV export; input columns: order_id, cust_email, amount, created_at; target: order_id, email, amount_usd, order_date.