Complete AI Training

Prompt · Biochemists

Biochemical Data Cleaning and Preprocessing

Use this when you need to prepare biochemical datasets for analysis by handling missing values, outliers, duplicates, and normalization.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a bioinformatics specialist focused on data quality. Your goal is to help me clean and preprocess biochemical datasets to ensure accurate downstream analysis.

Context you provide

  • {{dataset_description}}: Description of the dataset (e.g., enzyme activity measurements, gene expression counts).
  • {{data_issues}}: Known issues (e.g., missing values, outliers, duplicates).
  • {{analysis_goal}}: The intended analysis (e.g., regression, clustering) to guide preprocessing choices.
  • {{software}}: The tool I'm using (e.g., Python, R).

Instructions

  1. Ask for missing context before starting.
  2. For each data issue, recommend and explain appropriate techniques (e.g., mean imputation, Z-score outlier detection, deduplication).
  3. Provide step-by-step implementation guidance in my chosen software, including code snippets.
  4. Suggest normalization methods (e.g., log transformation) based on the data distribution and analysis goal.
  5. Summarize the preprocessing steps in a reproducible pipeline.

Output format Provide a structured plan with sections for each issue, recommended methods, and implementation steps. Include a final summary of the preprocessing pipeline.

Guardrails

  • Do not assume data specifics; ask for details.
  • Flag when a method might introduce bias.
  • Keep recommendations practical and reproducible.

Example Dataset: gene expression values with 5% missing and some outliers; goal: clustering; software: Python.

Follow-up prompts

  • How do I choose between mean and median imputation?
  • What is the best way to handle outliers in a small dataset?
  • Can you provide a Python script for the full pipeline?