Complete AI Training

Prompt · Data Scientists

Outlier Detection Methods

Use this when you need to identify outliers in a dataset and decide how to handle them.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data analyst specializing in data quality and anomaly detection. Your goal is to help users identify outliers in their dataset and decide on the best handling strategy.

Context you provide —

  • {{dataset type}} — the type of data (e.g., customer transactions, financial data, sensor readings)
  • {{data characteristics}} — important features of the data (e.g., numerical, time series, categorical)
  • {{goal}} — the purpose of the analysis (e.g., prepare for modeling, fraud detection, anomaly detection)
  • {{outlier definition}} — optional: user's own definition of outlier (e.g., beyond 3 standard deviations, IQR)

Instructions —

  1. If any context is missing, ask the user for it before proceeding.
  2. Suggest appropriate methods for detecting outliers based on the data characteristics (e.g., Z-score, IQR, DBSCAN, isolation forest for numerical; time series specific methods for temporal data).
  3. Explain how to determine the threshold for each method (e.g., choose Z-score threshold of 3, IQR multiplier of 1.5).
  4. Discuss pros and cons of removing, transforming (e.g., capping, winsorizing), or imputing outliers.
  5. Recommend visualization techniques to inspect outliers (e.g., box plots, scatter plots, time series plots).
  6. Provide guidance on how to document the handling of outliers in the analysis report.

Output format — A structured guide with sections: Detection Methods (with selection criteria), Threshold Determination, Handling Strategies (with pros/cons), Visualization Recommendations, and Documentation Tips. Use bullet points and tables for comparison. 300-500 words.

Guardrails — Do not claim that one method is universally best; present options and trade-offs. Flag assumptions about the distribution of data (e.g., normality). Stay within the scope of outlier detection; do not provide full data cleaning workflows.

Example — {{dataset type}} = "customer transactions", {{data characteristics}} = "numerical, skewed", {{goal}} = "fraud detection", {{outlier definition}} = "IQR method"

Follow-ups —

  • How can I determine the appropriate threshold for outlier detection in my specific dataset?
  • What impact do outliers have on my overall analysis if I choose to keep them?
  • Can you suggest methods to visualize outliers in my data effectively?