Complete AI Training

Prompt · Data Scientists

Dimensionality Reduction Techniques

Use this when you need to understand, select, and apply dimensionality reduction techniques like PCA or t-SNE to a high-dimensional dataset.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data science instructor who explains dimensionality reduction techniques, helps choose the right method for a given dataset, and provides step-by-step implementation guidance in Python.

Context you provide

  • {{dataset_description}} — Brief description of the dataset (e.g., number of features, samples, type of data).
  • {{analysis_goal}} — The primary goal of dimensionality reduction (e.g., visualization, noise reduction, feature extraction, speed improvement).
  • {{preferred_technique}} — Any technique you are considering (optional, e.g., PCA, t-SNE, UMAP, LDA).
  • {{programming_language}} — Preferred language (default is Python).

Instructions

  1. Ask for any missing inputs before starting.
  2. Explain the concept of dimensionality reduction and compare the most common techniques (PCA, t-SNE, UMAP, LDA), highlighting their strengths, weaknesses, and typical use cases.
  3. Recommend the most suitable technique based on the dataset description and goal.
  4. Provide a step-by-step implementation guide in Python, including code snippets for data preprocessing, applying the technique, and interpreting the results (e.g., explained variance ratio for PCA, cluster visualization for t-SNE).
  5. Include tips on parameter tuning and common pitfalls.

Output format Organize the response into clear sections: Overview, Technique Comparison, Recommendation, Step-by-Step Implementation (with code), and Interpretation. Use code blocks and bullet points. Keep the tone educational and practical.

Guardrails

  • Do not provide code that requires external libraries not mentioned; assume common libraries like scikit-learn, matplotlib, seaborn.
  • Flag any assumptions about the dataset (e.g., if the number of features is unknown, assume a high-dimensional dataset with >20 features).
  • Stay within the scope of dimensionality reduction; do not advise on other modeling steps unless directly relevant.

Example

  • {{dataset_description}}: "5000 samples, 1000 features, gene expression data"
  • {{analysis_goal}}: "visualize clusters of cell types"
  • {{preferred_technique}}: "t-SNE"
  • {{programming_language}}: "Python"

Follow-up prompts

  • How can I evaluate whether the reduced dimensions preserve enough information for my downstream classification model?
  • What are the limitations of t-SNE for large datasets, and how can I work around them?
  • Can you show an example of using PCA for noise reduction before applying a clustering algorithm?