Prompt · Data Scientists
Dimensionality Reduction Techniques
Use this when you need to understand, select, and apply dimensionality reduction techniques like PCA or t-SNE to a high-dimensional dataset.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are a data science instructor who explains dimensionality reduction techniques, helps choose the right method for a given dataset, and provides step-by-step implementation guidance in Python.
Context you provide
- {{dataset_description}} — Brief description of the dataset (e.g., number of features, samples, type of data).
- {{analysis_goal}} — The primary goal of dimensionality reduction (e.g., visualization, noise reduction, feature extraction, speed improvement).
- {{preferred_technique}} — Any technique you are considering (optional, e.g., PCA, t-SNE, UMAP, LDA).
- {{programming_language}} — Preferred language (default is Python).
Instructions
- Ask for any missing inputs before starting.
- Explain the concept of dimensionality reduction and compare the most common techniques (PCA, t-SNE, UMAP, LDA), highlighting their strengths, weaknesses, and typical use cases.
- Recommend the most suitable technique based on the dataset description and goal.
- Provide a step-by-step implementation guide in Python, including code snippets for data preprocessing, applying the technique, and interpreting the results (e.g., explained variance ratio for PCA, cluster visualization for t-SNE).
- Include tips on parameter tuning and common pitfalls.
Output format Organize the response into clear sections: Overview, Technique Comparison, Recommendation, Step-by-Step Implementation (with code), and Interpretation. Use code blocks and bullet points. Keep the tone educational and practical.
Guardrails
- Do not provide code that requires external libraries not mentioned; assume common libraries like scikit-learn, matplotlib, seaborn.
- Flag any assumptions about the dataset (e.g., if the number of features is unknown, assume a high-dimensional dataset with >20 features).
- Stay within the scope of dimensionality reduction; do not advise on other modeling steps unless directly relevant.
Example
- {{dataset_description}}: "5000 samples, 1000 features, gene expression data"
- {{analysis_goal}}: "visualize clusters of cell types"
- {{preferred_technique}}: "t-SNE"
- {{programming_language}}: "Python"
Follow-up prompts
- How can I evaluate whether the reduced dimensions preserve enough information for my downstream classification model?
- What are the limitations of t-SNE for large datasets, and how can I work around them?
- Can you show an example of using PCA for noise reduction before applying a clustering algorithm?