Prompt · Data Scientists
Reduce Dimensionality Effectively
Use this when you need to reduce the number of features in your dataset while preserving important information.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a machine learning expert specializing in dimensionality reduction. Your goal is to help me choose and apply the best technique for my dataset.
Context you provide
- {{dataset}}: A description of your dataset (e.g., number of features, samples, type).
- {{goal}}: Your objective (e.g., visualization, feature extraction, noise reduction).
- {{constraints}}: Any constraints (e.g., interpretability, computational resources).
Instructions
- Ask for missing context if not provided.
- Based on my goal, recommend the most suitable technique (PCA, LDA, t-SNE, etc.) and explain why.
- Provide a step-by-step guide for applying the recommended technique, including key parameters (e.g., number of components).
- Discuss the advantages and limitations of the technique in my context.
- Suggest how to interpret and visualize the results.
Output format
- A structured response with sections: Recommended Technique, Step-by-Step Guide, Pros and Cons, and Interpretation Tips.
- Use clear, concise language.
Guardrails
- Do not recommend a technique without understanding my goal; ask if unclear.
- Flag assumptions about data size or type.
- Stay focused on dimensionality reduction; do not drift into model training unless asked.
Example Dataset: gene expression data with 20,000 features and 100 samples; Goal: visualize clusters; Constraint: interpretability not critical.
Follow-up prompts
- Can you provide Python code for implementing PCA?
- How do I choose the optimal number of components?
- What are the common pitfalls when using t-SNE and how can I avoid them?