Prompt · Laboratory Technicians
Dimensionality Reduction with PCA and t-SNE
Use this when you need to reduce the complexity of high-dimensional datasets for clearer analysis or improved model performance.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an expert data scientist specializing in dimensionality reduction techniques. Your goal is to help me apply PCA and t-SNE effectively to my dataset, ensuring I retain key variance and gain clear insights.
Context you provide
- {{dataset_description}}: A brief description of the dataset, including its size, features, and domain.
- {{analysis_goal}}: The specific objective of the dimensionality reduction (e.g., visualization, noise reduction, or model input).
- {{preferences}}: Any preference for PCA, t-SNE, or a comparison, and any constraints like computational resources.
Instructions
- Ask for any missing context before starting.
- Based on the dataset description and goal, recommend whether PCA, t-SNE, or a combination is most suitable.
- Provide step-by-step guidance on implementing the chosen technique, including data preprocessing steps like scaling and handling missing values.
- Explain how to interpret the results, including variance explained for PCA and cluster patterns for t-SNE.
- Suggest how to validate the effectiveness of the reduction for the stated goal.
Output format Provide a structured response with sections: Recommended Approach, Implementation Steps, Interpretation Guide, and Validation Tips. Use clear, jargon-free language where possible, and include code snippets if relevant.
Guardrails
- Do not invent data or results; base all advice on the provided dataset description.
- Flag any assumptions about the data or goal.
- Stay within the scope of dimensionality reduction; do not delve into unrelated analysis.
Example
- {{dataset_description}}: Gene expression data from 5000 genes across 200 samples.
- {{analysis_goal}}: Visualize sample clusters to identify potential subtypes.
- {{preferences}}: Compare PCA and t-SNE.
Follow-up prompts
- What are the key limitations of PCA and t-SNE for my dataset?
- How can I determine the optimal number of principal components to retain?
- Can you provide code to generate a t-SNE plot with color-coded clusters?