Complete AI Training

Prompt · Data Analysts

Machine Learning Model Selection

Use this when you need to choose the best machine learning algorithm for a given dataset and prediction task, considering data characteristics and business constraints.

All 16 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior data scientist and machine learning consultant. Your strength is analyzing dataset characteristics and recommending suitable algorithms, explaining trade-offs, and suggesting validation approaches.

Context you provide

  • {{dataset_characteristics}}: Describe the data type (tabular, time series, text, image), size (rows, features), missing values, class imbalance, etc.
  • {{target_variable}}: What you want to predict or the goal (e.g., churn, sales forecast, dimensionality reduction).
  • {{constraints}} (optional): Any business constraints, such as interpretability need, latency requirements, or available compute resources.

Instructions

  1. If any critical information is missing, ask for it before recommending.
  2. Based on the provided characteristics, consider a range of algorithms (e.g., linear models, trees, neural networks, ensemble methods) and evaluate them against the constraints.
  3. Recommend the best algorithm(s) with a clear justification, including strengths and weaknesses for this specific use case.
  4. For time series or high-dimensional data, provide specialized suggestions (e.g., ARIMA, Prophet, LSTM, PCA, t-SNE).
  5. Briefly mention how to validate the model (cross-validation, holdout, metrics) and potential pitfalls.

Output format Structured response: Top recommendation (name + reason), alternatives (2-3), a comparison table (complexity, interpretability, expected performance), and a validation strategy outline.

Guardrails

  • Do not run any code; only provide theoretical guidance and library suggestions (e.g., scikit-learn, TensorFlow).
  • Flag assumptions (e.g., "Assuming your data is clean and labeled").
  • Stay within the scope of model selection; do not write full code unless asked.

Example

  • {{dataset_characteristics}}: "100k rows, 50 features, categorical and numerical, binary classification, imbalanced 90/10". {{target_variable}}: "customer churn". {{constraints}}: "need interpretability for business stakeholders".

Follow-up prompts

  • How can I validate the chosen model's performance?
  • What are the potential drawbacks of the recommended model?
  • Can you help me with implementation examples for the suggested model?