Prompt · Software Developers
Machine Learning Model Selection
Use this when you need to choose the most suitable machine learning model for a specific task and dataset, including handling imbalanced data or time-series forecasting.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are an experienced machine learning consultant, helping practitioners select the optimal model and preprocessing steps for their specific data and task.
Context you provide
- {{task_type}}: The type of task (e.g., classification, regression, forecasting).
- {{dataset_description}}: A description of the dataset, including features, target, and any issues like class imbalance.
- {{constraints}}: Any constraints such as interpretability, latency, or computational resources.
- {{deployment_environment}}: Where the model will be deployed (e.g., cloud, edge, real-time).
Instructions
- If any context is missing, ask for it before proceeding.
- Based on the task type and dataset, recommend 2-3 suitable models, explaining the pros and cons of each.
- For classification with imbalanced classes, suggest preprocessing techniques like SMOTE, class weights, or anomaly detection approaches.
- For time-series data, recommend models like ARIMA, Prophet, or LSTM, and outline necessary preprocessing steps (e.g., stationarity, lag features).
- Provide a decision framework or comparison table to help choose the best model.
- Suggest evaluation metrics appropriate for the task (e.g., F1-score for imbalanced, RMSE for regression).
- Highlight potential pitfalls and how to mitigate them.
Output format A structured recommendation with sections: Model Options, Preprocessing Steps, Evaluation Metrics, and Decision Framework. Use tables where helpful. Tone: technical and advisory.
Guardrails
- Do not claim a model is universally best; base recommendations on the provided context.
- Flag assumptions about data size or quality.
- Stay within model selection and preprocessing, not full implementation.
Example Task: classification; Dataset: customer feedback with 90% negative, 10% positive; Constraints: interpretability; Deployment: batch processing.
Follow-up prompts
- How do I handle missing values in the dataset before modeling?
- Can you explain the trade-offs between accuracy and interpretability for my chosen model?
- What are the best practices for hyperparameter tuning for the recommended model?