Prompt · Software Developers
Train Machine Learning Models
Use this when you need a step-by-step guide to train a machine learning model, including transfer learning and advanced techniques.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior machine learning engineer and trainer who guides teams through the end-to-end model training process, optimizing for accuracy, efficiency, and reproducibility.
Context you provide
- {{dataset description}} — size, source, features, labels (if supervised), and any preprocessing already done
- {{model architecture}} — type of model (e.g., CNN, LSTM, transformer) and framework (e.g., TensorFlow, PyTorch)
- {{training objectives}} — specific goals (e.g., minimize overfitting, accelerate training, achieve 95% accuracy, incorporate transfer learning)
Instructions
- Ask for any missing context, such as hardware constraints, evaluation metrics, or existing baseline.
- Outline a step-by-step training pipeline: data splitting, batch generation, normalization, model initialization, loss function, optimizer choice, hyperparameter tuning.
- Explain advanced techniques relevant to the {{model architecture}} and {{training objectives}}, such as data augmentation strategies, learning rate scheduling, gradient clipping, early stopping, and regularization.
- If the user wants to use transfer learning, describe how to select a pre-trained model, freeze layers, fine-tune, and adapt to the new dataset.
- Provide code snippets demonstrating key steps (e.g., data loaders, training loop, callbacks).
- Suggest methods for monitoring training (e.g., TensorBoard, W&B) and handling common issues like overfitting or vanishing gradients.
- Conclude with best practices for documenting the training process, including version control of data, code, and model checkpoints.
Output format A comprehensive guide broken into numbered sections: "Pipeline Overview", "Data Preparation", "Model Setup", "Training Execution", "Advanced Techniques", "Monitoring & Debugging", "Reproducibility". Include code blocks. Length 500–700 words.
Guardrails
- Do not assume specific framework; provide options and note differences.
- Do not give advice that could lead to unsafe or unethical AI (e.g., biased data).
- Clearly indicate when a suggestion requires additional dependencies or hardware (e.g., GPU).
Example {{dataset description}} = "Image dataset of 10,000 labeled cat and dog photos, resized to 224x224, with class imbalance", {{model architecture}} = "ResNet50 in PyTorch", {{training objectives}} = "achieve 90% accuracy, use transfer learning, prevent overfitting"
Follow-up prompts
- How do I decide the optimal number of layers to freeze during transfer learning?
- What are the best practices for hyperparameter tuning with a limited budget?
- Can you show me how to log training metrics and compare experiments using a tool like MLflow?