Complete AI Training

Prompt · CDOs (Chief Digital Officers)

Optimize AI Model Performance

Use this when you need to improve the speed, scalability, or resource efficiency of your AI models without sacrificing accuracy.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in AI model optimization, focused on improving inference speed, scalability, and resource efficiency while maintaining model accuracy.

Context you provide

  • {{model_details}}: Describe your AI model (type, framework, current performance metrics).
  • {{performance_goals}}: Specify what you want to improve (e.g., reduce latency, increase throughput, lower memory usage).
  • {{constraints}}: Mention any trade-offs you cannot accept (e.g., accuracy loss, cost limits).

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Analyze the provided model details and performance goals to identify potential bottlenecks.
  3. Recommend specific optimization techniques, such as quantization, pruning, or hardware acceleration, explaining how each impacts performance and accuracy.
  4. Prioritize recommendations based on effort, impact, and risk.
  5. Suggest methods for measuring the impact of each optimization.

Output format Provide a structured analysis with sections: Current Bottlenecks, Recommended Techniques (with expected benefits and risks), Implementation Steps, and Measurement Plan. Use bullet points and keep the tone technical and concise.

Guardrails

  • Do not invent performance metrics or benchmarks; use only provided data.
  • Flag any assumptions about the model or environment.
  • Stay within the scope of model optimization; do not suggest unrelated changes.

Example

  • {{model_details}}: "A BERT-based text classifier with 110M parameters, running on a single GPU, inference time 50ms per sample."
  • {{performance_goals}}: "Reduce inference time to under 20ms without losing more than 1% accuracy."
  • {{constraints}}: "Cannot increase hardware cost."

Follow-up prompts

  • What are the first three steps to implement quantization on my model?
  • How can I profile my model to identify the biggest bottlenecks?
  • What are the risks of pruning and how can I mitigate them?