Complete AI Training

Prompt · Data Scientists

Model Compression Techniques

Use this when you need to reduce your model's size and computational requirements for deployment on resource-constrained devices.

All 11 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in model compression and efficient AI deployment. Your goal is to help me reduce my model's size and inference cost while maintaining acceptable performance.

Context you provide

  • {{model_description}}: A description of my model, including architecture, size, and task.
  • {{deployment_target}}: The target environment (e.g., edge device, mobile, web) and its constraints (memory, compute, battery).
  • {{compression_goals}}: My specific goals (e.g., reduce size by 50%, speed up inference, maintain accuracy).
  • {{current_performance}}: The current performance metrics (e.g., accuracy, latency) to compare against.

Instructions

  1. Ask me for any missing context before starting.
  2. Provide an overview of the most suitable compression techniques for my model and deployment target, such as pruning, quantization, or knowledge distillation.
  3. For each technique, explain how it works, its potential benefits, and any trade-offs (e.g., accuracy loss, complexity).
  4. Recommend a practical approach, including which techniques to combine and in what order.
  5. Suggest metrics to monitor during compression to ensure the model remains effective.

Output format Structure your response with sections: Overview, Recommended Techniques, Implementation Plan, and Evaluation Metrics. Use bullet points and keep explanations clear.

Guardrails

  • Do not promise specific compression ratios or performance retention; instead, explain how to measure them.
  • Flag any assumptions about my model's architecture or framework.
  • Stay focused on compression; avoid general model optimization advice.

Example Model: ResNet-50 for image classification, deployment target: mobile phone with 1GB RAM, goals: reduce size by 75%, maintain >90% accuracy.

Follow-up prompts

  • How do I decide between pruning and quantization for my model?
  • What is the best way to evaluate a compressed model's performance?
  • Can you explain the process of knowledge distillation in more detail?