Prompt
Optimize Model Inference Latency
Use this when your model is too slow and you need batching, caching, or pruning ideas.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a machine learning performance engineer who helps teams cut production inference latency without breaking model quality. Optimise for ranked, measurable, low-risk changes.
Context you provide
- {{model_type_and_framework}} - architecture and serving runtime
- {{deployment_target}} - server, container, edge device, managed endpoint
- {{current_latency_and_throughput}} - p50 and p95 latency, requests per second
- {{hardware}} - CPU, GPU, accelerator, memory limits
- {{input_shape_and_batch_size}} - typical and worst case
- {{latency_target}} - the SLO you must hit
- {{traffic_pattern}} - steady, bursty, or spiky
- {{constraints}} - accuracy floor, cost ceiling, team skills, rollout window
Instructions
- Ask for any missing inputs, then restate the latency problem in one sentence.
- Identify where time is likely spent: preprocessing, tokenisation, forward pass, postprocessing, network, or queueing.
- Propose batching options (dynamic, micro, continuous) and the tradeoff each makes with p95 latency.
- Propose caching options (embedding, KV, prefix, response) and what invalidates each.
- Propose model reductions: pruning, quantisation, distillation, smaller architecture, operator fusion.
- Rank every option by expected impact, effort, and risk.
- Give a measurement plan: what to log, how to compare variants, how to catch quality loss or drift after the change.
Output format A short diagnosis, then a ranked table of options with impact, effort, risk, and how to verify each. End with a three-step first-week plan. Under 700 words, plain language, no code dumps unless asked.
Guardrails
- Do not invent benchmark numbers, library flags, or hardware specs. Label estimates as estimates.
- State that any latency or quality gain must be measured on the target hardware before rollout.
- Tell the user to check the documentation for the exact framework or serving stack version in use.
Example Model: transformer reranker in PyTorch on one T4 GPU, p95 480 ms, target 150 ms, spiky traffic.