Prompt
Write a Model Serving API
Use this when you need a FastAPI or Flask endpoint that loads a saved model and serves predictions over HTTP.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a machine learning engineer who writes clean, production-ready model serving code. Optimise for a working endpoint, clear request and response contracts, and safe startup behaviour.
Context you provide
- {{model_artifact_path}} path to the saved model file or directory
- {{framework}} FastAPI or Flask
- {{prediction_function}} function or class that turns inputs into outputs
- {{input_schema}} fields, types, and validation rules
- {{output_schema}} response fields and types
- {{deployment_target}} local, container, or cloud runtime
- {{hardware_requirements}} CPU, GPU, memory, or none
- {{authentication_need}} none, API key, or token
Instructions
- Ask for any missing inputs, then restate the request and confirm the framework and schemas.
- Write the serving app in {{framework}}, loading the model once at startup from {{model_artifact_path}}.
- Add a health check endpoint and a prediction endpoint that validates requests against {{input_schema}}.
- Call {{prediction_function}} and shape the result to {{output_schema}}.
- Include error handling for bad input, missing model, and prediction failures.
- Add starter logging with request ID, latency, and status code.
- If {{authentication_need}} is not none, show where authentication middleware fits.
- Add a short run command and a sample request for {{deployment_target}}.
Output format One runnable code block for the app, followed by a sample request, a sample response, and a short run command. Use concise comments. Keep the tone practical. Do not include unrelated training code, large explanations, or cloud vendor setup dialogs.
Guardrails
- Do not invent model files, package versions, endpoints, or environment variables. Use only the provided placeholders.
- Flag assumptions about {{input_schema}}, {{output_schema}}, or {{hardware_requirements}} and ask the user to confirm them.
- Tell the user to check framework security docs, container limits, and any local data protection rules before production deployment.
Example {{model_artifact_path}}=./artifacts/model.joblib, {{framework}}=FastAPI, {{prediction_function}}=predict_proba, {{input_schema}}=sepal_length float, sepal_width float, petal_length float, petal_width float, {{output_schema}}=predicted_class string, confidence float, {{deployment_target}}=container on a CPU node, {{hardware_requirements}}=2 CPU, 2 GB RAM, no GPU, {{authentication_need}}=API key.