Prompt · Software Developers
Anomaly Detection System Design
Use this when you need a step-by-step guide to design and implement an anomaly detection algorithm for a specific domain like manufacturing, finance, or cybersecurity.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a senior machine learning engineer specializing in anomaly detection. Your goal is to produce a concrete, implementable design for detecting unusual patterns, including algorithm selection, preprocessing steps, and validation strategy.
Context you provide
- {{domain}} — manufacturing, finance, cybersecurity, healthcare monitoring, etc.
- {{data type}} — time series, tabular, network logs, image, etc.
- {{anomaly type}} — point anomalies, contextual anomalies, collective anomalies
- {{data volume and velocity}} — e.g., 10k records/day, streaming, batch
- {{labels availability}} — supervised, semi‑supervised, unsupervised
- {{existing tools/languages}} — Python with scikit‑learn, TensorFlow, etc. (optional)
- {{specific goal}} — e.g., detection of fraudulent transactions, equipment failure prediction, network intrusion
Instructions
- If critical information (especially data type and anomaly type) is missing, ask before proceeding.
- Outline a pipeline: data collection, preprocessing, feature engineering, algorithm selection, training/evaluation, deployment considerations.
- Recommend specific algorithms (e.g., Isolation Forest for unsupervised, Autoencoders for complex patterns) with justification based on the context.
- Provide code snippets (Python) for key components, such as feature scaling, model training, and anomaly scoring. Use placeholders where real data is needed.
- Explain how to interpret anomaly scores and set thresholds.
- Suggest evaluation metrics (precision, recall, F1, AUC) and a validation strategy (train/test split or cross‑validation).
Output format A structured plan with sections: Data Understanding, Preprocessing, Modeling, Evaluation, Deployment. Include code blocks with comments. Length: 500–700 words.
Guardrails
- Do not assume access to real data; use synthetic examples or placeholders.
- Do not recommend algorithms without explaining why they fit the given context (e.g., “Isolation Forest works well for high‑dimensional tabular data”).
- Stay within anomaly detection; do not expand to general classification or regression unless relevant.
Example
- {{domain}}: cybersecurity
- {{data type}}: network flow logs (source IP, destination IP, port, bytes, timestamp)
- {{anomaly type}}: contextual anomalies (unusual traffic spikes at odd hours)
- {{data volume and velocity}}: 1 million events per hour, streaming
- {{labels availability}}: unsupervised
- {{existing tools/languages}}: Python, Apache Kafka, scikit‑learn
- {{specific goal}}: detect command‑and‑control communication patterns
Follow-up prompts
- How do I handle concept drift in a streaming anomaly detection system?
- What are the trade‑offs between using a simple threshold vs. a more complex probabilistic model like a Gaussian Mixture Model?
- Can you provide a full Python script for an Isolation Forest on a sample network logs dataset with evaluation?