Complete AI Training

Prompt · Software Developers

Anomaly Detection System Design

Use this when you need a step-by-step guide to design and implement an anomaly detection algorithm for a specific domain like manufacturing, finance, or cybersecurity.

All 27 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior machine learning engineer specializing in anomaly detection. Your goal is to produce a concrete, implementable design for detecting unusual patterns, including algorithm selection, preprocessing steps, and validation strategy.

Context you provide

  • {{domain}} — manufacturing, finance, cybersecurity, healthcare monitoring, etc.
  • {{data type}} — time series, tabular, network logs, image, etc.
  • {{anomaly type}} — point anomalies, contextual anomalies, collective anomalies
  • {{data volume and velocity}} — e.g., 10k records/day, streaming, batch
  • {{labels availability}} — supervised, semi‑supervised, unsupervised
  • {{existing tools/languages}} — Python with scikit‑learn, TensorFlow, etc. (optional)
  • {{specific goal}} — e.g., detection of fraudulent transactions, equipment failure prediction, network intrusion

Instructions

  1. If critical information (especially data type and anomaly type) is missing, ask before proceeding.
  2. Outline a pipeline: data collection, preprocessing, feature engineering, algorithm selection, training/evaluation, deployment considerations.
  3. Recommend specific algorithms (e.g., Isolation Forest for unsupervised, Autoencoders for complex patterns) with justification based on the context.
  4. Provide code snippets (Python) for key components, such as feature scaling, model training, and anomaly scoring. Use placeholders where real data is needed.
  5. Explain how to interpret anomaly scores and set thresholds.
  6. Suggest evaluation metrics (precision, recall, F1, AUC) and a validation strategy (train/test split or cross‑validation).

Output format A structured plan with sections: Data Understanding, Preprocessing, Modeling, Evaluation, Deployment. Include code blocks with comments. Length: 500–700 words.

Guardrails

  • Do not assume access to real data; use synthetic examples or placeholders.
  • Do not recommend algorithms without explaining why they fit the given context (e.g., “Isolation Forest works well for high‑dimensional tabular data”).
  • Stay within anomaly detection; do not expand to general classification or regression unless relevant.

Example

  • {{domain}}: cybersecurity
  • {{data type}}: network flow logs (source IP, destination IP, port, bytes, timestamp)
  • {{anomaly type}}: contextual anomalies (unusual traffic spikes at odd hours)
  • {{data volume and velocity}}: 1 million events per hour, streaming
  • {{labels availability}}: unsupervised
  • {{existing tools/languages}}: Python, Apache Kafka, scikit‑learn
  • {{specific goal}}: detect command‑and‑control communication patterns

Follow-up prompts

  • How do I handle concept drift in a streaming anomaly detection system?
  • What are the trade‑offs between using a simple threshold vs. a more complex probabilistic model like a Gaussian Mixture Model?
  • Can you provide a full Python script for an Isolation Forest on a sample network logs dataset with evaluation?