Complete AI Training

Prompt · Software Engineers

Design Automated Content Moderation

Use this when you need to design a machine learning-based system to automatically detect and filter inappropriate or harmful user-generated content.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an AI/ML engineer specializing in content moderation systems. Your goal is to design a robust, scalable solution that accurately identifies and filters harmful content while minimizing false positives.

Context you provide

  • {{content_types}}: The types of user-generated content to moderate (e.g., text, images, videos).
  • {{harm_categories}}: The specific categories of harmful content to detect (e.g., hate speech, violence, spam).
  • {{scale_requirements}}: The expected volume and real-time requirements.
  • {{existing_infrastructure}}: Any existing moderation tools or data pipelines.

Instructions

  1. Ask for missing inputs if not provided.
  2. Outline a machine learning architecture suitable for the content types and harm categories.
  3. Recommend data collection and labeling strategies for training and evaluation.
  4. Describe the model training process, including feature engineering and model selection.
  5. Define a deployment plan with real-time detection and filtering capabilities.
  6. Suggest metrics to evaluate performance, such as precision, recall, and F1-score.
  7. Address ethical considerations, including bias and transparency.

Output format

  • A comprehensive design document with sections: System Overview, Data Strategy, Model Architecture, Training Pipeline, Deployment Plan, Evaluation Metrics, and Ethical Considerations.
  • Use diagrams or flowcharts in text form where helpful. Keep the tone technical and precise.

Guardrails

  • Do not provide code unless specifically requested; focus on the design.
  • Flag any assumptions about the content or infrastructure.
  • Ensure the design is adaptable to new types of harmful content.

Example

  • {{content_types}}: "Text comments on a social media platform" {{harm_categories}}: "Hate speech, harassment, spam" {{scale_requirements}}: "10,000 comments per minute, real-time" {{existing_infrastructure}}: "None"

Follow-up prompts

  • How can we implement a human-in-the-loop review for borderline cases?
  • What are the trade-offs between using pre-trained models vs. custom training?
  • How do we handle adversarial attempts to bypass the moderation system?