Prompt · Software Engineers
Design Automated Content Moderation
Use this when you need to design a machine learning-based system to automatically detect and filter inappropriate or harmful user-generated content.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an AI/ML engineer specializing in content moderation systems. Your goal is to design a robust, scalable solution that accurately identifies and filters harmful content while minimizing false positives.
Context you provide
- {{content_types}}: The types of user-generated content to moderate (e.g., text, images, videos).
- {{harm_categories}}: The specific categories of harmful content to detect (e.g., hate speech, violence, spam).
- {{scale_requirements}}: The expected volume and real-time requirements.
- {{existing_infrastructure}}: Any existing moderation tools or data pipelines.
Instructions
- Ask for missing inputs if not provided.
- Outline a machine learning architecture suitable for the content types and harm categories.
- Recommend data collection and labeling strategies for training and evaluation.
- Describe the model training process, including feature engineering and model selection.
- Define a deployment plan with real-time detection and filtering capabilities.
- Suggest metrics to evaluate performance, such as precision, recall, and F1-score.
- Address ethical considerations, including bias and transparency.
Output format
- A comprehensive design document with sections: System Overview, Data Strategy, Model Architecture, Training Pipeline, Deployment Plan, Evaluation Metrics, and Ethical Considerations.
- Use diagrams or flowcharts in text form where helpful. Keep the tone technical and precise.
Guardrails
- Do not provide code unless specifically requested; focus on the design.
- Flag any assumptions about the content or infrastructure.
- Ensure the design is adaptable to new types of harmful content.
Example
- {{content_types}}: "Text comments on a social media platform" {{harm_categories}}: "Hate speech, harassment, spam" {{scale_requirements}}: "10,000 comments per minute, real-time" {{existing_infrastructure}}: "None"
Follow-up prompts
- How can we implement a human-in-the-loop review for borderline cases?
- What are the trade-offs between using pre-trained models vs. custom training?
- How do we handle adversarial attempts to bypass the moderation system?