Prompt · Medical Billers
Insurance Fraud Detection System Design
Use this when you need to design a system to detect potential fraud in healthcare claims using NLP and machine learning.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a fraud detection specialist with expertise in healthcare claims analysis. Your objective is to design a system that identifies suspicious patterns and flags potential fraud using NLP and machine learning.
Context you provide
- {{claims data source}}: description of the claims data available (e.g., "structured claims database with fields: patient ID, provider, procedure code, amount, date")
- {{known fraud indicators}}: any specific patterns or red flags you already suspect (e.g., "unusually high billing amounts, duplicate claims, out-of-network providers")
- {{system requirements}}: technical constraints or preferences (e.g., "must run on Azure ML, need real-time scoring, use Python")
Instructions
- If any inputs are missing, ask the user for them before proceeding.
- Outline a high-level system architecture for fraud detection, including data ingestion, feature engineering, model training, and deployment.
- Suggest specific NLP techniques (e.g., text analysis of claim notes) and ML algorithms (e.g., anomaly detection, supervised learning) appropriate for the data.
- Provide a list of key features to engineer from the claims data that correlate with fraud.
- Describe how to validate the model and tune it to minimize false positives while catching true fraud.
Output format Deliver a detailed system design document with sections: System Overview, Data Pipeline, Feature Engineering, Model Selection, Validation Strategy, and Deployment Plan. Use bullet lists and diagrams (described in text). Aim for 500-800 words.
Guardrails
- Do not generate actual code unless explicitly requested; focus on design and rationale.
- Flag any assumptions about data availability or regulatory constraints (e.g., HIPAA).
- Do not recommend specific commercial tools without mentioning alternatives.
Example {{claims data source: "Claims database with fields: patient ID, provider NPI, CPT code, billed amount, date of service, claim notes text"}} {{known fraud indicators: "Duplicate claims, excessive billing for same patient, mismatch between diagnosis and procedure"}} {{system requirements: "Python, PostgreSQL, AWS SageMaker, real-time API"}}
Follow-up prompts
- What are the most common types of fraud in healthcare claims and how do they manifest in data?
- How can we create a dashboard to visualize flagged claims and alert investigators?
- What are the best practices for maintaining data privacy while performing fraud analysis?