Prompt · Clinical Data Managers
Aggregate Clinical Data Sources
Use this when you need to combine multiple clinical or healthcare data sources into a unified dataset for analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data integration specialist with expertise in clinical data management, optimizing for accuracy and completeness in merged datasets.
Context you provide
- {{source_names}}: List of data sources to aggregate (e.g., EHR systems, lab databases, research studies).
- {{data_types}}: Types of data to combine (e.g., patient demographics, lab results, adverse events, genetic data).
- {{merge_criteria}}: Key fields for matching records (e.g., patient ID, study ID).
Instructions
- If any of the above inputs are missing, ask for them before proceeding.
- Outline a step-by-step plan to aggregate the specified data types from the given sources, including how to handle different formats and structures.
- Provide a schema for the unified dataset, defining fields, data types, and relationships.
- Describe methods to ensure data consistency and quality during aggregation, such as validation rules and conflict resolution.
- Suggest tools or scripts (e.g., Python, SQL) that can automate the aggregation process.
Output format Provide a structured response with sections: Aggregation Plan, Unified Schema, Quality Assurance, and Recommended Tools. Use bullet points and tables where helpful. Keep the tone professional and technical.
Guardrails Do not invent specific data values or source details; use placeholders. Flag any assumptions about data availability or format. Stay within the scope of data aggregation, not analysis or cleaning.
Example Sources: 'Hospital A EHR', 'LabCorp results', 'Clinical trial DB'; Data types: 'demographics, lab results'; Merge criteria: 'Patient ID'.
Follow-up prompts
- What are the best practices for handling missing data during aggregation?
- How can I ensure the aggregated dataset complies with privacy regulations?
- Can you provide a sample Python script for merging these sources?