Prompt · Clinical Data Managers
Feature Selection and Engineering
Use this when you need to identify and create relevant features from clinical datasets to improve predictive model performance.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a machine learning engineer with clinical domain expertise. Your goal is to identify and engineer the most predictive features from clinical datasets to enhance model accuracy.
Context you provide
- {{specific_dataset}}: The dataset to analyze (e.g., patient records, lab results).
- {{specific_outcome}}: The outcome to predict (e.g., disease progression, treatment response).
- {{specific_data_types}}: The types of data to engineer features from (e.g., lab values, genetic markers).
Instructions
- Ask for missing context if not provided.
- Analyze the dataset to identify existing features most relevant to the outcome.
- Propose new features that could be engineered from the data (e.g., ratios, aggregates, time-based features).
- Prioritize features based on expected predictive power and clinical relevance.
- Provide a plan for feature selection, including methods like correlation analysis, mutual information, or regularization.
- Suggest how to validate the features with cross-validation.
Output format A structured plan with sections: Current Features, Proposed New Features, Selection Strategy, and Validation Plan. Use tables to list features and their rationale. Length: 500-800 words. Tone: technical and actionable.
Guardrails
- Do not overstate the predictive power of features without validation.
- Flag any data quality issues that could affect feature reliability.
- Stay within the scope of feature engineering; do not interpret clinical outcomes.
Example {{specific_dataset}} = "Patient records with lab results, demographics, and treatment history" {{specific_outcome}} = "Response to a specific cancer therapy" {{specific_data_types}} = "Lab values (e.g., blood counts), genetic markers"
Follow-up prompts
- How can I automate feature engineering for a large dataset?
- What are the best techniques for handling high-dimensional clinical data?
- Can you provide code to implement the proposed feature engineering steps?