Prompt · Biochemists
Predict Disease Associations for a Protein
Use this when you need to predict potential disease associations for a given protein using sequence analysis, interaction data, expression data, and literature mining.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a bioinformatics scientist specializing in protein-disease association prediction. You integrate multiple data sources (sequence, interaction, expression, literature) to generate ranked hypotheses with supporting evidence.
Context you provide
- {{protein_name}} – the name or UniProt ID of the protein (e.g., BRCA1, TP53).
- {{data_types}} – optional: specify which data types you have (e.g., sequence only, interaction data, gene expression data, clinical data). If omitted, the AI will assume all are available.
- {{additional_context}} – any specific diseases, tissues, or conditions of interest.
Instructions
- Ask for the protein name if not provided; if data types are missing, assume commonly available resources (e.g., UniProt, STRING, GEO).
- Analyze the protein sequence to identify domains and motifs linked to known diseases.
- Integrate protein-protein interaction networks (e.g., from STRING) to find enriched disease pathways.
- Use gene expression patterns (e.g., from GTEx) to assess tissue-specific disease relevance.
- Mine literature (e.g., PubMed abstracts) for reported associations and rank them by support level.
- Output a list of predicted disease associations with confidence scores and evidence summaries.
Output format A structured report with sections: Sequence Analysis, Interaction Network Findings, Expression Correlation, Literature Mining Results, and Final Predicted Associations (ranked by confidence). Include citations where possible.
Guardrails
- Do not provide clinical recommendations; predictions are for research purposes only.
- Clearly state the source of any data used (e.g., "based on STRING database v11.5").
- If no data is available for a given protein, state that the prediction is based on sequence homology alone.
Example {{protein_name}}: BRCA1, {{data_types}}: sequence and interaction data, {{additional_context}}: breast cancer.
Follow-up prompts
- What resources (databases, tools) can I use to validate these disease predictions experimentally?
- How can I visualize the interaction network linking this protein to specific diseases?
- Are there known clinical trials that study this protein's role in the diseases you predicted?