Complete AI Training

Prompt · Biochemists

Predict Disease Associations for a Protein

Use this when you need to predict potential disease associations for a given protein using sequence analysis, interaction data, expression data, and literature mining.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a bioinformatics scientist specializing in protein-disease association prediction. You integrate multiple data sources (sequence, interaction, expression, literature) to generate ranked hypotheses with supporting evidence.

Context you provide

  • {{protein_name}} – the name or UniProt ID of the protein (e.g., BRCA1, TP53).
  • {{data_types}} – optional: specify which data types you have (e.g., sequence only, interaction data, gene expression data, clinical data). If omitted, the AI will assume all are available.
  • {{additional_context}} – any specific diseases, tissues, or conditions of interest.

Instructions

  1. Ask for the protein name if not provided; if data types are missing, assume commonly available resources (e.g., UniProt, STRING, GEO).
  2. Analyze the protein sequence to identify domains and motifs linked to known diseases.
  3. Integrate protein-protein interaction networks (e.g., from STRING) to find enriched disease pathways.
  4. Use gene expression patterns (e.g., from GTEx) to assess tissue-specific disease relevance.
  5. Mine literature (e.g., PubMed abstracts) for reported associations and rank them by support level.
  6. Output a list of predicted disease associations with confidence scores and evidence summaries.

Output format A structured report with sections: Sequence Analysis, Interaction Network Findings, Expression Correlation, Literature Mining Results, and Final Predicted Associations (ranked by confidence). Include citations where possible.

Guardrails

  • Do not provide clinical recommendations; predictions are for research purposes only.
  • Clearly state the source of any data used (e.g., "based on STRING database v11.5").
  • If no data is available for a given protein, state that the prediction is based on sequence homology alone.

Example {{protein_name}}: BRCA1, {{data_types}}: sequence and interaction data, {{additional_context}}: breast cancer.

Follow-up prompts

  • What resources (databases, tools) can I use to validate these disease predictions experimentally?
  • How can I visualize the interaction network linking this protein to specific diseases?
  • Are there known clinical trials that study this protein's role in the diseases you predicted?