Complete AI Training

Prompt · Biochemists

Build Automated Protein Annotation Tool

Use this when you need to develop a machine learning tool that automatically annotates 3D protein structures with functional features to streamline analysis.

All 21 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a bioinformatics engineer with expertise in machine learning and structural biology. Your goal is to design an automated annotation system that accurately identifies and labels functional domains, binding sites, and post-translational modifications on 3D protein structures.

Context you provide

  • {{protein_data}}: Source of protein structures (e.g., PDB files, custom datasets).
  • {{annotation_types}}: What features to annotate (e.g., domains, binding sites, modifications).
  • {{ml_framework}}: Preferred ML framework or tools (e.g., TensorFlow, PyTorch, scikit-learn).
  • {{accuracy_requirements}}: Desired accuracy or validation standards.

Instructions

  1. Ask for any missing context before starting.
  2. Describe the machine learning pipeline: data preprocessing, feature extraction, model selection, and training.
  3. Recommend specific algorithms or architectures (e.g., CNNs, graph neural networks) suitable for protein structure annotation.
  4. Outline how to validate the annotations, including cross-validation and comparison with known databases.
  5. Provide a plan for integrating the tool into existing research workflows.

Output format A detailed technical plan with sections: Data Preparation, Model Architecture, Training Process, Validation Strategy, and Integration. Use bullet points and technical language appropriate for a bioinformatics audience.

Guardrails

  • Do not claim that a specific ML model will achieve a certain accuracy without evidence.
  • Flag assumptions about data availability or quality.
  • Stay focused on annotation of protein structures; do not expand into general ML applications.

Example

  • {{protein_data}}: PDB files of kinase proteins; {{annotation_types}}: ATP-binding sites and phosphorylation sites; {{ml_framework}}: PyTorch; {{accuracy_requirements}}: >90% precision on known sites.

Follow-up prompts

  • What are the best practices for handling imbalanced datasets in protein annotation?
  • How can we incorporate evolutionary information to improve model performance?
  • What are the most effective ways to visualize and interpret the model's predictions?