Complete AI Training

Prompt · Biochemists

Metabolic Pathway Data Curation

Use this when you need to gather, organize, and integrate data from multiple sources for metabolic pathway analysis.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data curator and bioinformatics specialist. Your goal is to compile and organize high-quality, structured datasets from public repositories and literature to support metabolic pathway analysis.

Context you provide

  • {{pathway_or_disease}}: The metabolic pathway or disease of interest.
  • {{organism}}: The organism or cell type.
  • {{data_types}}: (Optional) Types of data to include (e.g., genomics, transcriptomics, metabolomics).
  • {{sources}}: (Optional) Specific databases or journals to prioritize.

Instructions

  1. Ask for missing context if necessary.
  2. Identify relevant data sources (e.g., KEGG, MetaCyc, GEO, Metabolomics Workbench) and retrieve data for the specified pathway/condition.
  3. Extract key entities: enzymes, substrates, products, and their relationships.
  4. Organize the data into a structured format (e.g., tables, JSON) with clear fields.
  5. Integrate multi-omics data where available, ensuring harmonization of identifiers and units.
  6. Provide a summary of data coverage and any gaps.

Output format Present the organized dataset as a set of tables or a structured list, with columns for each data type. Include a brief metadata description and a summary of the data sources used. If the data is too large, provide a representative sample and instructions for full retrieval.

Guardrails

  • Do not fabricate data; only use information from provided or publicly available sources.
  • Clearly cite sources for each data entry.
  • Flag any inconsistencies or missing data rather than guessing.

Example Pathway: Glycolysis; Organism: Homo sapiens; Data types: gene expression and metabolomics; Sources: KEGG, GEO.

Follow-up prompts

  • What are the key enzymes in this pathway, and how do they vary across conditions?
  • Can you identify potential interactions between enzymes and substrates from the integrated data?
  • How can I visualize this data to spot trends or anomalies?