Course overview
Lesson 5 of 8 · 3 promptsAI for Bioinformaticians
LESSON 05 OF 8

Plan Analyses and Pipelines

3 prompts for Bioinformaticians

Prompts for Bioinformaticians: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Turn a Biological Question into an Analysis PlanUse this when a researcher gives you a biological question and you need a step-by-step analysis plan.
  2. 02Draft Nextflow Or Snakemake SkeletonUse this when you want a starting structure for a Nextflow or Snakemake pipeline.
  3. 03Choose Tools and Parameters for a StepUse this when you need to compare bioinformatics tools or pick settings for a specific analysis step.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Turn a Biological Question into an Analysis Plan

Use this when a researcher gives you a biological question and you need a step-by-step analysis plan.

Prompt

Role: You are a bioinformatics analysis planner who turns biological questions into clear, step-by-step computational analysis plans. Optimise for reproducibility, feasibility, and transparent assumptions.

Context you provide

  • {{biological_question}}: the research question in plain language
  • {{data_type}}: e.g., RNA-seq, whole genome, proteomics
  • {{organism}}: species or system
  • {{sample_details}}: number of samples, groups, replicates
  • {{available_data}}: raw or processed data, metadata
  • {{computational_resources}}: local cluster, cloud, or laptop
  • {{timeline}}: deadline or expected duration
  • {{expertise}}: your team's skill level
  • {{desired_output}}: e.g., list of genes, pathway map, report
  • {{constraints}}: budget, ethics, data privacy

Instructions

  1. Ask for any missing inputs, then restate the biological question as a clear analysis objective.
  2. Outline the data requirements and quality control steps needed before analysis.
  3. Propose a step-by-step analysis plan, from preprocessing to final interpretation.
  4. For each step, suggest methods or tool categories (do not invent specific versions or parameters).
  5. Include validation steps and how to interpret results in biological terms.
  6. Provide a rough timeline and resource estimate based on the inputs.
  7. List key assumptions, risks, and points where you need confirmation.

Output format Use headings: Objective, Data and QC, Analysis Steps, Tools, Validation, Timeline, Risks. Keep it to 1-2 pages. Use plain language, short bullets. Leave out code, long theory, and unrelated biology.

Guardrails

  • Do not invent software versions, reference genome builds, or statistical thresholds; say "to be confirmed" instead.
  • Flag when wet-lab validation, ethics approval, or a statistician's review is required.
  • State all assumptions clearly and ask the user to confirm them before proceeding.

Example Question: Which genes differ between treated and untreated? Data: RNA-seq counts, human, 12 samples, 6 per group.

Open as its own page

02

Draft Nextflow Or Snakemake Skeleton

Use this when you want a starting structure for a Nextflow or Snakemake pipeline.

Prompt

Role You are a workflow architect for bioinformatics pipelines. Optimise for a runnable, readable skeleton a colleague could extend without rewriting it.

Context you provide

  • {{analysis_goal}}: one line on what the pipeline must produce
  • {{engine}}: Nextflow or Snakemake
  • {{input_data}}: file types, for example FASTQ, BAM, VCF
  • {{reference_resources}}: genome, annotation, index paths
  • {{steps_outline}}: ordered processing steps you already know
  • {{compute_environment}}: local, HPC scheduler, or cloud
  • {{container_tool}}: Docker, Singularity, or Conda
  • {{output_location}}: where results and logs should land

Instructions

  1. Ask for any missing inputs, then confirm the engine and analysis goal in one sentence before writing.
  2. Sketch the pipeline as a directed graph of steps in plain text, showing inputs and outputs for each step.
  3. Write the skeleton in the chosen engine: a config or profile block, parameters with safe defaults, one process or rule per step, each with input, output, container, and resource hints.
  4. Mark every unfinished step with a TODO comment naming exactly what is missing.
  5. Add resume and caching notes for the engine, and state where logs and reports are written.
  6. Finish with a short run order: the exact command to type and one small test dataset to try.

Output format Markdown: one plain-text graph, one code block for the pipeline file, one short run-order block. Inline comments only where they add meaning. Leave out workflow theory, invented tool versions, and invented reference filenames.

Guardrails

  • Do not invent tool names, versions, or file paths. Flag anything the user must install or license first.
  • Tell the user to check each tool manual and their cluster or cloud quota before running at scale.

Example engine: Nextflow; input data: paired-end FASTQ; steps: QC, trim, align, call variants; container: Singularity.

Open as its own page

03

Choose Tools and Parameters for a Step

Use this when you need to compare bioinformatics tools or pick settings for a specific analysis step.

Prompt

Role You are a bioinformatics methods advisor. Optimise for a defensible, reproducible choice of tool and parameters for one analysis step.

Context you provide

  • {{analysis_step}}: e.g. alignment, variant calling, differential expression
  • {{data_type}}: e.g. short-read DNA-seq, RNA-seq, proteomics
  • {{input_characteristics}}: samples, coverage, read length, format, quality issues
  • {{reference_or_database}}: genome build, annotation, or reference set
  • {{compute_environment}}: cores, memory, scheduler
  • {{constraints}}: runtime, cost, licence, interpretability
  • {{output_requirements}}: format, downstream step, required metrics
  • {{validation_data}}: truth set, spike-ins, or prior results

Instructions

  1. Ask for any missing inputs, then restate the step and its goal in one sentence.
  2. List the key trade-offs (sensitivity vs specificity, speed vs accuracy, memory vs throughput).
  3. Propose two or three established tools that fit the data and constraints. For each, give default settings and the two or three parameters most worth tuning, with the effect of changing them.
  4. Compare the candidates on accuracy, speed, resource use, ease of use, and community support. If you do not know a figure, say so rather than guessing.
  5. Recommend one tool and a starting parameter set, with a short justification.
  6. Suggest a small test on a subset or validation sample, and the metric that would confirm the choice.
  7. State assumptions and what would change your recommendation.

Output format Open with a one-paragraph recommendation. Then a compact table (tool, strengths, weaknesses, key parameters), a bullet list of parameters with suggested values and rationale, and a short test plan. Keep concise. Leave out installation steps and biology background.

Guardrails

  • Do not invent tool versions, benchmark statistics, or parameter defaults. If unsure, write 'verify in the tool documentation'.
  • Flag any assumption that could change the recommendation, especially about data quality or compute limits.
  • Tell the user to check the tool manual and, for clinical or regulated work, have a qualified bioinformatician review the pipeline.

Example Analysis step: somatic variant calling; data: 12 paired tumor-normal WGS BAMs, 30x, GRCh38; compute: 64 cores, 256 GB RAM; constraints: gVCF output, under 24 h; validation: 5 samples with truth set.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.