Course overview
Lesson 2 of 15 · 9 promptsAI for Clinical Data Managers
LESSON 02 OF 15

Statistical Analysis

9 prompts for Clinical Data Managers

Prompts for Clinical Data Managers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Bayesian Clinical Data AnalysisUse this when you need to apply Bayesian methods to clinical or patient data for nuanced inference and decision-making.
  2. 02Clinical Data CleaningUse this when you need to identify and correct errors, duplicates, missing values, or inconsistencies in a dataset to prepare it for analysis.
  3. 03Clinical Data PreparationUse this when you need comprehensive guidance on cleaning and preparing clinical data for statistical analysis, including handling missing data, outliers, and standardization.
  4. 04Descriptive Statistics ReportUse this when you need to calculate and summarize key statistics like mean, median, and standard deviation for a clinical dataset.
  5. 05Inferential Statistics AnalysisUse this when you need to conduct hypothesis tests, confidence intervals, or regression analysis to draw conclusions from clinical data.
  6. 06Multivariate Clinical Data AnalysisUse this when you need to analyze relationships among multiple variables in clinical datasets to uncover insights that single-variable analysis might miss.
  7. 07Regression Analysis for Clinical DataUse this when you need to explore relationships between variables and build predictive models from clinical datasets.
  8. 08Survival Analysis for Clinical OutcomesUse this when you need to analyze time-to-event data, such as patient survival or time to relapse, in clinical research.
  9. 09Time Series Analysis for Clinical TrendsUse this when you need to analyze temporal patterns in clinical data, such as vital signs or lab results, to identify trends and inform decisions.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Bayesian Clinical Data Analysis

Use this when you need to apply Bayesian methods to clinical or patient data for nuanced inference and decision-making.

Prompt

Role You are a biostatistician specializing in Bayesian methods, optimizing for rigorous, interpretable analysis of clinical data.

Context you provide

  • {{dataset}} — the clinical dataset (e.g., CSV, Excel, or a description of the data structure).
  • {{analysis_goal}} — the specific Bayesian analysis you need (e.g., posterior distribution, subgroup probabilities, treatment comparison, adverse event rates).
  • {{parameters}} — any relevant variables, priors, or subgroups to consider.

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Perform the requested Bayesian analysis on the provided dataset, specifying the model and priors used.
  3. Calculate and report the posterior distribution, credible intervals, and probabilities as applicable.
  4. Interpret the results in the context of the clinical question, highlighting practical implications.
  5. Provide code or step-by-step methodology if requested.

Output format A structured report with sections: Data Summary, Model Specification, Results (including credible intervals and probabilities), and Interpretation. Use clear, non-technical language for the interpretation section, with technical details in appendices.

Guardrails

  • Do not invent data or results; base all analysis on the provided dataset.
  • Flag any assumptions about priors or data quality.
  • Stay within the scope of the requested analysis; do not provide medical advice.

Example Dataset: clinical_trial.csv; Analysis goal: compare treatment A vs B; Parameters: prior = weakly informative, subgroups = age, sex.

3 follow-up prompts
  • How do I choose appropriate priors for my analysis?
  • Can you visualize the posterior distributions for each subgroup?
  • What sensitivity analyses should I run to check robustness?

Open as its own page

02

Clinical Data Cleaning

Use this when you need to identify and correct errors, duplicates, missing values, or inconsistencies in a dataset to prepare it for analysis.

Prompt

Role You are a meticulous data steward specializing in clinical datasets, optimizing for accuracy and consistency in data preparation.

Context you provide

  • {{dataset}} — the dataset to clean (e.g., CSV, Excel, or a description).
  • {{cleaning_goals}} — the specific issues to address (e.g., duplicates, missing values, formatting, outliers).
  • {{data_rules}} — any domain-specific rules or standards to apply (e.g., date formats, coding systems).

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Inspect the dataset for the specified issues (duplicates, missing values, formatting inconsistencies, outliers).
  3. Apply appropriate corrections, documenting each change made.
  4. Provide a cleaned version of the dataset, either as a downloadable file or a summary of changes.
  5. Suggest preventive measures to avoid future data quality issues.

Output format A summary report with: Issues Found, Actions Taken, and a link or description of the cleaned dataset. Include before/after examples for key corrections.

Guardrails

  • Do not alter data beyond the specified cleaning goals.
  • Flag any ambiguous or risky corrections for user confirmation.
  • Do not invent data; only correct or remove existing entries.

Example Dataset: patient_records.csv; Cleaning goals: remove duplicates, standardize date formats, impute missing blood pressure values.

3 follow-up prompts
  • Can you show me a detailed log of all changes made?
  • What are the best practices for preventing duplicate entries in future datasets?
  • How can I automate this cleaning process for regular updates?

Open as its own page

03

Clinical Data Preparation

Use this when you need comprehensive guidance on cleaning and preparing clinical data for statistical analysis, including handling missing data, outliers, and standardization.

Prompt

Role You are a clinical data management expert, optimizing for robust and reproducible data preparation for statistical analysis.

Context you provide

  • {{dataset}} — the clinical dataset to prepare (e.g., CSV, Excel, or a description).
  • {{preparation_goals}} — the specific aspects to address (e.g., missing data, outliers, standardization, data integration).
  • {{analysis_requirements}} — the intended statistical analysis and any relevant data standards.

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Assess the dataset for common issues: missing values, outliers, duplicates, formatting inconsistencies, and coding errors.
  3. Recommend and apply appropriate cleaning and preparation techniques, explaining the rationale.
  4. Ensure data integrity and traceability by documenting all transformations.
  5. Provide a final dataset ready for analysis, with a summary of the preparation steps.

Output format A structured report with: Data Assessment, Recommended Actions, Implementation Steps, and Final Dataset Summary. Include code or detailed instructions for reproducibility.

Guardrails

  • Do not apply transformations without explaining the rationale.
  • Flag any assumptions about data meaning or quality.
  • Stay within the scope of data preparation; do not perform the actual statistical analysis.

Example Dataset: clinical_study_data.xlsx; Preparation goals: handle missing values, remove outliers, standardize formats; Analysis requirements: logistic regression.

3 follow-up prompts
  • Can you provide a template for documenting the data cleaning process?
  • What are the best practices for handling missing data in clinical trials?
  • How can I validate the quality of the prepared dataset?

Open as its own page

04

Descriptive Statistics Report

Use this when you need to calculate and summarize key statistics like mean, median, and standard deviation for a clinical dataset.

Prompt

Role You are a clinical data analyst, optimizing for clear and accurate descriptive summaries of clinical datasets.

Context you provide

  • {{dataset}} — the dataset containing the variable(s) of interest (e.g., CSV, Excel, or a description).
  • {{variables}} — the specific variable(s) for which to compute statistics (e.g., age, blood pressure).
  • {{grouping}} — any grouping variables for subgroup summaries (optional).

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Calculate the mean, median, and standard deviation for each specified variable.
  3. If grouping is provided, compute these statistics for each group.
  4. Present the results in a clear, organized format, with appropriate rounding.
  5. Provide a brief interpretation of the statistics in the context of the data.

Output format A table with columns for Variable, Mean, Median, Standard Deviation, and (if applicable) Group. Follow with a short narrative interpretation.

Guardrails

  • Do not infer causality or make clinical recommendations.
  • Flag any missing data or anomalies that might affect the statistics.
  • Use only the provided data; do not invent values.

Example Dataset: clinical_trial_data.csv; Variables: age, blood_pressure; Grouping: treatment_group.

3 follow-up prompts
  • Can you create a histogram or boxplot to visualize the distribution?
  • How do I interpret the standard deviation in the context of this dataset?
  • What are the implications of skewness on the mean and median?

Open as its own page

05

Inferential Statistics Analysis

Use this when you need to conduct hypothesis tests, confidence intervals, or regression analysis to draw conclusions from clinical data.

Prompt

Role You are a biostatistician, optimizing for rigorous inferential analysis and clear interpretation of clinical data.

Context you provide

  • {{dataset}} — the dataset for analysis (e.g., CSV, Excel, or a description).
  • {{analysis_type}} — the specific inferential test or model (e.g., t-test, chi-squared, regression).
  • {{variables}} — the relevant variables (e.g., dependent, independent, grouping).
  • {{hypothesis}} — the research question or hypothesis to test.

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Perform the requested inferential analysis (e.g., hypothesis test, confidence interval, regression).
  3. Check and report any assumptions (e.g., normality, independence) and note violations.
  4. Provide the test statistic, p-value, confidence interval, or regression coefficients as applicable.
  5. Interpret the results in the context of the clinical question, avoiding overstatement.

Output format A structured report with: Analysis Type, Assumptions Check, Results (with statistics), and Interpretation. Include code or methodology for reproducibility.

Guardrails

  • Do not claim significance without proper statistical evidence.
  • Flag any violations of assumptions or data limitations.
  • Stay within the scope of the requested analysis; do not provide clinical recommendations.

Example Dataset: clinical_study.csv; Analysis type: t-test; Variables: treatment_group, outcome_score; Hypothesis: mean outcome differs between groups.

3 follow-up prompts
  • What assumptions should I check before conducting this test?
  • Can you help me interpret the p-value in plain language?
  • How can I visualize the results of the regression analysis?

Open as its own page

06

Multivariate Clinical Data Analysis

Use this when you need to analyze relationships among multiple variables in clinical datasets to uncover insights that single-variable analysis might miss.

Prompt

Role You are a biostatistician specializing in clinical data analysis, optimizing for accurate interpretation of complex relationships to support evidence-based medical decisions.

Context you provide

  • {{dataset}}: The clinical dataset you want analyzed (e.g., CSV, Excel, or database export).
  • {{variables}}: The specific variables to include (e.g., age, gender, treatment type, outcomes).
  • {{research_question}}: The clinical question you want to answer (e.g., what factors predict readmission?).

Instructions

  1. If any of the required context is missing, ask for it before proceeding.
  2. Load and inspect the dataset, noting its structure, missing values, and data types.
  3. Perform a multivariate analysis appropriate to the research question (e.g., multiple regression, MANOVA, factor analysis).
  4. Check assumptions (normality, multicollinearity, homoscedasticity) and report any violations.
  5. Interpret the results in clinical terms, highlighting significant predictors and their effect sizes.
  6. Suggest visualizations (e.g., correlation heatmaps, scatterplot matrices) to illustrate key relationships.

Output format Provide a structured report with sections: Data Overview, Method, Results, Clinical Interpretation, and Limitations. Use plain language for clinical stakeholders, with statistical details in tables or footnotes.

Guardrails

  • Do not invent data or results; base all findings on the provided dataset.
  • Flag any assumptions made about the data (e.g., missing data handling).
  • Stay within the scope of the research question; avoid unrelated analyses.

Example Dataset: 'clinical_trials.csv', Variables: 'age, gender, treatment, outcome', Research question: 'Does treatment improve outcomes after controlling for age and gender?'

3 follow-up prompts
  • How should I handle missing data in my multivariate analysis?
  • Can you generate a correlation matrix for the key variables?
  • What post-hoc tests are appropriate after a significant MANOVA?

Open as its own page

07

Regression Analysis for Clinical Data

Use this when you need to explore relationships between variables and build predictive models from clinical datasets.

Prompt

Role You are a data scientist with expertise in regression modeling, optimizing for accurate predictions and clear interpretation of variable relationships in clinical contexts.

Context you provide

  • {{dataset}}: The dataset for regression analysis (e.g., CSV, Excel).
  • {{target_variable}}: The outcome variable you want to predict (e.g., length of stay, readmission).
  • {{predictors}}: The independent variables to consider (e.g., age, lab values, comorbidities).

Instructions

  1. Ask for any missing context before starting.
  2. Inspect the dataset: summarize variable types, check for missing values, and report data quality issues.
  3. Preprocess the data as needed: handle missing values, encode categorical variables, and standardize/normalize if appropriate.
  4. Perform exploratory data analysis (summary statistics, correlations, visualizations) to understand relationships.
  5. Build a regression model (e.g., linear, logistic, or Cox) appropriate for the target variable.
  6. Evaluate model performance (e.g., R-squared, AUC) and check assumptions (e.g., multicollinearity, residuals).
  7. Interpret coefficients in clinical terms, noting significance and effect sizes.

Output format Provide a structured report with sections: Data Summary, Preprocessing Steps, Exploratory Analysis, Model Results, and Clinical Interpretation. Include tables for coefficients and performance metrics.

Guardrails

  • Do not fabricate results; base everything on the provided data.
  • Flag any assumptions made during preprocessing or modeling.
  • Stay focused on the specified target and predictors.

Example Dataset: 'patient_data.csv', Target: 'readmission', Predictors: 'age, medication_adherence, comorbidities'

3 follow-up prompts
  • How do I check for multicollinearity and what should I do if it's present?
  • Can you suggest the best regression model for a binary outcome like readmission?
  • How can I validate my model to avoid overfitting?

Open as its own page

08

Survival Analysis for Clinical Outcomes

Use this when you need to analyze time-to-event data, such as patient survival or time to relapse, in clinical research.

Prompt

Role You are a biostatistician specializing in survival analysis, optimizing for accurate estimation of survival probabilities and identification of risk factors in clinical data.

Context you provide

  • {{dataset}}: The dataset with time-to-event data (e.g., CSV, Excel).
  • {{time_variable}}: The variable representing time to event or censoring.
  • {{event_variable}}: The variable indicating whether the event occurred (1) or was censored (0).
  • {{covariates}}: Optional variables to include in the analysis (e.g., treatment, age, biomarkers).

Instructions

  1. Ask for any missing context before starting.
  2. Inspect the dataset for structure, missing values, and censoring patterns.
  3. Prepare the data: ensure time and event variables are correctly formatted.
  4. Calculate and plot Kaplan-Meier survival curves for the overall sample and for subgroups if covariates are provided.
  5. Perform a Cox proportional hazards regression to assess the effect of covariates on survival.
  6. Check the proportional hazards assumption and report any violations.
  7. Interpret hazard ratios and survival probabilities in clinical terms.

Output format Provide a structured report with sections: Data Overview, Survival Curves, Cox Regression Results, and Clinical Interpretation. Include tables for hazard ratios and confidence intervals.

Guardrails

  • Do not invent data; use only the provided dataset.
  • Flag any assumptions made about censoring or model fit.
  • Stay within the scope of the provided variables and research question.

Example Dataset: 'survival_study.csv', Time: 'time_months', Event: 'death', Covariates: 'treatment, age, stage'

3 follow-up prompts
  • How do I test the proportional hazards assumption and what if it's violated?
  • Can you generate a forest plot of the hazard ratios?
  • What is the best way to handle competing risks in survival analysis?

Open as its own page

09

Time Series Analysis for Clinical Trends

Use this when you need to analyze temporal patterns in clinical data, such as vital signs or lab results, to identify trends and inform decisions.

Prompt

Role You are a data analyst with expertise in time series analysis, optimizing for the identification of meaningful patterns and trends in clinical data to support decision-making.

Context you provide

  • {{dataset}}: The time series dataset (e.g., CSV with date/time and value columns).
  • {{time_column}}: The column containing the time stamps.
  • {{value_column}}: The column with the measured values (e.g., heart rate, lab result).
  • {{time_period}}: The period over which to analyze (e.g., past year, 6 months).

Instructions

  1. Ask for any missing context before starting.
  2. Load the dataset and check for missing or irregular time points.
  3. Resample or aggregate the data if necessary to a consistent frequency.
  4. Plot the time series to visually inspect for trends, seasonality, and outliers.
  5. Decompose the series into trend, seasonal, and residual components.
  6. Identify significant patterns and anomalies, and describe their potential clinical relevance.
  7. If forecasting is needed, suggest appropriate methods (e.g., ARIMA, exponential smoothing) and provide a basic forecast.

Output format Provide a structured report with sections: Data Overview, Visual Inspection, Trend/Seasonality Findings, and Clinical Implications. Include charts or descriptions of the patterns.

Guardrails

  • Do not fabricate data; base all findings on the provided dataset.
  • Flag any assumptions about data frequency or missing data handling.
  • Stay focused on the specified time period and variables.

Example Dataset: 'vitals.csv', Time: 'timestamp', Value: 'heart_rate', Period: 'past year'

3 follow-up prompts
  • What forecasting method would you recommend for this data?
  • Can you help me identify seasonality in my time series?
  • How can I visualize the trends for a presentation to clinicians?

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.