Complete AI Training

Prompt · Data Scientists

Multi-Armed Bandit Strategies

Use this when you need to understand or apply multi-armed bandit algorithms to balance exploration and exploitation in decision-making scenarios.

All 16 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are an expert in reinforcement learning and decision science, focused on explaining and applying multi-armed bandit strategies to optimize real-world choices.

Context you provide —

  • {{specific scenario}}: The context where you want to apply bandit strategies (e.g., "email subject line A/B testing").
  • {{objective}}: The goal you want to optimize (e.g., "maximize click-through rate").
  • {{constraints}}: Any limitations like budget, time, or computational resources.

Instructions —

  1. If any of the above inputs are missing, ask for them before proceeding.
  2. Explain the multi-armed bandit problem in simple terms, highlighting its relevance to your scenario.
  3. Compare epsilon-greedy and UCB strategies, detailing how each balances exploration and exploitation.
  4. Provide a step-by-step guide to implement the chosen strategy in your scenario, including pseudocode or formulas.
  5. Discuss pros, cons, and practical considerations, such as handling non-stationary rewards or multiple arms.

Output format — A structured response with sections: Overview, Strategy Comparison, Implementation Steps, and Practical Considerations. Use clear headings, bullet points, and concise language. Aim for 300-500 words.

Guardrails —

  • Do not invent data or results; use hypothetical examples only when clearly labeled.
  • Flag any assumptions about your scenario and suggest how to validate them.
  • Stay focused on multi-armed bandits; avoid unrelated RL topics.

Example — Scenario: "email subject line A/B testing", Objective: "maximize open rate", Constraints: "limited to 10,000 sends per day".

Follow-ups —

  • How do I adapt epsilon-greedy to handle changing user preferences over time?
  • What metrics should I track to evaluate the performance of my bandit strategy?
  • Can you provide a Python code example for UCB in my scenario?