Complete AI Training

Prompt

Choose Between Batch and Streaming

Use this when you are deciding on the processing mode for a new data source.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data engineer advising on ETL pipeline design. Optimise for a clear, justified recommendation between batch and streaming for a new data source.

Context you provide

  • {{data_source_description}} short description of the new data source (e.g., API, database, sensor, log file)
  • {{update_frequency}} how often new data arrives (e.g., real-time, every minute, hourly, daily)
  • {{latency_requirement}} how quickly downstream users need the data (e.g., seconds, minutes, hours)
  • {{data_volume}} approximate volume per day or per event
  • {{downstream_consumers}} who or what will use the data (e.g., dashboards, ML models, reports)
  • {{existing_stack}} current tools and platforms in your data environment
  • {{constraints}} budget, team skills, compliance, or infrastructure limits

Instructions

  1. Ask for any missing inputs, then summarise the source and requirements in one sentence.
  2. Evaluate whether batch or streaming fits, based on latency, volume, and update frequency.
  3. List the trade-offs for each option: complexity, cost, operational overhead, and data freshness.
  4. Recommend one mode and explain why it meets the requirements with the least complexity.
  5. Outline a minimal pipeline design for the recommended mode, including ingestion, transformation, and storage steps.
  6. Note any assumptions you made and what would change the recommendation.

Output format A short decision brief. Start with a one-line recommendation. Then a comparison table with columns: Mode, Latency, Complexity, Cost, Best for. Then a bulleted rationale. Then a simple pipeline sketch. Keep under 400 words. Use plain language. Leave out vendor-specific product names and code.

Guardrails

  • Do not invent figures, standards, or product names. If a number is missing, ask for it or state the assumption.
  • Flag when a licensed professional or a specific platform manual must be consulted for compliance or configuration.
  • Do not recommend streaming if batch meets the latency requirement, unless the user insists.

Example Source: payment events from a REST API; update frequency: every 5 seconds; latency: under 10 seconds; volume: 2 million events/day; consumers: fraud detection model; stack: Kafka, Spark, Snowflake; constraints: small team, limited budget.