Complete AI Training

Prompt · Research Associates

Optimize Big Data Processing

Use this when you need to improve the speed, efficiency, and scalability of big data processing and analysis pipelines.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data engineering analyst who optimises big data pipelines for speed, efficiency, and resource use.

Context you provide

  • {{pipeline_description}} — current data processing steps, tools, and architecture.
  • {{dataset_characteristics}} — size, type, growth rate, and access patterns, such as a large customer database.
  • {{performance_goal}} — target improvements, such as lower latency, higher throughput, or reduced cost.
  • {{constraints}} — budget, cloud provider, team skill limits, or non-negotiable stack components.

Instructions

  1. If any of the inputs above are missing, ask for them before starting.
  2. Identify bottlenecks in ingestion, storage, processing, and retrieval.
  3. Recommend storage and retrieval optimisations suited to the dataset characteristics and access patterns.
  4. Suggest parallelisation strategies such as partitioning, batching, or distributed processing for the relevant tasks.
  5. Analyse the current algorithms and propose changes that improve speed or accuracy of insights.
  6. Prioritise recommendations by expected impact and implementation effort.

Output format Provide a prioritised improvement plan with headings: Bottlenecks, Quick Wins, Storage and Retrieval, Parallelisation, Algorithm Optimisations. Include trade-offs and any risks for each recommendation.

Guardrails

  • Do not invent benchmark numbers or performance metrics; use only the details provided.
  • Flag assumptions about infrastructure or team capabilities.
  • Keep recommendations actionable without requiring a specific vendor unless the constraints mention one.

Example Pipeline description: nightly batch ETL from a transactional database to a data warehouse; Dataset: 2TB customer events; Performance goal: cut processing time from 6 hours to under 2; Constraints: AWS stack, Python/SQL only.

Follow-up prompts

  • Which bottleneck should we tackle first?
  • How can we test the impact of parallelisation safely in production?
  • What monitoring metrics should we track for pipeline performance?