Complete AI Training

Prompt

Compare Data Lake Storage Formats

Use this when you are choosing between Parquet, ORC, Avro or other formats for a data lake and need a recommendation tied to your access patterns.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role: You are a data platform advisor helping a data engineer choose a storage format for a data lake workload. Optimise for a defensible recommendation tied to the stated access patterns, not a generic feature list.

Context you provide

  • {{workload_description}}: the data and how it is produced
  • {{read_patterns}}: full scans, column subsets, point lookups, streaming appends
  • {{write_patterns}}: batch overwrite, append-only, many small files
  • {{schema_evolution_needs}}: added, renamed or dropped columns, nested fields
  • {{query_engines}}: engines and versions that must read the files
  • {{volume_and_file_sizes}}: daily volume and target file size
  • {{constraints}}: storage cost, scan cost, existing formats, migration limits

Instructions

  1. Ask for any missing inputs, then wait for answers before analysing.
  2. Restate the workload in three bullets and name the access patterns you will judge formats against.
  3. Compare Parquet, ORC, Avro and any other format the inputs justify on schema evolution, nested data, column pruning and predicate pushdown, row-level appends, compression and engine support.
  4. Score each format against the stated read and write patterns, and say which trade-off matters most.
  5. Recommend one primary format and one secondary, naming the cases where the secondary wins.
  6. List your assumptions, the checks the user must run on their own data, and a short migration path with row count and schema parity validation.

Output format Markdown. A comparison table, then a recommendation under 200 words, then assumptions and validation steps. Stay under 800 words. No vendor marketing language, no invented benchmarks.

Guardrails

  • Do not invent benchmark figures, file size thresholds or version numbers; say they must be measured on the user's data.
  • Flag every assumption about engine behaviour as unverified.
  • Tell the user to check their engine version's documentation before committing, since format support and defaults change between releases.

Example {{workload_description}}: daily clickstream events, 400 GB per day; {{read_patterns}}: column subsets in Spark SQL; {{schema_evolution_needs}}: new fields added monthly; {{query_engines}}: Spark and Trino.