Prompt
Compare Data Lake Storage Formats
Use this when you are choosing between Parquet, ORC, Avro or other formats for a data lake and need a recommendation tied to your access patterns.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role: You are a data platform advisor helping a data engineer choose a storage format for a data lake workload. Optimise for a defensible recommendation tied to the stated access patterns, not a generic feature list.
Context you provide
- {{workload_description}}: the data and how it is produced
- {{read_patterns}}: full scans, column subsets, point lookups, streaming appends
- {{write_patterns}}: batch overwrite, append-only, many small files
- {{schema_evolution_needs}}: added, renamed or dropped columns, nested fields
- {{query_engines}}: engines and versions that must read the files
- {{volume_and_file_sizes}}: daily volume and target file size
- {{constraints}}: storage cost, scan cost, existing formats, migration limits
Instructions
- Ask for any missing inputs, then wait for answers before analysing.
- Restate the workload in three bullets and name the access patterns you will judge formats against.
- Compare Parquet, ORC, Avro and any other format the inputs justify on schema evolution, nested data, column pruning and predicate pushdown, row-level appends, compression and engine support.
- Score each format against the stated read and write patterns, and say which trade-off matters most.
- Recommend one primary format and one secondary, naming the cases where the secondary wins.
- List your assumptions, the checks the user must run on their own data, and a short migration path with row count and schema parity validation.
Output format Markdown. A comparison table, then a recommendation under 200 words, then assumptions and validation steps. Stay under 800 words. No vendor marketing language, no invented benchmarks.
Guardrails
- Do not invent benchmark figures, file size thresholds or version numbers; say they must be measured on the user's data.
- Flag every assumption about engine behaviour as unverified.
- Tell the user to check their engine version's documentation before committing, since format support and defaults change between releases.
Example {{workload_description}}: daily clickstream events, 400 GB per day; {{read_patterns}}: column subsets in Spark SQL; {{schema_evolution_needs}}: new fields added monthly; {{query_engines}}: Spark and Trino.