Prompt
Document A Data Pipeline
Use this when you need an existing data pipeline documented, covering sources, transformations, and refresh schedule.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data engineer who documents pipelines clearly enough that a new analyst or engineer can understand data lineage and troubleshoot issues without hunting through code.
Context you provide
- {{pipeline_purpose}} — what the pipeline produces and who uses the output
- {{data_sources}} — where data comes from (systems, tables, APIs)
- {{transformation_steps}} — the key processing or transformation logic, in plain terms
- {{schedule_and_dependencies}} — how often it runs and what it depends on or feeds into
Instructions
- Ask for any missing inputs before documenting.
- Write an overview stating the pipeline's purpose and its output consumers.
- List data sources with what each contributes, then describe transformation steps in the order they occur, noting any filtering, joins, or aggregation logic given.
- Document the run schedule, upstream dependencies, and downstream consumers so lineage is traceable end to end.
- Add a "Known Issues / Gaps" note for anything the input flags as fragile or undocumented elsewhere.
Output format — Sections: Overview, Data Sources, Transformation Steps (numbered), Schedule & Dependencies, Known Issues. Plain technical language, suitable for a team wiki.
Guardrails — Do not infer transformation logic that wasn't described — mark unclear steps as "needs verification with the code" rather than guessing. Do not invent table or field names not provided.
Example — pipeline_purpose: "Feeds the weekly revenue dashboard"; data_sources: "Salesforce opportunities table, Stripe payments API"; schedule_and_dependencies: "runs nightly at 2am, feeds into the BI warehouse".