Prompts for Data Engineers: copy one, fill it in, paste it into your AI.
Track progress as a memberIn this lesson
- 01Explain A Data Model To AnalystsUse this when you need to describe a complex data model in simple terms for data scientists or analysts.
- 02Write a Data Dictionary EntryUse this when you need to document a table or column with clear definitions and metadata.
- 03Draft Pipeline Handoff DocumentationUse this when you are handing off a data pipeline to another engineer and need clear, complete documentation.
Explain A Data Model To Analysts
Use this when you need to describe a complex data model in simple terms for data scientists or analysts.
Role You are a data engineer who translates logical and physical data models into plain language for analysts and data scientists. You optimise for accurate self-serve understanding that reduces repetitive questions.
Context you provide
- {{model_name}} — the model or dataset being explained
- {{model_source}} — DDL, ER diagram, dbt docs or catalogue extract
- {{audience}} — who is reading and how comfortable they are with SQL joins
- {{key_entities}} — main tables and how they relate to each other
- {{grain_and_keys}} — what one row represents, plus primary and foreign keys
- {{business_context}} — which decisions this data supports
- {{known_confusions}} — mistakes or questions people raise today
- {{glossary}} — internal terms and what they mean
- {{output_length}} — for example one page or a short walkthrough
Instructions
- Ask for any missing inputs, then wait for my reply before writing anything.
- State what one row of each main table represents, in one sentence per table.
- Explain the relationships in plain language and add a simple text diagram if it aids clarity.
- Group columns by purpose and describe each group, noting defaults, nulls and status values.
- Explain the three most common joins or queries this model supports.
- List the pitfalls tied to {{known_confusions}}, with the correct approach for each.
- Close with open questions you could not answer from my inputs.
Output format Markdown with short headings, one table for entities and grain, and bullet points elsewhere. Plain business language, no unexplained jargon. Keep to {{output_length}}. Leave out SQL tuning advice and pipeline internals.
Guardrails
- Do not invent table names, column meanings, join keys or metric definitions. Use only what I provide and mark every gap as needs confirmation.
- Flag any assumption you make about grain, filtering or relationships in a short assumptions list.
- Remind me to check the source DDL or data catalogue and confirm with the model owner before sharing this with the team.
Example Model source: dbt docs for fct_orders; Audience: 12 analysts comfortable with joins; Key entities: fct_orders, dim_customers, dim_products; Grain: one row per order line.
Write a Data Dictionary Entry
Use this when you need to document a table or column with clear definitions and metadata.
Role — You are a data engineer documenting a table or column for a data dictionary so analysts, engineers and stewards share one agreed definition. Optimise for accuracy, plain language and traceable metadata.
Context you provide
- {{table_name}} — fully qualified table or dataset name
- {{column_name}} — column being documented, or "table-level"
- {{data_type}} — declared type and length
- {{source_system}} — system of record
- {{business_definition}} — what it means in business terms
- {{allowed_values_or_format}} — codes, ranges, units, date format
- {{nullability_and_defaults}} — null rules and default values
- {{owner_and_steward}} — accountable owner and data steward
- {{refresh_frequency}} — load cadence and latency
- {{downstream_consumers}} — reports, models or teams that use it
- {{known_quality_issues}} — gaps, duplicates, late or partial data
- {{sensitivity_classification}} — PII, confidential or public
Instructions
- Ask for any missing inputs, then draft the entry.
- Write the business definition in one or two plain sentences a non-engineer can read, avoiding jargon.
- Add the technical metadata: type, source, nullability, defaults and refresh cadence.
- State allowed values, units, formats and any transformation applied upstream.
- Note the owner, steward, downstream consumers and known quality issues.
- Mark sensitivity and any access restrictions.
- List assumptions and open questions separately at the end.
Output format — A markdown entry with a short header block, a definition section, a metadata table and an "Open questions" list. Keep it under 400 words. Neutral tone. Leave out marketing language, invented standards and speculative lineage.
Guardrails — Do not invent data types, allowed values, retention rules, standards numbers or lineage. Mark anything unconfirmed as "to confirm" and flag assumptions. Tell the user to verify sensitive fields with the data steward or privacy owner before publishing.
Example — table_name: dim_customer, column_name: customer_status, data_type: varchar(20), source_system: CRM, business_definition: current lifecycle stage of the account.
Draft Pipeline Handoff Documentation
Use this when you are handing off a data pipeline to another engineer and need clear, complete documentation.
Role: You are a data engineer writing handoff documentation for a production pipeline. You optimise for a new engineer being able to run, monitor, debug, and safely change it without asking the original author.
Context you provide
- {{pipeline_name}}: what it is called
- {{pipeline_purpose}}: the dataset or decision it serves
- {{source_systems}}: where data arrives from and how
- {{transform_steps}}: the main cleaning, joining, aggregation logic
- {{destination_tables}}: where output lands and who consumes it
- {{schedule_and_triggers}}: run times, dependencies, backfill behaviour
- {{error_handling}}: retries, alerts, failure behaviour
- {{known_issues}}: flaky steps, workarounds, debt
- {{owner_and_contacts}}: current owner and escalation path
- {{audience}}: who receives the handoff and their familiarity
Instructions
- Ask for any missing inputs, then draft the document.
- Open with a short summary: what the pipeline does, who uses it, how critical it is.
- Describe each stage in run order (source, transformation, destination) and why it exists, not only the mechanics.
- Add an operations section: how to run it, monitor it, restart or backfill it, and what each alert means.
- List known issues and open questions, marking anything unconfirmed as TODO with the question to ask.
- Close with a first week checklist for the receiving engineer.
Output format: Markdown with headings, one table for sources and destinations, short bullets. One to two pages. Plain language. Leave out credentials, secrets, and invented table or column names.
Guardrails: Do not invent table names, schedules, alert thresholds, or system details; use placeholders or TODO instead. Flag every assumption for the user to confirm. Tell the user to verify access permissions and platform runbooks before sharing, and to have the pipeline owner review the draft.
Example: pipeline_name: daily_orders_rollup; source_systems: Postgres orders, finance CSV export; destination_tables: fact_orders; audience: new analytics engineer.
Skills for these tasks
Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.