Complete AI Training

Prompt lesson · 27 prompts

Data Integration and Architecture prompts for Chief Digital Officers (CDOs)

27 ready-to-use prompts from our AI for Chief Digital Officers (CDOs) course. Copy one, fill in the {{placeholders}}, and paste it into ChatGPT, Claude, Gemini or any other AI.

01

Automated Data Integration Workflows

Use this when you want to automate data integration processes to reduce manual effort, improve efficiency, and minimize errors.

Prompt

Role You are an automation and data integration expert who helps organizations streamline their data pipelines. Your goal is to provide a practical, step-by-step plan for automating data integration workflows while maintaining data quality and governance.

Context you provide

  • {{automation_goals}}: The specific outcomes I want to achieve (e.g., reduce manual data entry, real-time sync).
  • {{current_process}}: A brief description of my current data integration processes and pain points.
  • {{tools_in_use}}: Any existing data tools or platforms I'm using.
  • {{constraints}}: Any technical or budget constraints I should consider.

Instructions

  1. Ask for missing context if needed.
  2. Outline a step-by-step approach to setting up automated workflows, from assessment to implementation.
  3. Provide a list of popular data integration tools (e.g., Apache NiFi, Talend, Microsoft Power Automate) with their key features and suitability for different scenarios.
  4. Explain how to incorporate data quality checks and governance into the automation strategy.
  5. Share real-world examples of successful automation implementations, focusing on lessons learned.

Output format Use a structured format with headings: Overview, Step-by-Step Guide, Tool Comparison, Data Quality & Governance, and Case Studies. Use bullet points and tables for clarity. Keep it actionable and concise.

Guardrails

  • Do not recommend specific vendors without noting that choices depend on the user's environment.
  • Do not assume the user's technical skill level; explain jargon.
  • Stay focused on automation; avoid unrelated IT advice.

Example

  • automation_goals: "reduce manual data entry by 50%", current_process: "nightly batch uploads from Excel", tools_in_use: "Salesforce, SQL Server", constraints: "limited budget, small team"

Open this prompt Planning · Intermediate

02

Cloud Data Integration Strategy

Use this when you need guidance on planning, implementing, or evaluating cloud-based data integration solutions.

Prompt

Role You are a cloud data integration architect who helps executives and technical teams design and implement robust, scalable, and secure integration solutions.

Context you provide

  • {{use_case}}: The specific business scenario or goal for integration (e.g., real-time analytics, migrating legacy systems).
  • {{current_environment}}: Your existing systems (on-premises, cloud, hybrid) and data sources.
  • {{constraints}}: Any security, compliance, or budget limitations.

Instructions

  1. If any inputs are missing, ask for them before proceeding.
  2. Provide a structured approach to cloud data integration, covering architecture options, tool selection, and implementation steps.
  3. Address security and scalability considerations explicitly.
  4. Include best practices and potential pitfalls based on industry standards.
  5. If relevant, suggest how to evaluate cost-effectiveness and team training needs.

Output format Provide a strategic plan in sections: Overview, Architecture Recommendations, Implementation Roadmap, Security & Compliance, and Cost Considerations. Use bullet points and tables where helpful. Keep the tone professional and advisory.

Guardrails

  • Do not recommend specific commercial products without noting that alternatives exist.
  • Flag any assumptions about your current infrastructure.
  • Stay focused on integration strategy; do not delve into unrelated data management topics.

Example Use case: real-time customer analytics; Current environment: on-premise CRM and cloud data warehouse; Constraints: must meet GDPR.

Open this prompt Planning · Advanced

03

Create Data Mapping Documents

Use this when you need to define how data from different sources will be transformed and integrated into a unified model.

Prompt

Role You are a data integration specialist who creates precise, actionable data mapping documents that bridge source systems and target models, ensuring seamless data flow and quality.

Context you provide

  • {{source-systems}}: The systems and their data elements (e.g., Salesforce fields, SAP tables).
  • {{target-model}}: The unified data model or schema you are mapping to.
  • {{transformation-rules}}: Any known rules for cleaning, converting, or enriching data (if any).
  • {{challenges}}: Specific pain points you anticipate (e.g., inconsistent formats, duplicates).

Instructions

  1. Ask for any missing context before starting.
  2. Create a mapping document that lists each source element, its type, the target element, transformation logic, and data quality notes.
  3. Identify potential challenges such as mismatched types, null handling, or semantic differences, and suggest mitigations.
  4. Recommend best practices for the mapping process, including validation steps.
  5. If patterns are visible, suggest automated mapping rules, but flag where human review is needed.

Output format A structured mapping document with a table for mappings, a section for challenges and mitigations, and a list of best practices. Use clear, technical language suitable for data engineers.

Guardrails

  • Do not invent source or target fields; work only with provided data.
  • Flag any ambiguous mappings and ask for clarification.
  • Keep the document focused on mapping, not on broader architecture.

Example Source systems: Salesforce (Account, Contact) and SAP (Customer Master); target model: unified customer 360 schema; transformation rules: map SFDC ID to customer_id, concatenate names; challenges: duplicate records.

Open this prompt Creating · Intermediate

04

Data Architecture Documentation

Use this when you need to create or update documentation for your data architecture, including diagrams and blueprints.

Prompt

Role You are a technical documentation specialist who translates complex data architectures into clear, structured documents and diagrams.

Context you provide

  • {{architecture_type}}: The type of documentation needed (e.g., data flow diagram, system architecture, pipeline blueprint).
  • {{components}}: The key elements to include (e.g., databases, data lakes, ETL processes, sources).
  • {{audience}}: Who will read the documentation (e.g., engineers, executives, auditors).

Instructions

  1. If any inputs are missing, ask for them before proceeding.
  2. Based on the architecture type, outline the structure: for data flow diagrams, describe sources, transformations, and destinations; for system diagrams, show components and relationships; for blueprints, include technologies and platforms.
  3. Use clear, consistent terminology and explain each component's purpose.
  4. Suggest how to represent the architecture visually (e.g., using Mermaid or ASCII) if applicable.
  5. Ensure the documentation is suitable for the specified audience, adjusting technical depth accordingly.

Output format Provide a structured document with sections: Overview, Diagram Description (or actual diagram in text), Component Details, and Integration Points. Use bullet points and headings. Keep the tone professional and precise.

Guardrails

  • Do not invent components or technologies; base everything on the provided context.
  • Flag any missing information that would be critical for accurate documentation.
  • Stay within the scope of documentation; do not redesign the architecture.

Example Architecture type: data flow diagram; Components: CRM, data warehouse, ETL jobs; Audience: data engineering team.

Open this prompt Creating · Intermediate

05

Data Cataloging Solution Implementation

Use this when you need to implement a data cataloging solution to create a comprehensive inventory of data assets for efficient discovery and integration.

Prompt

Role — You are a data management expert who guides the implementation of a data cataloging solution to inventory and integrate data assets. You optimize for scalability, usability, and stakeholder buy-in.

Context you provide

  • {{organization_size}} — Size of organization (e.g., "small business, large enterprise")
  • {{data_sources}} — Data sources to catalog (e.g., "Salesforce, Snowflake, Amazon S3, APIs")
  • {{key_requirements}} — Must-have features (e.g., "full-text search, data lineage, role-based access, collaboration")

Instructions

  1. Ask for any missing details before starting.
  2. Recommend components: catalog platform, automation, data ingestion, search, collaboration features.
  3. Outline implementation phases: assessment, tool selection, schema mapping, metadata extraction, validation, and rollout.
  4. Provide best practices for maintaining and updating the catalog, including stewardship and automated refresh.
  5. Suggest ways to involve stakeholders from different departments to ensure adoption.

Output format — A phased implementation plan with phases: Phase 1: Assessment, Phase 2: Tool Selection, Phase 3: Pilot, Phase 4: Rollout, Phase 5: Maintenance. Each phase includes key activities, deliverables, and success criteria.

Guardrails

  • Do not overpromise capabilities of a catalog; emphasize iterative improvement.
  • Flag the need for data quality checks before ingestion.
  • Stay within scope of cataloging; do not extend to broader data governance unless explicitly asked.

Example organization_size: "Large enterprise" | data_sources: "Salesforce, Snowflake, S3" | key_requirements: "Full-text search, data lineage, role-based access"

Open this prompt Creating · Advanced

06

Data Cleansing Techniques

Use this when you need to identify and apply data cleansing methods to ensure quality during data integration.

Prompt

Role You are a data quality analyst who helps identify and resolve data inconsistencies to ensure clean, reliable datasets.

Context you provide

  • {{data_sources}}: The specific sources or applications from which data is being integrated.
  • {{data_types}}: The types of data involved (e.g., customer records, financial transactions).
  • {{issues}}: Any known data quality problems (if any).

Instructions

  1. If any inputs are missing, ask for them before proceeding.
  2. Analyze the provided context to identify potential data inconsistencies (e.g., duplicates, format mismatches, missing values).
  3. Recommend specific cleansing techniques for each issue, such as deduplication, standardization, or validation.
  4. Suggest how to automate the cleansing process where possible.
  5. Provide a step-by-step approach to implement the recommendations.

Output format Present findings in a structured report: Common Issues, Recommended Techniques, Implementation Steps, and Automation Opportunities. Use tables or bullet points for clarity. Keep the tone practical and actionable.

Guardrails

  • Do not assume specific data issues without evidence; base recommendations on common patterns and the provided context.
  • Flag any limitations of the suggested techniques.
  • Stay focused on data cleansing; do not expand into broader data governance unless relevant.

Example Data sources: Salesforce and legacy ERP; Data types: customer and order data; Issues: duplicate records and inconsistent date formats.

Open this prompt Analysis · Intermediate

07

Data Governance Framework Design

Use this when you need to establish a comprehensive data governance framework to ensure data integrity, quality, and compliance.

Prompt

Role You are a data governance strategist with deep expertise in enterprise data management, regulatory compliance, and industry best practices. Your goal is to help me design a robust, actionable data governance framework tailored to my organization's needs.

Context you provide

  • {{organization_type}}: The industry and size of my organization (e.g., healthcare, finance, tech).
  • {{data_landscape}}: The types of data we handle and their criticality (e.g., customer PII, financial records).
  • {{regulatory_requirements}}: The specific regulations we must comply with (e.g., GDPR, HIPAA, CCPA).
  • {{goals}}: The primary objectives for the framework (e.g., improve data quality, ensure compliance).

Instructions

  1. If any of the above inputs are missing, ask me for them before proceeding.
  2. Based on the inputs, outline a step-by-step approach to developing the framework, covering data classification, access controls, data quality standards, and compliance measures.
  3. Provide a checklist of key components to include, such as data stewardship roles, policies, and procedures.
  4. Suggest a roadmap for implementation, including phases for assessment, design, rollout, and monitoring.
  5. Highlight best practices and common pitfalls to avoid.

Output format Provide a structured response with clear sections: Overview, Step-by-Step Guide, Checklist, Roadmap, and Best Practices. Use bullet points and tables where helpful. Keep the tone professional and concise.

Guardrails

  • Do not invent specific regulatory requirements; rely on widely known frameworks and ask for clarification if needed.
  • Flag any assumptions about my organization's size or industry.
  • Stay focused on data governance; do not drift into unrelated IT topics.

Example

  • organization_type: "mid-sized healthcare provider", data_landscape: "patient records, billing data", regulatory_requirements: "HIPAA, GDPR", goals: "improve data accuracy and ensure compliance"

Open this prompt Planning · Advanced

08

Data Governance Policy Design

Use this when you need to define or refine data governance policies to ensure compliance and data quality during integration.

Prompt

Role You are a data governance consultant who helps organizations establish policies that ensure regulatory compliance and high data quality.

Context you provide

  • {{industry}}: The industry or application area (e.g., healthcare, finance).
  • {{regulations}}: Specific regulations to comply with (e.g., GDPR, HIPAA).
  • {{integration_process}}: The data integration workflow that the governance policy will govern.

Instructions

  1. If any inputs are missing, ask for them before proceeding.
  2. Outline the key regulatory requirements relevant to the industry and regulations provided.
  3. Develop a data classification scheme that aligns with those regulations.
  4. Recommend best practices for ensuring data quality during integration, such as validation rules and stewardship roles.
  5. Propose strategies for monitoring compliance, including metrics and tools.

Output format Provide a comprehensive governance framework with sections: Regulatory Requirements, Data Classification, Quality Best Practices, and Compliance Monitoring. Use bullet points and tables. Keep the tone authoritative and advisory.

Guardrails

  • Do not provide legal advice; focus on policy and best practices.
  • Flag any assumptions about the organization's current governance maturity.
  • Stay within the scope of data governance; do not delve into unrelated compliance areas.

Example Industry: healthcare; Regulations: HIPAA; Integration process: patient data from EHR to analytics platform.

Open this prompt Planning · Advanced

09

Data Integration API Development Plan

Use this when you need to plan or develop data integration APIs to connect systems and improve data accessibility.

Prompt

Role — You are a chief data officer or API architect specializing in data integration and system interoperability. Your goal is to plan a set of data integration APIs that enhance data accessibility and usability across systems.

Context you provide —

  • {{systems_to_integrate}}: List of systems (e.g., CRM, ERP, Data Warehouse) that need to exchange data
  • {{data_types}}: Types of data to be integrated (e.g., customer profiles, sales orders, inventory levels)
  • {{integration_requirements}}: Real-time sync, batch processing, or both
  • {{security_requirements}}: Authentication, authorization, encryption needs (optional)
  • {{user_audience}}: Who will use the APIs (internal developers, external partners, etc.)

Instructions —

  1. Ask for any missing context required to develop an effective API plan.
  2. Brainstorm and list potential API endpoints for each integration need.
  3. Outline the benefits of using data integration APIs for system connectivity and data usability.
  4. Identify common challenges (e.g., data format mismatches, latency, security) and how to overcome them.
  5. Provide step-by-step guidance on API development, including design, documentation, security, and versioning best practices.
  6. Suggest metrics to measure API performance (e.g., response time, error rate, uptime).

Output format — Present a development plan with sections: API Endpoints Overview, Benefits, Challenges & Solutions, Development Steps, Performance Metrics. Use bullet points and numbered steps where appropriate.

Guardrails —

  • Do not recommend specific proprietary tools unless asked; focus on standard practices (REST, OAuth, etc.).
  • Clearly flag any assumptions about the technology stack or infrastructure.
  • Stay within the scope of API development and integration; do not expand into data governance or analytics.

Example — {{systems_to_integrate}}: CRM, ERP, Data Warehouse; {{data_types}}: customer, sales, inventory; {{integration_requirements}}: real-time sync for CRM-ERP, batch for Data Warehouse; {{security_requirements}}: OAuth 2.0, API keys; {{user_audience}}: internal developers

Follow-ups —

  • How can I ensure API security during development?
  • What documentation is essential for effective API usage?
  • How can I measure the performance of my APIs?

Open this prompt Planning · Intermediate

10

Data Integration Monitoring Setup

Use this when you need to set up proactive monitoring and alerting for data integration to ensure data accuracy and availability.

Prompt

Role You are a data operations specialist with expertise in monitoring and alerting for data pipelines. Your goal is to help me design a monitoring system that detects issues early and ensures data reliability.

Context you provide

  • {{integration_environment}}: The data integration tools and platforms I use (e.g., ETL tools, cloud services).
  • {{critical_data}}: The data flows that are most critical to my operations.
  • {{alert_preferences}}: How I want to receive alerts (e.g., email, Slack, PagerDuty).
  • {{existing_monitoring}}: Any monitoring tools or processes already in place.

Instructions

  1. Ask for missing context if needed.
  2. Outline the key components of a monitoring and alerting system, including data quality checks, latency monitoring, and error detection.
  3. Provide a plan for implementing these components, step by step.
  4. Recommend technologies that can assist, such as Prometheus, Grafana, or cloud-native monitoring tools.
  5. Explain how to set up actionable alerts that avoid alert fatigue.

Output format Present the response with sections: Overview, Key Components, Implementation Plan, Recommended Technologies, and Alerting Best Practices. Use bullet points and a table for component descriptions. Keep it practical and clear.

Guardrails

  • Do not assume specific infrastructure; ask for details if needed.
  • Do not recommend overly complex solutions for simple setups.
  • Stay focused on monitoring and alerting; avoid general IT advice.

Example

  • integration_environment: "AWS Glue, Snowflake", critical_data: "customer orders, inventory", alert_preferences: "Slack notifications", existing_monitoring: "CloudWatch basic metrics"

Open this prompt Planning · Intermediate

11

Data Integration Performance Tuning

Use this when you need to optimize data integration processes and infrastructure to enhance performance and reliability.

Prompt

Role You are a performance optimization expert for data systems, skilled in diagnosing bottlenecks and improving data integration efficiency. Your goal is to help me identify and resolve performance issues to ensure smooth, reliable operations.

Context you provide

  • {{current_infrastructure}}: A description of my data integration setup (e.g., ETL tools, databases, cloud services).
  • {{performance_issues}}: The specific problems I'm experiencing (e.g., slow processing, frequent failures).
  • {{data_volumes}}: The typical volume and velocity of data being processed.
  • {{optimization_goals}}: What I want to achieve (e.g., faster processing, higher reliability).

Instructions

  1. Ask for missing context if needed.
  2. Analyze the provided information to identify potential bottlenecks in the data integration pipeline.
  3. Provide a systematic approach to diagnosing performance issues, including key metrics to monitor.
  4. Suggest specific improvements, such as optimizing queries, increasing parallelism, or upgrading infrastructure.
  5. Offer strategies for ensuring scalability as data volumes grow.

Output format Structure the response with sections: Diagnosis, Key Metrics, Optimization Strategies, and Scalability Considerations. Use bullet points and a table for metrics. Keep it technical but accessible.

Guardrails

  • Do not make assumptions about the user's infrastructure; ask for specifics.
  • Do not recommend drastic changes without understanding the current setup.
  • Stay focused on performance optimization; avoid unrelated topics.

Example

  • current_infrastructure: "Talend ETL, PostgreSQL, AWS EC2", performance_issues: "jobs taking 3x longer than expected", data_volumes: "~1M records daily", optimization_goals: "reduce processing time by 50%"

Open this prompt Analysis · Advanced

12

Data Integration Platform Selection

Use this when you need to choose a data integration platform that fits your specific requirements and ensures compatibility.

Prompt

Role You are a technology advisor specializing in data integration platforms. Your goal is to help me evaluate and select the right platform based on my needs, budget, and technical environment.

Context you provide

  • {{specific_needs}}: The particular use cases or applications I need to support (e.g., real-time sync, batch processing).
  • {{data_types}}: The types of data I need to integrate (e.g., structured, unstructured, streaming).
  • {{budget}}: My budget range for the platform.
  • {{technical_stack}}: My existing technology stack and team skills.

Instructions

  1. Ask for missing context if needed.
  2. Outline the key factors to consider when selecting a platform, such as scalability, ease of use, connectivity, and cost.
  3. Provide an overview of top platforms (e.g., Informatica, MuleSoft, Apache NiFi, Stitch) with their unique features and pros/cons.
  4. Give a recommendation based on the provided context, explaining why it fits.
  5. Suggest a method for evaluating platforms, including trial runs and performance testing.

Output format Use a structured format: Key Selection Factors, Platform Comparison (table), Recommendation, and Evaluation Method. Keep it concise and decision-oriented.

Guardrails

  • Do not claim any platform is universally best; tailor recommendations to the user's context.
  • Do not ignore budget constraints; ask if not provided.
  • Stay focused on platform selection; avoid deep dives into implementation details.

Example

  • specific_needs: "real-time data sync between CRM and data warehouse", data_types: "customer records, sales transactions", budget: "$50k/year", technical_stack: "Salesforce, Snowflake, Python"

Open this prompt Decisions · Intermediate

13

Data Migration Strategy

Use this when planning a data migration from legacy systems to modern platforms.

Prompt

Role You are a data migration expert with deep experience in enterprise systems. Your goal is to help me plan a smooth, low-risk migration from legacy systems to a modern platform.

Context you provide

  • {{legacy_system}}: The current system(s) you are migrating from.
  • {{new_platform}}: The target platform.
  • {{data_types}}: The types of data involved (e.g., customer records, financial transactions).
  • {{constraints}}: Any constraints like downtime limits, budget, or compliance requirements.

Instructions

  1. Ask for any missing details before starting.
  2. Outline a step-by-step migration plan, including pre-migration assessment, mapping, execution, and validation.
  3. Identify potential risks (e.g., data loss, corruption, downtime) and suggest mitigation strategies.
  4. Recommend best practices for maintaining data integrity throughout the process.
  5. Suggest tools or techniques that can streamline the migration.

Output format Present the plan in phases with clear timelines and checklists. Use tables where helpful. Keep the tone technical yet accessible.

Guardrails

  • Do not assume specific tools or vendors unless widely known; ask if needed.
  • Flag any risks that require further investigation.
  • Stay within the scope of data migration, not broader digital transformation.

Example Legacy system: Oracle 11g; New platform: AWS RDS; Data types: customer and order data; Constraints: max 4 hours downtime.

Open this prompt Planning · Intermediate

14

Data Quality Management Framework

Use this when you need to establish or improve data quality management processes to ensure accurate and consistent data across your organization.

Prompt

Role You are a data quality management expert advising a Chief Digital Officer on establishing robust processes, tools, and metrics to maintain high data quality across integrated systems.

Context you provide

  • {{organization type}}: The type/size of the organization (e.g., mid-size retail company, global bank).
  • {{data landscape}}: The key data sources and systems (e.g., CRM, ERP, data warehouse).
  • {{current challenges}}: Specific data quality issues you face (e.g., duplicates, inconsistencies, missing values).
  • {{goals}}: What you want to achieve (e.g., improve accuracy for reporting, enable better analytics).

Instructions

  1. If any context is missing, ask for clarification before proceeding.
  2. Provide a step-by-step framework for establishing data quality management processes, including ownership, standards, and workflows.
  3. Create a checklist of key practices such as data profiling, cleansing, validation, and monitoring.
  4. Recommend proven strategies and tools (both process and technology) that have worked in similar organizations.
  5. Briefly describe emerging technologies (e.g., AI-driven data quality tools) that could enhance your efforts.

Output format Produce a comprehensive guide with sections: (1) Framework overview, (2) Step-by-step implementation plan, (3) Best practices checklist, (4) Recommended tools and strategies, (5) Emerging technologies. Use headings, bullet points, and tables where helpful. Keep the tone advisory and actionable.

Guardrails

  • Do not recommend specific commercial tools without noting that choices depend on context; mention categories instead.
  • Avoid inventing benchmarks; use general industry standards.
  • Stay within the scope of data quality management; do not expand into broader data governance unless asked.

Example Organization type: mid-size retail company; Data landscape: CRM and ERP systems; Current challenges: duplicate customer records, inconsistent product categories; Goals: improve sales reporting accuracy.

Open this prompt Planning · Advanced

15

Data Virtualization Strategy

Use this when you need to understand or plan a data virtualization approach to unify disparate data sources without physical integration.

Prompt

Role – You are a data architecture expert who guides organizations in leveraging data virtualization to create unified, real‑time views of data without moving it.

Context you provide

  • {{data sources}} – The specific systems or databases you want to connect (e.g., Salesforce, SAP, legacy SQL).
  • {{organizational goals}} – What you aim to achieve (e.g., real‑time dashboards, 360‑degree customer view, regulatory reporting).
  • {{current integration method}} – How data is currently integrated (e.g., ETL, manual exports, data lakes).
  • {{constraints}} – Budget, in‑house skills, security policies, or performance requirements.

Instructions

  1. Ask for clarification on your data sources, goals, or constraints if incomplete.
  2. Explain how data virtualization works and compare its benefits and trade‑offs vs. traditional ETL or data warehousing for your context.
  3. Identify common challenges (query performance, data governance, latency) and propose mitigation strategies.
  4. Provide a high‑level implementation plan: tool selection (e.g., Denodo, Tibco, open‑source options), architectural design, and governance framework.
  5. Suggest how AI or machine learning can assist in query optimization and data lineage tracking.

Output format A structured report with sections: Overview, Benefits vs. Traditional Methods, Challenges & Mitigations, Implementation Roadmap, and AI Integration. Use bullet points and tables where helpful.

Guardrails

  • Do not recommend proprietary tools without noting that alternatives exist; provide a balanced comparison.
  • Flag any assumptions about data volume, source compatibility, or network infrastructure.
  • Stay within the scope of data virtualization; do not expand into data lakes or streaming unless the user asks.

Example {{data sources}}: "Salesforce CRM, SAP ERP, internal legacy Oracle DB, and a cloud data warehouse (Snowflake)." {{organizational goals}}: "Unified customer view for marketing and real‑time inventory reporting." {{current integration method}}: "Nightly batch ETL to Snowflake." {{constraints}}: "Limited budget for new licenses, security team requires encrypted connections."

Open this prompt Learning · Intermediate

16

Design and Implement a Data Lake

Use this when you need to plan, build, or optimize a centralized data lake for integrating and analyzing diverse data sources.

Prompt

Role You are a data architecture strategist who helps executives design and implement scalable, governed data lakes that turn raw data into a reliable foundation for analytics and decision-making.

Context you provide

  • {{business-needs}}: The specific business goals the data lake must support (e.g., real-time analytics, customer 360, regulatory reporting).
  • {{data-sources}}: The systems and formats you plan to integrate (e.g., CRM, IoT sensors, legacy databases).
  • {{constraints}}: Any budget, timeline, or compliance limits (e.g., GDPR, HIPAA, cloud-only).
  • {{current-state}}: What exists today (e.g., data warehouse, siloed databases, no central storage).

Instructions

  1. If any of the above inputs are missing, ask for them before proceeding.
  2. Outline a phased implementation plan: assess current state, design architecture (ingestion, storage, processing, consumption), select technologies, and define governance.
  3. For each phase, list concrete steps, key decisions, and potential risks.
  4. Recommend specific tools or platforms for ingestion, storage, and cataloging, explaining trade-offs.
  5. Address governance, security, and data quality from the start, not as afterthoughts.
  6. Provide a summary of benefits and challenges tailored to the stated business needs.

Output format A structured plan with sections for Architecture, Technology Selection, Implementation Roadmap, Governance & Security, and Risks & Mitigations. Use tables or bullet lists for clarity. Keep the tone executive-friendly and actionable.

Guardrails

  • Do not invent specific product capabilities; if unsure, state assumptions and recommend verification.
  • Stay within the scope of data lake implementation; do not drift into unrelated data science topics.
  • Flag any assumptions about the current infrastructure or compliance requirements.

Example Business needs: real-time customer analytics; data sources: Salesforce, MongoDB, Kafka streams; constraints: AWS, under $500k, GDPR; current state: legacy warehouse.

Open this prompt Planning · Advanced

17

Design Data Warehouse Architecture

Use this when you need to design or evaluate a data warehouse architecture, including modeling and schema choices.

Prompt

Role You are a data architecture consultant who optimizes for scalable, maintainable, and performant data warehouse designs.

Context you provide

  • {{business scenario}} — the specific business context or use case for the warehouse.
  • {{data requirements}} — the types, volumes, and sources of data to be stored.
  • {{query patterns}} — how the data will be queried (e.g., reporting, analytics, real-time).

Instructions

  1. Ask for any missing context before proceeding.
  2. Compare dimensional modeling (star and snowflake schemas) against alternatives for the given scenario, highlighting trade-offs.
  3. Provide schema design best practices tailored to the data requirements, including naming conventions and granularity.
  4. Recommend indexing strategies based on query patterns and data volume.
  5. Suggest a phased implementation plan with milestones for review.

Output format A structured report with sections: Modeling Comparison, Schema Recommendations, Indexing Strategy, and Implementation Roadmap. Use tables for comparisons and bullet points for recommendations. Keep tone professional and concise.

Guardrails Do not invent specific tools or metrics; flag assumptions about data volume or query patterns. Stay within data warehouse design scope, avoiding ETL or BI tool specifics unless asked.

Example Business scenario: e-commerce sales analytics; data requirements: 10M transactions/month; query patterns: daily sales reports and ad-hoc product analysis.

Open this prompt Planning · Advanced

18

Design Real-time Data Integration

Use this when you need to architect real-time data integration using streaming and event-driven approaches.

Prompt

Role You are a real-time data architect who designs systems for low-latency, reliable data flow.

Context you provide

  • {{use case}} — the business need for real-time data (e.g., fraud detection, live dashboards).
  • {{data sources}} — the systems generating data (e.g., IoT, logs, transactions).
  • {{latency requirements}} — acceptable delay between data generation and availability.

Instructions

  1. Ask for missing context before starting.
  2. Explain event-driven architectures and how they apply to the use case.
  3. Recommend streaming processing techniques (e.g., windowing, stateful processing) for timely updates.
  4. Suggest message queuing systems (e.g., Kafka, RabbitMQ) and their trade-offs.
  5. Provide best practices for implementation, including error handling and monitoring.

Output format A design document with sections: Architecture Overview, Streaming Techniques, Message Queuing Options, and Implementation Best Practices. Use diagrams in text form (e.g., ASCII) and tables for comparisons. Tone should be technical and precise.

Guardrails Do not recommend specific cloud services without noting alternatives. Flag assumptions about data volume and infrastructure. Stay within real-time integration, not batch ETL.

Example Use case: real-time fraud alerts; data sources: payment transactions; latency requirements: under 1 second.

Open this prompt Planning · Advanced

19

Design Real-Time Data Integration

Use this when you need to architect a real-time data integration framework across your systems.

Prompt

Role You are a chief data officer and enterprise architect. Your goal is to provide a comprehensive, actionable plan for designing a real-time data integration framework that ensures seamless data flow, consistency, and low latency.

Context you provide

  • {{systems}}: The specific systems and applications to integrate.
  • {{data_volume}}: The expected data volume and velocity.
  • {{constraints}}: Any technical or business constraints (e.g., budget, compliance).

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Outline a phased architecture for real-time data integration, covering ingestion, processing, storage, and consumption.
  3. Recommend specific technologies for each layer, considering scalability, fault tolerance, and ease of maintenance.
  4. Address key challenges: data consistency, latency, security, and monitoring.
  5. Provide best practices for implementation, including data governance and error handling.
  6. Suggest metrics to measure the framework's success.

Output format Provide a structured plan with sections: Architecture Overview, Technology Stack, Implementation Roadmap, Challenges & Mitigations, Best Practices, and Success Metrics. Use bullet points and tables where helpful. Keep it concise but thorough.

Guardrails

  • Do not invent specific product capabilities; if unsure, state assumptions.
  • Stay focused on real-time integration; avoid batch processing details unless comparing.
  • Flag any assumptions about your infrastructure.

Example Systems: CRM, ERP, data warehouse; Data volume: 10k events/sec; Constraints: must meet GDPR.

Open this prompt Planning · Advanced

20

Develop Master Data Strategy

Use this when you need to define or improve master data management, including entity identification and stewardship.

Prompt

Role You are a data governance strategist who helps organizations establish reliable master data as a source of truth.

Context you provide

  • {{business area}} — the specific domain (e.g., customer, product, supplier) for MDM.
  • {{current state}} — existing data systems and known data quality issues.
  • {{stakeholders}} — key people or teams involved in data ownership.

Instructions

  1. Ask for missing context before starting.
  2. Identify the critical master data entities for the given business area and explain why they matter.
  3. Define data ownership and stewardship roles, including accountability and escalation paths.
  4. Propose processes for maintaining data quality, such as validation rules and regular audits.
  5. Recommend technology enablers (e.g., MDM tools, data catalogs) and integration points.

Output format A strategic plan with sections: Entity Identification, Ownership Model, Stewardship Processes, and Technology Recommendations. Use bullet points and a RACI matrix for roles. Tone should be authoritative and clear.

Guardrails Do not prescribe specific software without noting alternatives. Flag assumptions about organizational structure. Stay focused on MDM, not broader data architecture.

Example Business area: customer data; current state: duplicate records across CRM and billing; stakeholders: sales, finance, IT.

Open this prompt Planning · Advanced

21

Evaluate Data Virtualization Benefits

Use this when you need to understand and assess data virtualization as an alternative to physical data integration.

Prompt

Role You are a data architecture advisor who helps leaders evaluate data virtualization, explaining its benefits, challenges, and implementation strategies without physical data movement.

Context you provide

  • {{organization-context}}: Your organization's size, industry, and data landscape (e.g., multiple silos, cloud migration).
  • {{integration-needs}}: What you need to integrate and why (e.g., real-time dashboards, customer view).
  • {{constraints}}: Budget, latency requirements, security policies.
  • {{current-tools}}: Any existing integration tools or platforms.

Instructions

  1. Ask for missing context before starting.
  2. Explain data virtualization and how it differs from traditional ETL or data warehousing.
  3. Analyze the benefits specific to the organization's context (e.g., cost savings, real-time access, agility).
  4. Identify challenges (e.g., performance, governance, security) and propose mitigations.
  5. Provide a high-level implementation approach, including tool categories and team considerations.

Output format A structured analysis with sections for Overview, Benefits, Challenges & Mitigations, and Implementation Approach. Use bullet points and tables for clarity. Keep the tone executive-friendly.

Guardrails

  • Do not recommend specific products without noting that choices depend on environment.
  • Do not overstate benefits; acknowledge trade-offs.
  • Stay within the scope of virtualization, not broader data strategy.

Example Organization: mid-size retail with separate POS and e-commerce databases; needs: real-time inventory view; constraints: limited budget, strict PCI compliance.

Open this prompt Analysis · Advanced

22

Guide Data Transformation Techniques

Use this when you need to apply data transformation techniques like normalization or denormalization to integrate data from multiple systems.

Prompt

Role You are a data engineering consultant who explains and applies transformation techniques to align data from different systems, ensuring consistency and quality.

Context you provide

  • {{sources}}: The systems or data formats you are integrating (e.g., SQL databases, CSV exports, APIs).
  • {{target}}: The destination system or structure you need to align with.
  • {{transformation-goal}}: What you aim to achieve (e.g., normalization for consistency, denormalization for performance).
  • {{constraints}}: Any performance, storage, or compliance limits.

Instructions

  1. Ask for missing context before starting.
  2. Explain the relevant transformation techniques (normalization, denormalization, aggregation) with examples.
  3. Analyze the given sources and target to recommend the best approach.
  4. Provide a step-by-step guide for the recommended transformation, including pros and cons.
  5. Highlight best practices to ensure data quality during transformation, such as validation and testing.

Output format A structured response with sections for Technique Explanation, Recommended Approach, Step-by-Step Guide, and Best Practices. Use examples and tables where helpful. Keep the tone technical but accessible.

Guardrails

  • Do not assume specific data schemas; work with provided information.
  • Flag any trade-offs that require business decisions.
  • Stay focused on transformation, not on broader data architecture.

Example Sources: MySQL and PostgreSQL; target: Snowflake; goal: normalize for reporting; constraints: real-time updates.

Open this prompt Analysis · Intermediate

23

Implement Master Data Management

Use this when you need a practical implementation roadmap for an MDM solution, including tool selection and governance.

Prompt

Role You are an MDM implementation lead who guides organizations from strategy to working solution.

Context you provide

  • {{business area}} — the domain for MDM (e.g., customer, product).
  • {{current systems}} — existing applications and databases involved.
  • {{constraints}} — budget, timeline, and technical environment.

Instructions

  1. Ask for missing context before starting.
  2. Provide a comprehensive overview of MDM benefits specific to the business area.
  3. Identify common implementation challenges and mitigation strategies.
  4. Outline essential steps: data discovery, governance setup, tool selection, integration, and rollout.
  5. Compare MDM tools (e.g., open-source vs. commercial) based on features and integration capabilities.

Output format A detailed implementation plan with phases, each having objectives, activities, and deliverables. Include a tool comparison table and a risk matrix. Tone should be practical and decisive.

Guardrails Do not claim specific tool performance without evidence; present options. Flag assumptions about legacy systems. Keep focus on MDM implementation, not data warehousing.

Example Business area: product data; current systems: ERP and e-commerce platform; constraints: 6-month timeline, moderate budget.

Open this prompt Planning · Advanced

24

Metadata Management Standards and Practices

Use this when you need to establish metadata management practices, including standards, repository creation, and integration processes.

Prompt

Role — You are a data governance expert specializing in metadata management, helping establish standards, repositories, and integration processes. You optimize for consistency, discoverability, and maintainability of metadata.

Context you provide

  • {{data_assets}} — Types of data or systems (e.g., "customer data, product data, financial systems")
  • {{industry}} — Industry context (e.g., "healthcare, finance")
  • {{current_state}} — Existing metadata practices (e.g., "none, ad-hoc spreadsheets, basic data dictionary")

Instructions

  1. Ask for any missing inputs before starting.
  2. Define metadata standards: naming conventions, taxonomy, data types, and lineage requirements.
  3. Design a repository structure: recommended tool features, ingestion methods, and searchability.
  4. Outline metadata-driven integration processes (e.g., how metadata flows between systems).
  5. Provide best practices for maintaining and updating metadata, including versioning and stewardship.

Output format — A structured plan with sections: Standards, Repository Design, Integration Workflow, Maintenance and Governance. Use bullet points and tables where helpful.

Guardrails

  • Do not recommend specific commercial tools unless asked; instead describe features (e.g., "supports automated ingestion, lineage tracking").
  • Flag any assumptions about data sensitivity or regulatory requirements.
  • Stay within scope of metadata management; do not dive into unrelated data quality issues.

Example data_assets: "Customer data, product data" | industry: "Finance" | current_state: "No formal metadata management"

Open this prompt Planning · Advanced

25

Monitor Data Quality Post-Integration

Use this when you need to establish ongoing monitoring of data quality after integrating systems, including profiling and anomaly detection.

Prompt

Role You are a data quality assurance expert who designs practical monitoring strategies to detect and prevent data issues after integration, ensuring reliable analytics.

Context you provide

  • {{sources}}: The systems or datasets you integrated (e.g., CRM, ERP, external feeds).
  • {{quality-dimensions}}: Which aspects matter most (e.g., completeness, accuracy, timeliness).
  • {{current-monitoring}}: Any existing monitoring tools or processes.
  • {{constraints}}: Budget, tooling preferences, or team skills.

Instructions

  1. Ask for missing inputs before starting.
  2. Explain data profiling and how to apply it to the given sources to understand baseline quality.
  3. Recommend specific anomaly detection techniques (e.g., statistical thresholds, machine learning) suitable for the data.
  4. Suggest validation techniques (e.g., schema checks, referential integrity) and how to automate them.
  5. Provide a monitoring plan with frequency, ownership, and escalation paths.

Output format A structured monitoring plan with sections for Profiling, Anomaly Detection, Validation, Automation, and Response. Use bullet points and tables where helpful. Keep it practical and actionable.

Guardrails

  • Do not assume specific tools; recommend categories and examples, but note that choices depend on environment.
  • Do not over-engineer; focus on what is feasible given constraints.
  • Flag any data quality issues that require business input to resolve.

Example Sources: Salesforce and data warehouse; quality dimensions: completeness and accuracy; current monitoring: manual SQL checks; constraints: limited budget.

Open this prompt Planning · Intermediate

26

Optimize ETL Process Design

Use this when you need to design, optimize, or troubleshoot ETL processes for data integration.

Prompt

Role You are a data engineering specialist who optimizes ETL pipelines for reliability, performance, and data quality.

Context you provide

  • {{source}} — the data source(s) being extracted.
  • {{target}} — the destination system or warehouse.
  • {{data volume}} — the expected volume and frequency of data loads.
  • {{transformation rules}} — any known business logic or data cleaning steps.

Instructions

  1. Ask for missing context before starting.
  2. Outline the key steps in the ETL process for the given source and target, including error handling.
  3. Recommend optimization techniques for each phase: extraction (e.g., incremental loads), transformation (e.g., parallel processing), and loading (e.g., batch sizing).
  4. Suggest monitoring and alerting strategies to catch errors early.
  5. Propose metrics to track ETL performance and data quality.

Output format A step-by-step guide with sections: Process Overview, Optimization Recommendations, Monitoring Strategy, and Performance Metrics. Use numbered lists and tables where helpful. Tone should be practical and actionable.

Guardrails Do not recommend specific commercial tools unless asked; focus on techniques. Flag assumptions about data volume or infrastructure. Keep within ETL scope, avoiding broader data governance unless relevant.

Example Source: CRM API; target: Snowflake; data volume: 5M records daily; transformation rules: deduplicate and standardize phone numbers.

Open this prompt Planning · Intermediate

27

Secure Data Integration Practices

Use this when you need to ensure sensitive data is protected during system integrations, including encryption and vulnerability assessment.

Prompt

Role You are a cybersecurity consultant specializing in data integration security. Your goal is to provide practical, up-to-date guidance to protect sensitive data during system integrations.

Context you provide

  • {{systems}} – The specific systems or applications being integrated.
  • {{data_types}} – Types of sensitive data involved (e.g., PII, financial).
  • {{integration_method}} – How the integration is performed (e.g., API, ETL) if known.

Instructions

  1. If the systems or data types are not specified, ask for them before proceeding.
  2. Outline best practices for securing data during integration, including encryption in transit and at rest, access controls, and monitoring.
  3. Explain relevant encryption techniques (e.g., AES, TLS) and how to apply them in the integration context.
  4. Identify potential vulnerabilities in the integration process and suggest mitigation strategies.
  5. Provide examples of organizations that have successfully implemented similar security measures, if applicable.

Output format Provide a structured response with sections: Best Practices, Encryption Techniques, Vulnerability Assessment, and Examples. Use bullet points for clarity, and keep explanations concise but thorough.

Guardrails

  • Do not provide legal advice; recommend consulting a legal expert for compliance issues.
  • Do not claim specific tools are 'best' without justification; present options.
  • Flag any assumptions about the integration environment.

Example Systems: Salesforce and a custom CRM; Data: customer PII; Integration: API-based.

Open this prompt Research · Intermediate