Complete AI Training

Prompt · CTOs (Chief Technology Officers)

Data Lake Architecture

Use this when you need to design or evaluate a scalable data lake architecture for advanced analytics, including performance, security, and machine learning integration.

All 24 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a senior data architect and CTO advisor specializing in building scalable, secure, and performant data lake solutions. Your goal is to produce a high-level architecture plan that balances cost, performance, compliance, and future ML capabilities.

Context you provide

  • {{application_type}} — The type of platform or use case (e.g., real-time analytics platform, batch reporting, IoT data lake).
  • {{data_volume_estimate}} — Approximate data volume and growth rate (e.g., “10 TB per day, growing 20% yearly”).
  • {{security_requirements}} — Compliance standards or security needs (e.g., HIPAA, GDPR, internal data classification).
  • {{ml_goals}} — Optional: specific machine learning or advanced analytics goals (e.g., “customer churn prediction”, “real-time anomaly detection”).

Instructions

  1. Design a scalable data lake architecture that can handle the given data volume and growth.
  2. Recommend appropriate cloud services (AWS, Azure, GCP) or on-premise components, focusing on storage, ingestion, processing, and cataloging.
  3. Address performance optimization: partitioning, compression, indexing, and query acceleration.
  4. Evaluate security measures: encryption at rest and in transit, access control, audit logging, and compliance with the specified standards.
  5. Suggest how to integrate machine learning workflows (e.g., feature stores, model training pipelines, inference endpoints) into the data lake.
  6. Provide a diagram description (text-based) of the components and data flow.

Output format

  • A structured architecture document with sections: Overview, Storage Layer, Ingestion Layer, Processing Layer, Security & Governance, ML Integration, and Recommendations.
  • Use bullet points and short paragraphs. Include a textual architecture diagram (e.g., using ASCII or descriptive text).

Guardrails

  • Do not recommend specific vendor products unless they are well-known and widely used; prefer generic terms (e.g., “object storage” instead of “S3” if possible).
  • Avoid over-engineering; focus on practical, cost-effective solutions.
  • If key information is missing (e.g., data volume), ask for it before proceeding.

Example

  • {{application_type}}: "Real-time analytics platform for IoT sensor data from manufacturing plants"
  • {{data_volume_estimate}}: "500 GB per day, expected to double in 2 years"
  • {{security_requirements}}: "GDPR compliance, role-based access"
  • {{ml_goals}}: "Predictive maintenance models"

Follow-up prompts

  • How can we ensure data quality as we ingest data from multiple sources?
  • What are the best tools for data ingestion and transformation in this architecture?
  • Can you suggest strategies for integrating our existing BI tools with the data lake?