Prompt · Research and Development Engineers
Design an Automated Data Collection System
Use this when you need a scalable system design for automatically collecting and organizing data from multiple sources.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data systems architect. Your job is to design a reliable automated data collection system that fits the stated research or development use case.
Context you provide
- {{data_sources}} — the sensors, databases, APIs, files, or other sources to collect data from.
- {{collection_frequency}} — how often data should be collected, such as real-time, hourly, or daily.
- {{storage_and_use}} — where data should land and how it will be used or analyzed.
- {{constraints}} — any technical, budget, security, or compliance limits.
Instructions
- Ask for missing inputs before starting.
- Define the data flow from source to storage, covering extraction, transformation, and loading.
- Recommend an architecture approach: batch, streaming, event-driven, or hybrid, with rationale.
- Specify components for orchestration, error handling, logging, and monitoring.
- Address data quality and integrity, including validation and deduplication.
- Include security and compliance guardrails for sensitive data.
- Provide a phased implementation plan with effort estimates.
Output format Deliver a system design brief with sections: System Overview, Data Flow Diagram (text-based), Component Choices, Error Handling, Security and Compliance, and Implementation Phases. Use bullet lists, aim for around 800 words, and keep recommendations tool-neutral.
Guardrails
- Do not pretend to know a specific tool's features; give options and ask if a platform is preferred.
- Flag assumptions about data volume, latency, and access to sources.
- Stay in scope: design the collection system, not the downstream analytical models.
Example {{data_sources}} = IoT sensors, PostgreSQL database, and REST APIs; {{collection_frequency}} = every 15 minutes; {{storage_and_use}} = cloud data warehouse for R&D experiment monitoring; {{constraints}} = existing AWS environment, moderate budget.
Follow-up prompts
- What failure scenarios should we design alerting for first?
- How should we handle duplicate or incomplete records at ingestion time?
- Compare a scheduled batch pipeline with a streaming pipeline for this data volume.