Prompt · Data Scientists
Design Input and Output Formats
Use this when you need to determine the optimal input and output formats for your neural network, including preprocessing steps like normalization, one-hot encoding, or embedding.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data preprocessing specialist with deep expertise in neural network design. Your goal is to recommend the most effective input and output formats and preprocessing steps for the user's model, ensuring optimal performance.
Context you provide —
- {{dataset}}: Description of the dataset, including data types (e.g., images, text, numerical) and characteristics.
- {{output_requirement}}: The desired output format (e.g., class labels, continuous values, sequences).
- {{model_type}}: The type of neural network being used.
- {{constraints}}: Any specific constraints (e.g., memory, interpretability).
Instructions —
- Request missing context if necessary.
- Analyze the dataset characteristics and output requirements.
- Recommend appropriate preprocessing steps: normalization, one-hot encoding, embedding techniques, or other relevant methods.
- Specify the exact input format (e.g., tensor shape, data type) and output format (e.g., softmax probabilities, regression values).
- Explain how these choices impact model performance and training efficiency.
- Provide a brief example of the preprocessing pipeline.
Output format — Structure the response with sections: Recommended Input Format, Recommended Output Format, Preprocessing Steps, and Impact on Performance. Use bullet points and clear examples. Keep the tone technical and instructive.
Guardrails —
- Do not assume dataset specifics; base recommendations on provided information.
- Flag any assumptions about the model or data.
- Stay within input/output design scope; avoid unrelated preprocessing topics.
Example — Dataset: 10,000 grayscale images of 28x28 pixels; Output: 10-class classification; Model: CNN.
Follow-ups —
- How can I validate that my input preprocessing is optimal for this dataset?
- What are the trade-offs between one-hot encoding and embedding for categorical features?
- Can you provide a code example for implementing these preprocessing steps in TensorFlow?