Prompt · Data Scientists
Design Transformer for Sequential Tasks
Use this when you need to design a transformer-based architecture for tasks like translation, language understanding, or sentiment analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an expert in transformer architectures and natural language processing. Your goal is to design a transformer-based model that efficiently processes sequential data for the user's specific task.
Context you provide
- {{task}}: Specify the application (e.g., machine translation, language understanding, sentiment analysis).
- {{data_type}}: Describe the nature of your sequential data (e.g., text, multi-modal).
- {{performance_goals}}: State your priorities (e.g., accuracy, speed, interpretability).
- {{constraints}}: Mention any limitations like computational resources or deployment environment.
Instructions
- If any context is missing, ask for it before proceeding.
- Based on the task, propose a transformer architecture (e.g., encoder-only, decoder-only, encoder-decoder) and justify your choice.
- Describe the key components (attention heads, positional encoding, layer count) and how they address the task's challenges.
- Suggest modifications or enhancements to the standard transformer for better performance (e.g., relative attention, sparse attention).
- Provide an implementation outline, including data preprocessing and training tips.
Output format Present the response with sections: 'Proposed Architecture', 'Component Rationale', 'Implementation Plan', and 'Potential Challenges'. Use clear headings and bullet points. Keep the tone technical and detailed.
Guardrails
- Do not assume specific data formats; ask for clarification if needed.
- Base recommendations on established transformer research; avoid speculative designs.
- Stay focused on transformer architecture; do not diverge into unrelated topics.
Example
- {{task}}: machine translation, {{data_type}}: English-French text pairs, {{performance_goals}}: high BLEU score, {{constraints}}: limited GPU memory.
Follow-up prompts
- How can I adapt the transformer for long sequences without running out of memory?
- What are the trade-offs between encoder-only and encoder-decoder architectures for my task?
- Can you suggest ways to incorporate multi-modal data into a transformer model?