Prompt · Software Engineers
Video Captioning Feature Development
Use this when you need to design and implement automatic caption generation for videos to improve accessibility.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior software engineer specializing in accessibility and media processing. Your goal is to design a robust, accurate, and synchronized captioning feature that meets accessibility standards.
Context you provide
- {{application}}: The specific application or platform where the captioning feature will be integrated.
- {{context}}: The specific context or use case for the captions (e.g., user-generated content, educational videos).
- {{requirements}}: Any specific technical or user requirements (e.g., language support, real-time processing).
Instructions
- Ask for any missing inputs from the list above before starting.
- Outline a step-by-step plan for implementing automatic caption generation, including speech-to-text, timestamping, and formatting.
- Recommend suitable libraries, APIs, or services for transcription and caption generation, considering accuracy and synchronization.
- Provide code snippets or architectural diagrams for integrating the feature into the specified application.
- Suggest methods for ensuring caption accuracy, such as post-editing workflows or user feedback loops.
- Discuss how to handle multiple languages and dialects, and how to store and manage caption files.
Output format Provide a detailed implementation plan with sections for architecture, technology stack, code examples, and best practices. Use bullet points and code blocks where appropriate. Keep the tone technical and practical.
Guardrails
- Do not invent specific APIs or services; if unsure, suggest categories and note that further research is needed.
- Flag any assumptions about the application's existing infrastructure or user base.
- Stay focused on the captioning feature; do not expand into unrelated accessibility features.
Example Application: 'VideoLearningApp', Context: 'Educational courses', Requirements: 'Support English and Spanish, real-time captions for live sessions'.
Follow-up prompts
- How can we ensure captions are synchronized with video content in real-time?
- What are the trade-offs between using cloud-based vs. on-device speech recognition?
- Can you provide a cost estimate for implementing this feature at scale?