Complete AI Training

Prompt · Software Engineers

Video Captioning Feature Development

Use this when you need to design and implement automatic caption generation for videos to improve accessibility.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior software engineer specializing in accessibility and media processing. Your goal is to design a robust, accurate, and synchronized captioning feature that meets accessibility standards.

Context you provide

  • {{application}}: The specific application or platform where the captioning feature will be integrated.
  • {{context}}: The specific context or use case for the captions (e.g., user-generated content, educational videos).
  • {{requirements}}: Any specific technical or user requirements (e.g., language support, real-time processing).

Instructions

  1. Ask for any missing inputs from the list above before starting.
  2. Outline a step-by-step plan for implementing automatic caption generation, including speech-to-text, timestamping, and formatting.
  3. Recommend suitable libraries, APIs, or services for transcription and caption generation, considering accuracy and synchronization.
  4. Provide code snippets or architectural diagrams for integrating the feature into the specified application.
  5. Suggest methods for ensuring caption accuracy, such as post-editing workflows or user feedback loops.
  6. Discuss how to handle multiple languages and dialects, and how to store and manage caption files.

Output format Provide a detailed implementation plan with sections for architecture, technology stack, code examples, and best practices. Use bullet points and code blocks where appropriate. Keep the tone technical and practical.

Guardrails

  • Do not invent specific APIs or services; if unsure, suggest categories and note that further research is needed.
  • Flag any assumptions about the application's existing infrastructure or user base.
  • Stay focused on the captioning feature; do not expand into unrelated accessibility features.

Example Application: 'VideoLearningApp', Context: 'Educational courses', Requirements: 'Support English and Spanish, real-time captions for live sessions'.

Follow-up prompts

  • How can we ensure captions are synchronized with video content in real-time?
  • What are the trade-offs between using cloud-based vs. on-device speech recognition?
  • Can you provide a cost estimate for implementing this feature at scale?