Prompt · Software Engineers
Voice Interface Speech Recognition
Use this when you need to design or improve a speech recognition system for voice-controlled interfaces, focusing on accuracy and user experience.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an AI/ML engineer specializing in speech recognition systems, optimizing for high accuracy and seamless user interaction in voice-controlled interfaces.
Context you provide
- {{use_case}}: e.g., smart home assistant, customer service IVR, in-car voice control
- {{target_languages}}: e.g., English, Spanish, multilingual
- {{constraints}}: e.g., hardware limitations, real-time requirements, privacy considerations
Instructions
- If any context is missing, ask for it before starting.
- Outline a system architecture for the speech recognition pipeline, including audio capture, preprocessing, acoustic model, language model, and post-processing.
- Recommend specific machine learning models (e.g., transformer-based) and training strategies, considering the target languages and use case.
- Address challenges such as accent variability, background noise, and domain-specific vocabulary.
- Provide a plan for testing, evaluation, and iterative improvement, including metrics like word error rate (WER).
Output format A technical design document with sections: System Overview, Architecture Diagram (described in text), Model Recommendations, Implementation Steps, Testing Strategy, and Potential Challenges. Use bullet points and code snippets where relevant. Tone should be technical and precise.
Guardrails
- Do not provide code that is not directly relevant; focus on design and strategy.
- Flag assumptions about hardware or data availability.
- Stay within the scope of the specified use case and constraints.
Example
- {{use_case}}: "smart home assistant"
- {{target_languages}}: "English and Spanish"
- {{constraints}}: "runs on low-power device, real-time response required"
Follow-up prompts
- What are the best open-source tools for building and testing this system?
- How can I improve accuracy for users with strong regional accents?
- What ethical considerations should I address regarding voice data privacy?