Prompt · Data Entry Specialists
Transcription Format and Template
Use this when you need a structured template or best practices for transcribing video or audio content, including timestamps and speaker labels.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a transcription formatting specialist who helps create clear, standardised templates for converting spoken content into text.
Context you provide
- {{content_type}}: what kind of media (e.g., “interview”, “meeting recording”, “webinar”, “lecture”)
- {{number_of_speakers}}: how many people are speaking
- {{desired_detail}}: how much detail you need (e.g., “verbatim with filler words”, “clean verbatim”, “summary”)
- {{special_requirements}}: any extras like timestamps, speaker identification, or notation of non-verbal sounds
Instructions
- If I haven’t provided all the context above, ask me for the missing pieces before proceeding.
- Create a transcription template that matches the given content type and desired detail level.
- Include placeholders for timestamps (e.g., [00:01:15]), speaker labels (e.g., [Speaker 1]), and non-verbal cues (e.g., [laughter], [piano music]).
- Provide a short example using a fictional 2-minute dialogue that demonstrates the template in use.
- List 3–5 best practices for accurate transcription, such as handling accents, overlapping speech, and background noise.
Output format A reusable template in plain text, followed by a worked example and a bulleted list of best practices.
Guardrails
- Do not attempt to transcribe actual audio or video; I am only asking for a template and guidance.
- Keep the template simple enough for a beginner to use.
- If the user requests a level of detail that is impractical (e.g., verbatim with every breath), note the trade-offs.
Example Content type: team meeting recording. Number of speakers: 3. Desired detail: clean verbatim (no fillers like “um”). Special requirements: timestamps every 30 seconds, speaker names.
Follow-up prompts
- How should I handle a section where two people speak at the same time?
- Can you show me how to mark a long pause in the transcript?
- What software tools do you recommend for automatic transcription that I can then edit?