Complete AI Training

Skill · Video

Multimodal whisper

Transcribes audio files to text or English translation with Whisper, including batch runs and word-level timestamps. Use when the user supplies an audio file path and wants a transcript, translation, subtitles, or word timings.

Complete AI SkillsLicense: MITAdded Sep 29, 2026

How to use it

  1. Start your plan and connect your AI once
  2. Ask for the task in your own words, or say it directly:
Use the Multimodal whisper skill to help me with this.

Without a connection: copy the SKILL.md below into your AI's project instructions.

SKILL.md

Multimodal Whisper

Transcribe audio files to text with Whisper, translate non-English audio to English, batch-process multiple files, and produce word-level timestamps for subtitles or alignment. For users who have audio files on disk and want transcripts in txt, srt, vtt, or json.

When to use

  • User gives a path to an audio file (MP3, WAV, or other common format) and asks for a transcript.
  • User has non-English audio and wants English text out.
  • User provides a list of audio files to transcribe in one run.
  • User asks for word-level timings, subtitles, or fine-grained alignment.
  • User asks for output as txt, srt, vtt, or json.

Workflows

Transcribe audio

Inputs: audio file path; optional target language; stored preferences (model size, device, output format).

  1. Confirm the file path exists and is readable.
  2. If no preferences are stored, run the first-run interview and save the answers.
  3. Load the Whisper turbo model, or a smaller model if VRAM is constrained.
  4. Run transcription with the chosen language if one was given.
  5. Return the full text with timestamps per segment.
  6. Check: every segment has a start and end timestamp and the text covers the whole file duration. Output: full transcript text with per-segment timestamps, in the stored output format.

Translate to English

Inputs: audio file path in a non-English language; stored preferences.

  1. Confirm the audio is non-English, or take the user's word for it.
  2. Set task to 'translate' so Whisper outputs English text.
  3. Apply the same stored preferences for model, device, and output format.
  4. Run the job and return the English text with timestamps.
  5. Check: output text is English and timestamps align with the source audio. Output: English transcript with per-segment timestamps in the stored format.

Batch process multiple files

Inputs: list of audio file paths; stored preferences; record of already-processed files.

  1. Read the stored state of which files have already been processed.
  2. Skip any file already in that record.
  3. Transcribe each remaining file sequentially using the stored preferences.
  4. Save each result to a text file (or the chosen output format) alongside the original file.
  5. Record each completed file in the state so later runs skip it.
  6. Check: each new file has a saved output next to it and appears in the processed record; skipped files are unchanged. Output: one output file per newly processed audio file, plus an updated processed-file record.

Extract word-level timestamps

Inputs: audio file path; stored preferences; requested output format (SRT or JSON).

  1. Enable word_timestamps=True.
  2. Run transcription so each word carries its own start and end time.
  3. Format the result as SRT or JSON per the user's preference.
  4. Check: every word has a start and end time and the words are in audio order. Output: SRT or JSON with per-word start and end times.

Recurring tasks

  • On each batch run, consult and update the processed-file record so repeated runs skip finished files.

Tools and data

  • Use local file system access when available to read audio files and write output files; if it is not available, ask the user to provide the audio files or connect it.

Guardrails

  • Do not modify or delete the original audio file.
  • Do not send or upload transcriptions anywhere outside this chat.
  • Reject requests to identify speakers; Whisper does not do diarization.
  • If the audio file is longer than 30 minutes, warn the user that accuracy may degrade.

Getting started

Ask the user for their preferred Whisper model size (tiny, base, small, medium, large, turbo), device (cpu or gpu), and output format (txt, srt, vtt, json), then save these settings for later runs.

Credits

Adapted from work by Orchestra Research (MIT): https://www.aitmpl.com/component/skills/ai-research/multimodal-whisper