AI app for writers · no coding needed
Source-linked desktop dictation and meeting transcription console
Reduce tool switching and keep spoken work on the user's machine while producing source-linked text.
Made for: Writers, consultants and small teams who dictate text and transcribe meetings on their own computers

What it does for you
The problem
Spoken work is scattered across separate dictation and meeting tools, and audio often leaves the machine before it becomes usable text.
What it gives you
Reviewed transcripts and dictation text linked to their audio sources
What you give it
Microphone audiomeeting recordingsuser-defined vocabularypermitted reference documents
Build your own version of Aqua Voice, Whisperstream and more
One app with what these 4 AI tools do, yours to keep and change: Aqua Voice, Whisperstream, Notto, NexTalk.
Everything these tools do, in one app
- Voice-to-text dictation Converts spoken words into text that can be pasted into any application.Found in Whisperstream, Notto, NexTalk
- Meeting transcription Captures spoken content from meetings and creates searchable notes.Found in Notto
- On-device processing Processes audio locally on the user's computer without sending data to external servers.Found in Whisperstream, NexTalk
- Push-to-talk hotkey Allows users to activate dictation with a customizable keyboard shortcut.Found in Whisperstream
- Overlay interface Displays a transparent layer over the screen for AI assistance without switching apps.Found in Notto
- Invisible chat overlay Provides an overlay for asking questions or brainstorming while keeping the current screen visible.Found in Notto
- Quick note-taking Enables capturing notes and action items directly on top of other applications.Found in Notto
- Cross-platform desktop support Works on multiple desktop operating systems such as macOS, Windows, and Linux.Found in Notto
- Low latency Provides near-instant feedback while dictating, with sub-20ms latency.Found in NexTalk
- Open-source Allows users to inspect, modify, or contribute to the software's code.Found in NexTalk
- Native Linux integration Integrates with Fcitx5 via Unix sockets for compatibility with Wayland and common Linux apps.Found in NexTalk
- Minimalist UI Offers a clean, low-distraction visual interface that sits over the desktop.Found in NexTalk
- Multiple voice styles Provides various voice styles and customizable parameters to tailor speech output.Found in Aqua Voice
- Multilingual support Supports various languages and accents for global usability.Found in Aqua Voice
- Real-time generation Enables quick iterations and immediate results for voice output.Found in Aqua Voice
- API integration Offers integration options with popular platforms and APIs for seamless workflow incorporation.Found in Aqua Voice
- Free trial Provides a limited free version to explore core capabilities without commitment.Found in Aqua Voice, Whisperstream, Notto
How it works, step by step
- Convert spoken words into text that can be pasted into any application
- Capture meeting audio and produce searchable notes
- Process audio locally on the user's computer without sending it to external servers
- Activate dictation with a customizable push-to-talk hotkey
- Show a transparent overlay for assistance without switching apps
- Provide an invisible chat overlay for questions while the current screen stays visible
- Capture notes and action items on top of other applications
- Run on macOS, Windows and Linux desktops
- Provide near-instant feedback while dictating
- Allow users to inspect, modify or contribute to the software's code
- Integrate with Fcitx5 via Unix sockets for Wayland and common Linux apps
- Offer a clean, low-distraction interface over the desktop
- Provide multiple voice styles and customizable parameters for speech output
- Support various languages and accents
- Generate results in real time for quick iterations
- Offer API integration with common platforms and workflows
- Provide a limited free version to explore core capabilities
- Compare the reviewed result with the recorded baseline and value assumptions
- Capture corrections and named-owner approval before consequential use
- Export a versioned reviewed transcript and dictation text linked to its audio sources with source references and unresolved questions
Build it yourself with your AI system
Build this app yourself, no coding needed
Start with a quick version you can try in a few minutes. Like it? Then build the full app by copying and pasting our step-by-step instructions: everything is prepared for you.
Sign in to see how to build it yourself
Build a quick version to try, or get the full app pack for Source-linked desktop dictation and meeting transcription console with the step-by-step building instructions. You don't need any technical skills: you copy, paste and answer a few questions. Both are included in the membership.
4 Have it built for you days to a few weeks
Rather not do it yourself, or want it fully tailored to your data, your way of working and your brand? Nexibeo builds Source-linked desktop dictation and meeting transcription console with you.
What's in the app pack
Included in the Complete AI Training membership.
- The building instructions your AI follows, step by step
- The questions your AI will ask you about your business before it starts
- A clickable demo you can open in your browser, to see how it should work
- A detailed blueprint of the screens, the information it keeps and the checks it runs
Become a member to get the app packAlready a member? Sign in
The files, for the technically curious
- START-HERE.mdHow to build it with your own AI (read first)3 KB
- README.mdOverview and links4 KB
- questions.mdQuestions to answer before you build2 KB
- prompt-cloudflare.mdThe full build prompt, hosted on Cloudflare25 KB
- prompt-vps.mdThe same build on your own server (Docker)25 KB
- spec.jsonData model, API, AI pipeline, acceptance criteria12 KB
- demo/index.htmlThe working demo on sample data195 KB
Questions
Do I need to know how to code?
No. You copy and paste the prompts on this page into ChatGPT or Claude, and the AI does the building. When it asks you something, you answer in your own words.
What does it cost?
The quick version, the app pack and the step-by-step instructions are for members: you pay the membership price, not a price per app (see the plans). Building the full app uses your own ChatGPT or Claude subscription. Putting it online is often cheap or no cost at the start, and your AI tells you before anything costs money.
How long does it take?
The quick version: about two minutes. The real app: an afternoon for a first version you can use, longer if you want every feature.
Can I change it to fit my business?
Yes. Tell your AI what to change in plain words, like “add a column for the price” or “use our logo and colours”. Or have Nexibeo build and customise it for you.
More detailsHow the AI works, safeguards and what to build first
Reduce tool switching and keep spoken work on the user's machine while producing source-linked text. For writers, consultants and small teams who dictate text and transcribe meetings on their own computers, convert microphone audio, meeting recordings and user-defined vocabulary into reviewed transcripts and dictation text linked to their audio sources. The benefit is a testable hypothesis, measured through accepted transcript words per editing hour and corrections after insertion; do not assume that AI output alone produces business value.
Confirm the buyer's problem and scope, collect microphone audio, meeting recordings, user-defined vocabulary and permitted reference documents, then follow this sequence: 1. Convert spoken words into text that can be pasted into any application. 2. Capture meeting audio and produce searchable notes. 3. Process audio locally on the user's computer without sending it to external servers. Resolve uncertain cases with qualified reviewers, approve reviewed transcripts and dictation text linked to their audio sources, and measure accepted transcript words per editing hour and corrections after insertion against a documented baseline.
How the AI works
Use AI to interpret permitted inputs, suggest structured mappings and generate candidate outputs for the three stated task modules. Use deterministic code for arithmetic, schema validation, hard constraints and reproducible tests. Review source-linked explanations and uncertainty before accepting results. Local processing on the user's machine; final accuracy and meaning checks remain human. A model suggestion is never a verified fact, professional decision or authorization to act.
Safeguards
Preserve speaker meaning, source attribution, quotation accuracy and usage permissions. Users approve substantive changes and publication scope. One desktop operating system and one language pair; final accuracy and meaning checks remain human. Keep all consequential actions under authorized human control and do not fabricate missing inputs, permissions, professional judgments or market evidence.
What to build first
Pilot scope: One desktop operating system and one language pair; final accuracy and meaning checks remain human. Implement one approved input format, a bounded representative case set and the first two task modules: convert spoken words into text that can be pasted into any application; capture meeting audio and produce searchable notes. Support the third module with operator review: process audio locally on the user's computer without sending it to external servers. Include source references, corrections, basic organization access, approval states, export and value measurement. Use managed operator assistance for unresolved exceptions. The cost estimate covers this narrow prototype, not unrestricted multi-tenant scale, complex production integrations, specialist certification or physical operations.
What it can connect to
User-owned audio files, permitted meeting recordings and authorized reference documents. Cloud storage, document editors and note destinations. Start with file exchange and validate destination specifications before promising direct publishing. Start with authorized file exchange. Validate current provider access, usage rights and schema behavior before promising a connector.
The screens in detail
Primary screens: Session setup and vocabulary, Live dictation and transcript review, Meeting notes and export. Use a session list for recordings and dictations, a large central transcript canvas with audio timestamps, and a right-hand panel for vocabulary, speakers and comments. Let users compare transcript versions side by side. Display draft, changes requested and approved states. Provide a client preview link with comments anchored to the relevant transcript segment. Make the task-specific outcome reviewed transcripts and dictation text linked to their audio sources visible beside its evidence, review state and value baseline.





