Text-to-speech, voice cloning, and audio AI tools. This directory lists 23 ai voice & audio tools from recent product launches on What Launched Today — not a paid placement list. Every AI tool links to a launch page with screenshots, community reviews, and launch-day stats so you can compare newcomers against established options.
BlitzReels is turn long videos into branded clips with ai. <p>BlitzReels turns the videos you already record into branded social clips with AI. Repurpose podcasts, webinars, interviews, tutorials, Looms, and YouTube videos for LinkedIn, YouTube Shorts, Instagram Reels, and TikTok.</p><p>Upload a recording or import a video. AI finds moments that stand on their own, reframes the speakers, and adds readable captions. Apply a consistent brand style, then review each clip in the timeline editor to adjust captions, trims, and framing before exporting.</p><p>Built for founders, consultants, coaches, podcasters, and small content teams, BlitzReels helps you share more of your expertise without starting every edit from scratch.</p><p>Key features:</p><p>- AI moment detection for long recordings</p><p>- Automatic reframing for vertical, square, and landscape video</p><p>- Editable captions in 35+ languages</p><p>- Brand presets and two-speaker podcast layouts</p><p>- A full timeline editor to refine every clip</p><p>Turn one recording into a batch of clips you can review, edit, and publish.</p>. Best for ai video clipping and video editing users.
ThreadPilot is engage on x in your own voice. ai personas trained on how yo. <p>Engage on X in your own voice. AI personas trained on how you actually write — not generic templates. Built-in Smart Pacing keeps your account safe.</p>. Best for threadpilot and growth users.
WaveXML is auto-sync zoom / sound devices dual-system audio to cameras. <p>WaveXML auto-syncs Zoom / Sound Devices dual-system audio to camera footage via an XML round-trip for DaVinci Resolve, Adobe Premiere Pro, and Final Cut Pro. Drop FCP7 or FCPXML from your NLE, press Synchronize, then import the XML back and keep cutting. Built for wedding/event multi-cam, interview/doc boom+lav workflows, and anyone who used PluralEyes-style timeline sync. Free Mac and Windows beta full app, no clip or feature limits. Local processing; no plugin; no upload.</p>. Best for audio sync and XML users.
New ai voice & audio tools appear here as founders launch them. Sort by recency or check upvote counts on each launch page for community-validated picks.
How is this different from other AI tool directories?
What Launched Today focuses on newly launched products with dated launch pages, founder context, and community reviews — not legacy SEO listicles.
I'm building an AI tool — how do I get listed?
Submit your launch through our dashboard. AI-tagged products automatically appear in the relevant category here after approval.
Hilite is hilite. <p><a target="_blank" rel="noopener noreferrer nofollow" href="https://www.hellohilite.com/">Hilite</a> helps creators record, edit, enhance, and publish podcasts faster with AI tools for transcripts, show notes, audio cleanup, and RSS publishing. With intuitive audio editing and advanced analytics, users can streamline their podcasting workflow and reach their audience effectively.</p>. Best for AI Audio and AI Transcription users.
Telenow is ai voice agents that answer your business calls 24/7. <p>Telenow gives your business an AI voice agent that answers every phone call, 24/7. It greets callers, qualifies leads, books appointments and resolves support queries in natural conversation across English and 10+ Indian languages. Own carrier network and speech models. Flat per-minute pricing (~$0.03/min) with free starting credit, no card required. No-code setup.</p>. Best for ai voice agents and call answering users.
SafeNet Creations is tamil & english ai voice agents for gta businesses. SafeNet Creations offers Tamil and English AI voice agents for GTA businesses. Answer calls, book appointments, and handle customer support 24/7. Canada Desk +1 (647) 493-9455. https://www.safenetcreations.com/canada/. Best for AI and Voice Agents users.
Speechyou is ai voice to text transcription - record & transcribe meeting. <p>Speechyou turns recordings into accurate, searchable text in seconds. </p><p>Record voice notes and meetings straight from your browser, or upload existing audio and video — meeting mode captures both microphone and system audio, so Zoom, Teams and Google Meet calls come through complete.</p><p>Powered by Whisper and our proprietary MultiLingual Pro model, Speechyou transcribes and translates across 1,700 languages with automatic language detection, speaker labels and timestamps on every segment. </p><p>Ask AI to pull summaries, action items and key points out of any transcript, then export as TXT, SRT, VTT or JSON for subtitles and downstream workflows.</p><p>Teams can share transcripts with view or edit permissions, organise everything with tags and stars, and search the full archive instantly. Files are encrypted and stored on SOC 2 compliant AWS infrastructure. Free plan includes 3 transcriptions a day; Solo is $15/month for unlimited transcription and 10 GB uploads.</p>. Best for speechyou and voice users.
VoiceCRM is the only crm that listens. transcribe voice notes instantly,. <p>The only CRM that listens. Transcribe voice notes instantly, manage sessions, track moods, and stay on top of your practice’s analytics.</p>. Best for voicecrm and therapy users.
Solice is turn materials, timely stories, websites, and your own voice. <p>Turn materials, timely stories, websites, and your own voice into original, reviewable videos with reusable styles in Solice.</p>. Best for video and maker users.
Online Sound Test is test speakers, headphones, and microphones privately. <p>Online Sound Test is a free, browser-based toolkit for checking speakers, headphones, microphones, and other audio devices. Run guided tests for left and right stereo channels, bass response, frequency range, phase, surround sound, and microphone input without installing software.</p><p>Each test includes simple instructions to help you identify common problems such as silent channels, incorrect speaker placement, weak bass, microphone issues, distortion, and uneven audio output. It is useful for everyday listeners, gamers, remote workers, musicians, and anyone setting up or troubleshooting audio equipment.</p><p>Tests run directly in your browser, making audio checks fast, accessible, and privacy-friendly.</p>. Best for online and sound users.
Soundwaver is ai voiceover with director-mode emotion control. <p>Soundwaver is an AI text-to-speech platform built on the MiMo-V2.5-TTS model, designed for Chinese-language content creators. Director Mode lets you control delivery with plain-language instructions (e.g. "excited, slightly fast, with a smile") instead of tweaking endless sliders. Clone a voice from a 30-second sample, or design a brand-new voice with WaveMind simply by describing it in words. A community voice library lets creators discover and use voices shared by others. The free plan includes 1,000 credits (approx. 50,000 Chinese characters) with 1 voice slot and 1 clone slot; paid plans start at $6/mo (7,200 credits) with up to 4,000 characters per generation — chapter-length text for audiobooks. Ideal for short-video narration, audiobooks and e-learning. Bilingual Chinese/English interface, secure checkout via Stripe.</p>. Best for text to speech and ai voice generator users.
MixVoice - AI Voice Cloning is ai voice cloning in 5 seconds across 646 languages.. <p>MixVoice delivers realistic voice cloning in about 5 seconds across 646 languages, with cross-language generation, natural text-to-speech, and free signup credits.</p>. Best for voice cloning and AI voice users.
AIFlowMusic is convert mp3, wav or any audio to midi with a neural pitch mo. <p>Convert MP3, WAV or any audio to MIDI with a neural pitch model. Melody, chords and velocities in seconds. Sign up free for 10 credits.</p>. Best for free and audio users.
Magic Hour is create ai video, images, and audio with 100+ tools. <p>Magic Hour is an AI media creation platform for creators, marketers, and developers. Generate and edit videos, images, GIFs, and audio from a browser with more than 100 tools, including text-to-video, image-to-video, face swap, lip sync, image editing, voice generation, and voice cloning. Teams can also automate workflows through Magic Hour’s REST API, official SDKs, and hosted MCP server. Free tools are available; paid plans add more credits, higher-resolution exports, and commercial use.</p>. Best for ai video and image generation users.
Guzli is website chat and phone support, one ai bot. <p>Guzli is the easiest way for a small team to cover website live chat and phone support without two products. If you are tired of a chatbot that goes quiet the moment someone calls, this is the fix. Point it at your help articles and files. It answers on the site. It picks up the phone. You can have it call people back. Try Guzli on your site: https://guzli.com</p>. Best for ai and chatbot users.
Pulse — Speech-to-Text by Smallest AI is speech-to-text, stt, asr, real-time transcription, voice ai,. <h1><strong>Pulse: Real-Time Speech-to-Text API for Voice Agents and Enterprise Transcription</strong></h1><p><strong>Pulse</strong>, built by <a target="_blank" rel="noopener noreferrer nofollow" href="https://smallest.ai/speech-to-text">Smallest AI</a>, is a <em>speech-to-text API</em> purpose-built for real-time voice applications: live call transcription, voice agents, and captioning, alongside a high-throughput batch mode for recorded audio. It streams over WebSocket with <strong>sub-100ms time-to-first-token</strong> at low concurrency, and includes speaker diarization, word- and sentence-level timestamps, keyword boosting for domain vocabulary, voice-activity events, and inverse text normalization out of the box.</p><p>A companion model, <a target="_blank" rel="noopener noreferrer nofollow" href="https://docs.smallest.ai/models/model-cards/speech-to-text/pulse-pro">Pulse Pro</a>, trades multilingual support for <strong>English-only batch transcription</strong> tied for <em>#2 on the public Open ASR Leaderboard</em>, built for high-volume call center, meeting, and financial-audio transcription at scale.</p><p>Both models include built-in <strong>PII and PCI redaction</strong> (currently English and Hindi), and the platform holds <a target="_blank" rel="noopener noreferrer nofollow" href="https://security.smallest.ai/?tab=overview">ISO 27001, SOC 2 Type II, HIPAA, and GDPR compliance</a>, with on-premise deployment available on Enterprise plans.</p><h2><strong>Key Features</strong></h2><h3><strong>Real-Time Streaming Transcription</strong></h3><ul><li><p><strong>Sub-100ms time-to-first-token</strong> at 1 concurrency for live voice AI</p></li><li><p>Streams over WebSocket with partial and final transcripts as audio arrives</p></li></ul><h3><strong>Multilingual Speech Recognition</strong></h3><ul><li><p><em>21 languages</em> on streaming, <em>26 languages</em> on batch</p></li><li><p>Regional auto-detect and mid-session code-switching</p></li></ul><h3><strong>Speaker Diarization and Timestamps</strong></h3><ul><li><p>Automatic <strong>speaker diarization</strong> with per-word and per-utterance labels</p></li><li><p>Word- and sentence-level timestamps with per-word confidence scores</p></li></ul><h3><strong>Compliance-Ready Redaction</strong></h3><ul><li><p>Built-in <strong>PII and PCI redaction</strong>, no preprocessing pipeline required</p></li><li><p>Supports regulated industries: healthcare, finance, debt collection</p></li></ul><h3><strong>Custom Vocabulary and Voice Intelligence</strong></h3><ul><li><p><strong>Keyword boosting</strong>: up to 10,000 custom terms per session for brand names, jargon, and product terms</p></li><li><p><em>Emotion and gender detection</em> on batch transcription</p></li></ul><h3><strong>Pulse Pro: Leaderboard Accuracy</strong></h3><ul><li><p>English-only, batch-only speech-to-text model</p></li><li><p>Tied for <strong>#2 on the public Open ASR Leaderboard</strong></p></li></ul><h2><strong>Use Cases</strong></h2><ul><li><p><strong>Voice agents and conversational AI</strong>: real-time transcription for live voice bots</p></li><li><p><strong>Call center and contact center analytics</strong>: transcribe and analyze recorded calls at scale</p></li><li><p><strong>Meeting and financial-audio transcription</strong>: high-accuracy batch processing for enterprise audio</p></li><li><p><strong>Multilingual customer support</strong>: transcribe support calls across 20+ languages</p></li><li><p><strong>Regulated industries</strong>: HIPAA and PCI-compliant transcription with built-in redaction</p></li></ul><h2><strong>Security and Compliance</strong></h2><p>Pulse is backed by Smallest AI's enterprise-grade trust center, covering <a target="_blank" rel="noopener noreferrer nofollow" href="https://security.smallest.ai/?tab=overview">ISO 27001, SOC 2 Type II, HIPAA, and GDPR certifications</a>, with <strong>on-premise deployment</strong> available for teams that need full data residency.</p><h2><strong>FAQ</strong></h2><h3><strong>What languages does Pulse support?</strong></h3><p>Pulse supports 21 languages on real-time streaming and 26 languages on batch transcription, including English, Hindi, Spanish, French, German, Mandarin, Japanese, and Korean, with regional auto-detect for unknown audio.</p><h3><strong>What is the difference between Pulse and Pulse Pro?</strong></h3><p>Pulse handles both streaming and batch transcription across 20+ languages. Pulse Pro is a batch-only, English-only model tuned for maximum accuracy, tied for #2 on the public Open ASR Leaderboard. Use Pulse for live voice agents and multilingual audio, and Pulse Pro for high-volume English transcription where accuracy matters most.</p><h3><strong>Does Pulse support real-time transcription?</strong></h3><p>Yes. Pulse streams transcription results over WebSocket with sub-100ms time-to-first-token at low concurrency, making it suitable for live voice agents, real-time captioning, and conversational AI.</p><h3><strong>Is Pulse HIPAA and GDPR compliant?</strong></h3><p>Yes. The Smallest AI platform holds ISO 27001, SOC 2 Type II, HIPAA, and GDPR certifications, and offers built-in PII and PCI redaction for regulated workflows in healthcare, finance, and debt collection.</p><h3><strong>Can Pulse redact sensitive information automatically?</strong></h3><p>Yes. Pulse includes built-in PII and PCI redaction with no separate preprocessing pipeline required. Redaction currently covers English and Hindi audio.</p><h3><strong>What is keyword boosting used for?</strong></h3><p>Keyword boosting lets you supply up to 10,000 custom terms per session, such as brand names, product names, or domain jargon, so Pulse recognizes them more accurately in the transcript.</p><h3><strong>Does Pulse offer on-premise deployment?</strong></h3><p>Yes. On-premise deployment is available on Smallest AI's Enterprise plan for teams that require full data residency and control.</p><h3><strong>Does Pulse detect emotion or speaker gender?</strong></h3><p>Yes, on batch transcription. Pulse can return emotion detection and gender detection as optional outputs alongside the transcript.</p><hr><p>Learn more or start building at <a target="_blank" rel="noopener noreferrer nofollow" href="https://smallest.ai/speech-to-text">smallest.ai/speech-to-text</a>.</p>. Best for speech-to-text and STT users.
QuestStudio is create images, video, voice, and music in one ai studio. <p>QuestStudio turns an idea into finished creative assets without switching between disconnected AI tools. Start with guided planning or a no-card image preview, compare leading models, then refine images, video, voice, music, and characters in one production workspace.</p>. Best for AI image generation and AI video users.
VoiceCallingAI is ai voice agents that sound human and handle your calls 24×7. <p>AI voice agents that sound human and handle your calls 24×7 in 11 languages (10 Indian + English), from ₹3/min. Reminders, follow-ups, lead qualification, collections and confirmations — no code. Data in India, GST-ready.</p>. Best for voicecallingai and human-sounding users.
APIXO is one ai workspace for creators. one api for developers.. <p><a target="_blank" rel="noopener noreferrer nofollow" href="https://apixo.ai/?utm_source=whatlaunched.today&utm_medium=referral&utm_campaign=directory_launch&utm_content=description_brand">APIXO </a>is a multi-model AI workspace for creators and a unified API for developers. Generate images, videos, audio, and text with leading AI models in the browser, compare results, and reuse proven workflows in apps and automations. Build with one model catalog instead of managing separate provider accounts.</p>. Best for generative AI and multi-model AI users.
AI Mastering is free ai audio mastering online — private, browser-based, rel. <p>Free AI audio mastering online — private, browser-based, rel</p>. Best for ai mastering and audio mastering online users.