Text-to-speech, voice cloning, and audio AI tools. This directory lists 31 ai voice & audio tools from recent product launches on What Launched Today — not a paid placement list. Every AI tool links to a launch page with screenshots, community reviews, and launch-day stats so you can compare newcomers against established options.
Soundwaver is ai voiceover with director-mode emotion control. <p>Soundwaver is an AI text-to-speech platform built on the MiMo-V2.5-TTS model, designed for Chinese-language content creators. Director Mode lets you control delivery with plain-language instructions (e.g. "excited, slightly fast, with a smile") instead of tweaking endless sliders. Clone a voice from a 30-second sample, or design a brand-new voice with WaveMind simply by describing it in words. A community voice library lets creators discover and use voices shared by others. The free plan includes 1,000 credits (approx. 50,000 Chinese characters) with 1 voice slot and 1 clone slot; paid plans start at $6/mo (7,200 credits) with up to 4,000 characters per generation — chapter-length text for audiobooks. Ideal for short-video narration, audiobooks and e-learning. Bilingual Chinese/English interface, secure checkout via Stripe.</p>. Best for text to speech and ai voice generator users.
MixVoice - AI Voice Cloning is ai voice cloning in 5 seconds across 646 languages.. <p>MixVoice delivers realistic voice cloning in about 5 seconds across 646 languages, with cross-language generation, natural text-to-speech, and free signup credits.</p>. Best for voice cloning and AI voice users.
AIFlowMusic is convert mp3, wav or any audio to midi with a neural pitch mo. <p>Convert MP3, WAV or any audio to MIDI with a neural pitch model. Melody, chords and velocities in seconds. Sign up free for 10 credits.</p>. Best for free and audio users.
Guzli is website chat and phone support, one ai bot. <p>Guzli is the easiest way for a small team to cover website live chat and phone support without two products. If you are tired of a chatbot that goes quiet the moment someone calls, this is the fix. Point it at your help articles and files. It answers on the site. It picks up the phone. You can have it call people back. Try Guzli on your site: https://guzli.com</p>. Best for ai and chatbot users.
New ai voice & audio tools appear here as founders launch them. Sort by recency or check upvote counts on each launch page for community-validated picks.
How is this different from other AI tool directories?
What Launched Today focuses on newly launched products with dated launch pages, founder context, and community reviews — not legacy SEO listicles.
I'm building an AI tool — how do I get listed?
Submit your launch through our dashboard. AI-tagged products automatically appear in the relevant category here after approval.
QuestStudio is create images, video, voice, and music in one ai studio. <p>QuestStudio turns an idea into finished creative assets without switching between disconnected AI tools. Start with guided planning or a no-card image preview, compare leading models, then refine images, video, voice, music, and characters in one production workspace.</p>. Best for AI image generation and AI video users.
VoiceCallingAI is ai voice agents that sound human and handle your calls 24×7. <p>AI voice agents that sound human and handle your calls 24×7 in 11 languages (10 Indian + English), from ₹3/min. Reminders, follow-ups, lead qualification, collections and confirmations — no code. Data in India, GST-ready.</p>. Best for voicecallingai and human-sounding users.
AI Mastering is free ai audio mastering online — private, browser-based, rel. <p>Free AI audio mastering online — private, browser-based, rel</p>. Best for ai mastering and audio mastering online users.
AudioToText.run is faithful multilingual transcription for audio and video.. <p><a target="_blank" rel="noopener noreferrer nofollow" href="http://AudioToText.run">AudioToText.run</a> is a multilingual transcription workspace for long audio and video. It supports MP3, WAV, M4A, MP4 and MOV, and exports editable TXT, SRT and VTT files. It helps podcasters, researchers, educators, journalists and teams turn recordings into accurate, searchable text while preserving names, numbers, speakers and context.</p>. Best for transcription and speech to text users.
Qwen Audio 3.0 TTS is create multilingual text to speech with qwen audio 3.0 tts.. <p>Create multilingual text to speech with <a target="_blank" rel="noopener noreferrer nofollow" href="https://qwenaudio3.com/">Qwen Audio 3.0 TTS</a>. Control emotion and pacing, compare Plus and Flash, hear samples, and generate professional audio.</p>. Best for qwen and audio users.
PodTyper is podcast transcription from podcast urls. <p>Podtyper transcribes YouTube, Spotify, and Apple Podcasts, with AI summary, key takeaways, and exportable files. First 30 minutes free.</p>. Best for podcast and transcription users.
Qwen3 TTS Voice Studio is controllable multilingual text-to-speech workspace. <p>Qwen3 TTS is presented as a browser-based text-to-speech workspace for creators and teams preparing multilingual voice drafts. According to the product page, it supports a practical workflow for turning written scripts into reviewable speech, comparing delivery choices, and checking pronunciation, pacing, and tone before final production decisions.</p>. Best for qwen3 tts and text to speech users.
KidVoice is ai kid voices with natural child speech. <p>KidVoice is an AI-powered voice generation platform designed for creating natural child, teen, and character voices.</p><p>It offers more than 50 voice styles, multilingual text-to-speech, permitted voice cloning, instant browser previews, and downloadable audio—all in one focused workspace.</p><p>KidVoice is ideal for storytelling, animation, games, educational content, short videos, product demos, and creative prototypes. Simply enter your script, choose a voice, generate the audio, preview the result, and download your preferred take.</p><p>Voice cloning is only supported with proper consent and permission.</p>. Best for ai voice generator and text to speech users.
VoiceDrop is ai-powered ringless voicemail and voice cloning for smarter. <p>VoiceDrop AI is an AI-powered ringless voicemail platform that helps businesses reach prospects without interrupting them with traditional phone calls. Users can create personalized voicemail campaigns using AI voice cloning, dynamic variables, and automated workflows, then deliver messages directly to recipients' voicemail inboxes without making their phones ring.</p><p>The platform supports bulk voicemail campaigns, AI-generated scripts, CRM integrations, phone number validation, detailed analytics, multilingual voice generation, and API access for custom automation. Businesses can also use VoiceDrop's AI Callback Agent (Beta) to automatically answer returned calls, qualify leads, schedule appointments, and improve customer engagement.</p><p>VoiceDrop AI is designed for sales teams, marketing agencies, real estate professionals, insurance brokers, recruiters, healthcare providers, and other organizations that rely on outbound communication. With built-in compliance tools, scalable infrastructure, and integrations with platforms like HubSpot, Salesforce, GoHighLevel, Zapier, and Twilio, VoiceDrop helps businesses automate personalized outreach while increasing response rates and saving time.</p>. Best for AI and Sales users.
Sing Test is a free online vocal assessment tool providing instant scores. <p>Sing Test is a private, browser-based application that allows users to evaluate their singing abilities in just two minutes. By analyzing microphone input in real-time, the tool provides detailed feedback on pitch accuracy, identifies vocal range, and determines voice type without requiring any signups or data uploads.</p>. Best for singing and vocal training users.
Riffloop is isolate any instrument, loop the hard parts, slow it down and change key on youtube. <p>Riffloop is an all-in-one music practice studio that runs right under the YouTube video you are already watching.<br><br>AI stem separation splits any song into six parts: vocals, drums, bass, guitar, piano and more. Solo the guitar to learn a riff note for note, or mute the vocal and sing it yourself. Drag a region over the bar that trips you up and loop it, snapped to the beat, until it sticks. Pull the tempo down to 25% to catch every note with the pitch held steady, then walk it back up to full speed. Shift the whole track up or down by as much as 12 semitones to fit your voice or instrument.<br><br>It works on your own files too. Upload an MP3, WAV or FLAC to the web Studio and get the same stems, loops, speed and key controls.<br><br>Most tools give you half of this. A stem splitter hands you files and walks away. A looper forgets the stems. Riffloop keeps all of it in one place, on the song you are already learning.<br><br>Free to start, no signup needed to try it.</p>. Best for stem splitter and vocal remover users.
EzTranscript is free ai transcription for instagram reels and tiktok videos. <p>EzTranscript is a free AI-powered transcription tool built for short-form video. Paste any Instagram Reel or TikTok link and get an accurate, timestamped transcript in seconds. No signup required for basic use.</p><p></p><p>Key features include AI-generated summaries (TL;DR), translation into 12 languages, and downloadable SRT subtitle files. The tool handles videos up to 60 minutes and works entirely in the browser.</p><p></p><p>Built for content creators who need to repurpose video into text — blog posts, captions, subtitles, or social media threads. Also useful for students, journalists, and social media managers who work with short-form video daily.</p><p></p><p>Freemium model: free daily transcripts for all users, with a paid tier for higher volume and priority processing.</p>. Best for ai and transcription users.
MP3 to MIDI is mp3 to midi converter — free, no signup. <p><a target="_blank" rel="noopener noreferrer nofollow" href="https://mp3tomidi.io/"><strong>MP3 to MIDI Converter</strong></a> is a free online tool that uses AI to transcribe MP3 audio into editable MIDI files—all directly in your browser.</p><p><strong>How it works</strong></p><p>Upload an MP3 file, choose the closest audio type (voice, piano, guitar, bass, another solo instrument, a full song, or Not Sure), and the AI transcribes the notes. The entire process runs locally in your browser—your audio stays on your device and is never uploaded to a server.</p><p><strong>Key Features</strong></p><ul><li><p><strong>AI-powered note detection</strong> — Detects pitch, timing, duration, and velocity, including chords and polyphonic audio</p></li><li><p><strong>100% browser-based</strong> — No uploads, no servers, no signup required</p></li><li><p><strong>Built-in piano roll editor</strong> — Move, resize, join, or delete notes with undo, redo, and optional grid snapping</p></li><li><p><strong>Preview before downloading</strong> — Compare the original audio with the synthesized MIDI playback</p></li><li><p><strong>Standard MIDI output</strong> — Download a .MID file and continue editing in any DAW or MIDI editor</p></li><li><p><strong>Works on desktop and mobile</strong> — Use it in any modern browser</p></li></ul><p><strong>Limitations</strong></p><ul><li><p>Supports MP3 files up to 20 MB and 3 minutes</p></li><li><p>Works best on solo instruments or clear melodies; full songs are experimental and may need cleanup</p></li><li><p>Does not export PDF sheet music, MusicXML, tabs, or separate instrument tracks</p></li></ul><p><strong>Who it's for</strong></p><ul><li><p>Musicians and producers who want to transcribe audio into their DAW (Ableton Live, Logic Pro, FL Studio, etc.)</p></li><li><p>Music learners who want to see the notes of a song</p></li><li><p>Anyone who needs a quick, private, and free MP3 to MIDI conversion</p></li></ul><p></p>. Best for mp3 to midi and AI users.
Amical is type 4x faster with your voice. <p>AI-powered speech-to-text that auto-formats your dictation for any app — emails, messages, code and more. Open source, private and free.</p>. Best for ai and saas users.
VideoAny Brazil is vídeos, imagens e áudio com ia em uma só plataforma. <p>VideoAny is an AI creative platform available in Brazilian Portuguese for producing videos, images, and audio from text prompts or visual references. Its tools include text-to-video, image-to-video, video-to-video, face swap, lip sync, image generation and editing, and audio creation in one online studio. New accounts can try the service with promotional credits.</p>. Best for AI and video generation users.
Hallodesk is the 24/7 ai voice receptionist. <p>Hallodesk's AI Phone Voice Agent answers your business calls 24/7 — booking appointments, routing emergencies, and handling FAQs automatically. No more missed leads.<br>Every missed call is a missed customer. For small and medium-sized businesses — dental clinics, law firms, real estate agencies, hair salons — the phone is still the primary way customers make contact. Yet the average SME misses 60% of inbound calls outside business hours. That's not just lost revenue. That's a competitor's gain.<br>Hallodesk solves this with an AI Phone Voice Agent that works 24 hours a day, 7 days a week, 365 days a year — without breaks, sick days, or holidays.<br>Explore our [AI Phone Voice Agent](https://hallodesk.de/funktionen), [AI Website Voice Widget](https://hallodesk.de/en/website-voice-agent), and [WhatsApp Chat Automation](https://hallodesk.de/whatsapp-ki-agent) to transform your customer communication.</p>. Best for AI and Voice AI users.