S2.1 Pro
Fish Audio · Available for text to speech
Create natural, expressive narration for videos, products, assistants, and stories. Choose a voice or clone one you have permission to use, then shape the pace and export production-ready audio.
Voice studio
Text to speech preview
Your script
Every idea has a voice. Give yours the clarity, warmth, and rhythm it deserves.
Speed
1.0×
Generation mode
Fast and balanced
Supported models
New speech models appear here as they become available, so you can choose the right voice technology for each project.
Fish Audio · Available for text to speech
Built for real production
Move from a finished script to usable audio without recording gear or a complicated editing workflow.
Advanced speech models preserve phrasing, pauses, and expression so long narration stays clear and engaging.
Upload a clean voice sample and its transcript to reproduce an authorized voice without a training step.
Adjust speaking speed and choose a generation mode that favors responsiveness or stability.
Export MP3, WAV, PCM, or Opus for editing, streaming, apps, and telephony workflows.
Three simple steps
Start with text, pick the right voice, and generate audio that is ready to use.
Use punctuation and paragraph breaks to guide pauses, pacing, and emphasis naturally.
Select a saved voice ID or upload an authorized 10 to 30 second reference sample with its transcript.
Choose a supported model, speed, and format, then preview or download the result.
Voice cloning
A short, clean recording can become a reusable voice for narration, product demos, lessons, and branded content. An accurate transcript helps preserve pronunciation and consistency.
Reference voice sample
Voice ready
Generate new speech in the authorized voice
Create consistent speech for the channels your audience already uses.
Produce narration for videos, explainers, podcasts, courses, and audiobooks.
Give assistants and product experiences a more natural, recognizable voice.
Turn written information into clear spoken content for more people.
Create phone prompts, alerts, updates, and reusable product messages.
FAQ
Answers about supported models, formats, voice cloning, and production use.
AI text to speech converts written text into spoken audio. HeyMarmot uses advanced speech models to create natural phrasing, pauses, and expression for narration and product experiences.
The speech models currently available on HeyMarmot are: S2.1 Pro.
Yes. Upload a clean 10 to 30 second recording and the exact transcript. You must own the voice or have explicit permission from the voice owner.
You can generate MP3, WAV, PCM, and Opus audio. MP3 is convenient for everyday publishing, WAV suits high quality editing, and PCM or Opus work well in real-time pipelines.
Use clear punctuation, paragraph breaks, and a well-written script. You can also adjust speed and use a clean reference recording with an accurate transcript when cloning a voice.
Create natural speech with the model that fits your project and keep every voiceover in your HeyMarmot workspace.
Explore more creation guides