How to create text to speech
Text to Speech is for turning written scripts into usable voice audio. By the end of this guide, you will have a generated audio card in Generation History that you can review, download, or reuse in your production workflow.
Before you start
- Check the account menu and confirm whether Personal or a Workspace is selected. That selection determines which credits are used and where the generation is saved.
- Have a finished script or one short section of a longer script ready.
- Save or choose the voice you want to test before generating a long read.
- Decide how the audio will fit into your production workflow.
Create it in Giggy
- In the left sidebar, open Text to Speech.
- Enter the script in the text area.
- Choose the voice shown in the voice card or open the voice selector.
- Adjust Speed only if the delivery is too slow or too fast.
- Before Generate, review the mode selector. Creative opens with Fast selected; Fast consumes credits, while Batch costs 0 credits.
- Switch to Batch before generating when you want zero-cost generation. The Generate button shows the applicable credit cost. Batch text and final audio are removed within seven days of request creation, or sooner under the 12-hour privacy option; download audio before it expires.
- Select Generate.
- Review the finished audio card in Generation History.
Improve the result
- Write scripts in complete sentences so the AI voice has clear phrasing.
- Use punctuation to guide pauses, emphasis, and sentence rhythm.
- Break long scripts into shorter sections when you need more control.
- Spell out numbers, symbols, prices, acronyms, names, and technical terms when pronunciation matters.
- Add short pronunciation hints directly in the script for words that are easy to misread.
- Preview the selected voice before generating long audio.
- These are examples, not the complete set. Add a tag before the sentence that needs the expression. Use the in-product selector to see the controls currently available for your workflow, and keep the script readable by using tags only when the cue changes the delivery.
- Adjust wording first when timing feels wrong; use Speed only after the script reads correctly.
Examples
- Good script: "Welcome to the product demo. In the next 30 seconds, we will show how the dashboard turns raw feedback into clear priorities."
- Pronunciation helper: write "twenty-nine dollars" instead of "$29" when the exact read matters.
- Short-form voiceover: keep the message direct and easy to review before exporting.
Common mistakes
- Generating a long script before testing the selected voice on one short paragraph.
- Using Speed to fix awkward timing when the script itself needs cleaner punctuation.
- Leaving acronyms, names, prices, or technical terms in a format the voice may misread.
Use cases
Creator voiceover
- Create narration for videos, ads, explainers, tutorials, social clips, podcasts, and product demos.
- Generate one clean section at a time for long-form scripts that need repeatable pacing.
Drafting and iteration
- Generate draft narration before recording final human audio.
- Test different voices against the same script before choosing one for production.
Training and product content
- Make training reads, onboarding modules, course snippets, product walkthroughs, and announcement audio.
- Use complete sentences and punctuation so instructional content has clear phrasing.
Short-form production audio
- Create concise voiceover audio for creator projects.
- Keep short scripts focused so each generated take is easy to review.
Expression tags
[laughter], [sigh], [dissatisfaction]