How to narrate a video in your own voice

Voice cloning is not really about sounding convincing — it is about what happens when a sentence changes. The value is the re-record you never have to do, and the setup takes about thirty seconds of clear speech.

The reason to clone your voice is not that a synthetic read sounds impressive. It is that on the fourth revision of a script, the narration is the only part of the video that still costs you a session at a microphone.

Everything else about an edit is cheap to change. A cut is free. A title is free. A colour grade is free. Narration is the one element where changing four words means finding a quiet room, warming up, matching your energy from three days ago, and re-recording — and it is exactly the element that changes most, because copy is what people have opinions about.

A cloned voice removes that. Change the sentence, click, and the line is read again in your own voice, at the same pitch and pace as everything around it.

What you actually record

About thirty seconds of clear speech. A passage is written for you, so you are not standing there inventing something to say, which is harder than it sounds and produces a stilted read.

Three things matter and nothing else does:

  • Quiet. Not a studio — a room with the window shut and the fan off. Whatever is in the background of your recording is in the background of the voice.
  • Your normal pace. Read it the way you would narrate, not the way you would read aloud in a classroom. A careful, over-enunciated sample clones into a careful, over-enunciated narrator.
  • Length. Aim for thirty seconds; under twenty is refused, because every engine clones badly from less than that and it is better to say so up front than to hand back something that almost sounds like you.

If you already have a clean recording — an old voiceover, an interview, a podcast take — upload that instead. A take with music under it, or a voice buried in other sound, is refused when you create the voice rather than quietly cloned into nonsense later.

The recording is the ceiling. Everything downstream inherits its room tone, its sibilance and its mood, and no amount of engine choice climbs above it. Thirty focused seconds is worth more here than at any other point in a project.

Permission is not a footnote

You are asked whose voice it is: your own, or someone else's with their explicit written permission — which asks for their name. The answer is recorded with the exact wording you agreed to.

There is also a liveness check: a phrase is generated, you read it aloud, and we confirm it was spoken just now rather than downloaded from somewhere. That defeats lifting a voice off a podcast. It is not identity verification, and we do not claim it is — it cannot prove the person reading is the person the voice belongs to.

Cloning a voice without permission is unlawful in many places and a criminal offence in some. If you are cloning a colleague for company narration, get it in writing before you record, not after someone hears the result.

Draft cheap, deliver good

Three engines sit behind one voice, and they differ in more than price.

One works the instant your recording exists and stores nothing anywhere. It is the draft engine: use it while the script is still moving, when you want to hear whether a paragraph lands, not whether it sounds broadcast.

One hands us a voice file we keep, which means deleting it is entirely ours to honour — and it is.

The best-sounding one is held by the provider, where it lapses after a week without use. We keep it alive for you automatically, and your original recording can always rebuild it. Use this one for anything anyone else will hear.

Set one recording up on all three and switch per generation. The cost difference is per word, which sounds trivial until you are on the fifth pass of a nine-minute explainer.

Script first, voice second

This is the part that changes how you work, and the reason to bother at all.

Write the script as timed segments, pick your voice once, and read each segment with its own take. Every take lands on an audio lane at the segment it belongs to, already in position — you are not dragging forty files onto a timeline and nudging them into place.

Then, when paragraph six is wrong, you re-read paragraph six. Not the narration. Not the surrounding paragraphs to match the energy. One line, one click, back in position.

That is the whole argument. A single clip's voiceover is a nice convenience; a forty-segment script that stays editable to the end is a different way of making the video.

Mixing it so it does not sound generated

A cloned read arrives clean and level, which is not the same as finished.

  • Duck the music. Set the voice as your reference, bring the bed in around 12 dB below it, and flag the music clip as Duck so it dips under speech automatically and recovers in the pauses.
  • Leave the breaths in the timing. Machine narration has no natural gasp, so the pauses have to come from the edit. A beat before an important sentence is what makes it read as a person.
  • Check it on phone speakers. Music that sits politely under a voice on monitors routinely swallows it on a phone, and a phone is where this will be watched.
  • Caption against the final audio. Captions are generated from the audio as it exists — recut or re-read afterwards and they are out of step.

Where it goes wrong

Punctuation is direction. Engines read commas and full stops as timing. A sentence that runs on in the script runs on in the read. If a line lands flat, re-punctuate it before you blame the engine.

Names and jargon. Product names, place names and acronyms are where synthetic reads most often break character. Spell them phonetically in the script and nobody will know.

Reading is not performing. A cloned voice reproduces how you sounded for thirty seconds. If you recorded the sample tired, everything you narrate for the next year sounds tired. It is worth re-recording the source once you have heard yourself narrate a real script.

Deleting is real. Deleting a voice removes your recording, the voice files we hold and every engine set up from it. The record that you confirmed permission is kept, marked as withdrawn — which is the point of having it.

The order that works

  1. Record thirty seconds somewhere quiet, at narration pace.
  2. Say whose voice it is and, if it is not yours, have the permission before you record.
  3. Set up all three engines so you can switch per generation.
  4. Write the script as segments before you read a word of it.
  5. Draft on the cheap engine while the copy is still moving.
  6. Deliver on the recommended one once it is locked.
  7. Duck the music, then caption, in that order.