Week starting 17th August: Good availability with most voices this week Studios reopen 9am Thursday 20th August Need help choosing a voice? Send us your brief
Get a quote

Notable AI Tech (2 of 3)

Updated

Intro

We think it's important to highlight some of the AI tech that directly affects our industry. This page focuses on Speech-to-Speech systems, where a different vocal likeness is mapped onto a 'donor' recording.

In the following example, Gary recorded the donor recording first, and we've edited to show a short clip line by line this get mapped using a female voice and male voiceover artist likeness. Neither the AI female or male voice are a likeness of any real person - the AI tech provider randomly generated a likeness for this process.

2 of 3 - Speech-To-Speech Systems

These systems are often developed by makers of text-to-speech models.

There are different types:
a) Some are marketed as 'speech to speech' but are actually:
1) 'clone generated from speech' -> 2) 'speech-to-text (transcription)' -> 3) 'translation into another language' -> 4) 'translated text-to-speech', the latter using the clone created earlier.

They are for automated translation and dubbing, such as English to German, whilst retaining the 'tonality' of the speaker. The results are currently less than perfect but are seen as a useful tool by some content creators. I'd say the tech doesn't really encroach much on services our voices would be offering.
Technically, you could voice a script and offer, say, a German version of the recording. But unless you are a German speaker, you may not be able to perform your own quality control. Plus, you would be effectively advocating the use of poor-quality text-to-speech audio, which is probably not your aim.

b) Then there is pure 'speech-to-speech' tech which maps a likeness onto a voice recording, largely retaining the language, accent, pace, and stress of the 'donor track'.

These systems have been around for a while but have become significantly more realistic in Q4 2023, especially when mapping English to English.
Some providers market these products as 'pro voiceover'.
For example, a versatile voiceover artist may work on a computer game project, perform lots of donor recordings in differing accents and styles, which are then remapped using the likenesses of different characters. When remapping, you can change all aspects of the timbre, thus changing gender and age. There is a grey area here in whether this is good for the industry, though many users of the tech would argue that budgets don't allow for employing hundreds, if not thousands, of artists, and they would use the tech for ancillary characters, not the leads. They argue it's just like having a good character voice artist with a superpower to manipulate their voice.

As well as remapping the recording using different likenesses, you can map your recording using your own cloned likeness, which gives a varied take, perhaps a brighter read.

Pros for voices

Some voices may consider using the tech in the following way:
- Opens up possibilities for continuing to voice when your voice is croaky, for example, by mapping onto your own clone.
- Create a clone using recordings from when you were younger and effectively produce a youthful take.
- Clone your children and voice scripts on their behalf.
- Have a clone-buddy, where you voice scripts for your voiceover buddy when they are unavailable.
- Could a person leave a clone to a family member in their will? Realistically, it might be too difficult to hear a representation of a loved one who is no longer with us.

Some of these we've mentioned in another article, and may need more discussion to see if they are acceptable uses in the industry.
They are not necessary our views on what is acceptable.

If your main strength is your ability to perform a script more than your timbre, then you have an advantage over some other voice talent.

Cons for pro voices

- The technology ALSO allows remapping a recording using a likeness with no human owner. This could mean voices compete in a way not seen before.
- Some clients or producers may feel confident making the donor recording and then remapping it using a timbre of their choosing.

If your main strength is your timbre, rather than your ability to perform a script, then you are at a disadvantage compared to other voice talent.

Limitations

Audio quality is not amazing and can heavily depend on the clone used, and the quality of the donor voice performance.
Not until someone runs the donor recording through the Speech to Speech process will you know if it's worked. The donor voice recording may need to be re-recorded in places through no fault of the reader.

Ways for us to attack the tech and differentiate yourself

- Privacy grounds. Make more of your privacy policy when handling client data.
However, this can also backfire as many of the tools we all use now are slightly grey in that regard. It may stop us from using Adobe tools, Zoom for link-up sessions, etc.

- Waterproofing recordings (there will be a separate section on this).
- Highlight that it doesn't sound as good as the real thing.

Summary

In summary, we see some pros in this tech for the industry if used in a restricted way.
We need to work out what crosses the line to protect voiceovers' interests.