Week starting 17th August: Good availability with most voices this week Studios reopen 9am Thursday 20th August Need help choosing a voice? Send us your brief
Get a quote

Notable AI Tech (3 of 3)

Voiceovers Team
Updated

Intro

We think it's important to highlight some of the AI tech which directly affects our industry.
This section covers Text to Speech.

3 of 3 - Text-To-Speech Systems

Text to speech is potentially the number one AI tech which directly affects voiceover artists.

This is because it can take the voice artist out of the equation completely. You can generate fairly realistic speech with a likeness that may not even be based on a real voice.

This could hit the industry like a sledgehammer if it's adopted more and more by business.

The text to speech models are created using lots of different recordings voiced by different people. This gives then them the ability to be able to clone a persons voice.

You can 'rapid-clone' a voice with a short dry recording of less than a minute.
Some providers also offer a more 'in-depth cloning' service where you can provide a lot of recordings, and it will aim to clone you perfectly.

Let's be clear here, when rapid-cloning a person it's pretty much only cloning your timbre of voice and approximating your accent if that accent has already been cloned. The likely result of a quick clone is your approximate timbre, but with a neutral British accent (or American accent if you're American). In some cases this is enough for it to sound like you, for others it can be wildly different. Also if the system has 'ingested' your voice already then it's likely to create a realistic clone. Some of the Stephen Fry clones are extremely accurate, as these systems have ingested enough of him through various websites, intereviews, audio books.

With in-depth cloning it will collect more information about your voice, and may also be able to clone your accent relatively well.
However these systems are not able to clone your ability to bring out all the right elements of a script, or interpret the delivery like you would. So it's not a clone of you, just something who sounds a bit like you, reading the script in some kind of aggregated common way.

These systems have mainly been trained on neutral accented data. When we tested them they were unable to accurately do accents. The companies are itching for pro voices to go ahead and be professionally cloned, and I wonder if one reason is so their models can learn new accents.

Pros for voices

You may be able to use the tech to offer services you don't normally. Such as extremely long audio books. It's also likely in the future that more companies will have their own 'siri-eque' operators, instead of speaking with a human call handler. These companies may want to want to license the likenesses of some pro voices for these systems, who are complementary to their brand image.

Cons for voices

Whether your likeness has been used or not, Text to Speech systems may start to replace voice artists. In some cases the content won't be considered particular brand enhancing, more functional. Perhaps a how-to use a product, or take medicine. The client or production company may think the general public either won't realise or won't care. They don't need to use a likeness of any known human. (Though in reality this probably sounds like someone).

Limitations of TTS systems (we can exploit)

Please note tech is developing all the time so this is a snapshot of current limitations.

Trained on mainly audio book and Youtube data. Typically does not understand how to read a telephone number, or a web address convincingly.

Pronunciations can vary for the same word in the same script.

The system will seemingly randomly apply stresses to words. Sometimes it gets it right, most of the time it gets it wrong.

Realistically to use this tech and get an acceptable result you'd need to pump in paragraph by paragraph, regenerating each a few times until you get a 'take' which is acceptable. And then edit together all the paragraphs.

This is a time consuming process requiring some audio editing skills, and is probably not the easy solution as many people may think it is.

Listen to these awful (and funny) samples of it trying to say 'WWW'.

An Aside

We have to hope that the more AI voices are used by hobbyists and small companies, the more the larger companies will need to do to stand out. If they have the budget of course.