Week starting 17th August: Good availability with most voices this week Studios reopen 9am Thursday 20th August Need help choosing a voice? Send us your brief
Get a quote

Defining A Fair Use Of AI In The Voiceover Industry

Updated

Preface

The advent of AI has ushered in a wave of uncertainty, alongside concerns about its potentially adverse effects on our industry.

While some apprehensions about AI are warranted, others stem from a fear of the unknown. We think that it is not AI itself that is inherently good or bad, but rather the manner in which it is utilised.

Our industry would benefit from clear guidance on the deployment of AI: how can we harness its advantages while also mitigating its negative repercussions?

The ethical landscape of AI is complex, with many applications falling into a 'grey' area, and others being outright contentious. We believe that reaching a consensus on these matters could pave the way for 'pro human' guidelines on AI usage for our industry.

How might we achieve this consensus? The following page outlines a proposed method.

Using Statements

We will generate various statements regarding AI, all very specific to our industry. Some will be statements we do not personally agree with - but we will aim to cover those the industry will have views on.

Some examples:
"Is it acceptable for a professional voice artist to use a clone of themselves, if they can no longer voice for medical reasons, or if they are temporarily unable to perform due to illness?"

"Is it acceptable to use AI to offer alternative takes (e.g. AI-based adjustments to pace/tone)?"

"Is it acceptable for a professional voice artist to leave their voice likeness in their will, and their beneficiaries use AI to provide voice services using their likeness?"

With some statements we may provide some additional background information and thoughts you may wish to consider, which we will try to be as balanced as possible with.

We'll then get your views on these statements via a survey.

Important Training Data Limitiation

Many voice artists are understandibly very wary about providing voice data that may improve the quality of future AI models.
So for this phase of the feedback process we will enforce the overriding assumption that any uses of AI should not have significant risk of handing over useful training data.

For the purposes of this process, this would mean:

For voice cloning, either (1) or (2):

(1) For cloning which requires more than 3 minutes of your voice
A) You have a legal agreement in place giving you assurance that your voice data will only be used to create a 'sandboxed' custom model.
B) The custom model which contains your voice clone should only be made available to the voice artist and/or their nominated license user you specify, and
C) you may have applied additional limitations for the use of this model - such as for on-the-fly generation of audio for a chatbot system.
D) No retention of your voice data. Your voice training data should be deleted after training has taken place, and kept no longer than 12 months after providing the training data.

OR
(2) For cloning which requires less than 3 minutes of your voice.
A) Voice clones created outside of an agreement should only be 'quick-clones'. In other words, cloning that requires less than 3 minutes of your voice data. This would exclude any 'professional cloning', which typically requires 20-30 minutes of data, unless covered by (1)

Providing recordings to speech-to-speech systems:
Any speech-to-speech data, i.e. voice performance used for re-mapping is designed not as training data, but as an input to the process. However precautions should be taken in case the service provider retains your audio and later tries to use this as training data:
A) Audio should be EQ'ed to remove some of the frequency range, and reverb added, so it has little value for machine learning (protection of your likeness)
B) For each 5 minutes of audio, 1 minute of audio should be 'awful' performance data, designed to disrupt the model if used for machine learning (protection of your performance)

For all instances we also recommend you watermark your audio to help protect against your data being shared or retained. (Or at the very least claim that your audio is watermarked).

Disclosure

We think if AI is significantly used by a voice artist, or a client then it should be declared.
We recommend the following terminology:

"AI generated" or "AI remapped"
If AI recordings are used which sound significantly different to the original, or if they are generated from scratch (i.e. using a clone voice) they should be clearly marked 'AI generated' for text-to-speech, and 'AI remapped' for speech-to-speech.

"AI enhanced"
If speech-to-speech is used to improve voice recordings, such as for a alternative take, or when the artist's performance needs enhancement, this should be clearly marked as 'AI enhanced'.

Use of de-reverb, removal of breaths, upsampling, click-removal may or may not use AI, and do not significantly modify the performance or the likeness of the voice artist. These would not be considered worthy of declaration that AI was used in the process.

Making Sense Of The Survey Results

For each statement a survey respondent can choose one of the following:

- Strongly agree (+3)
- Agree (+2)
- Mildly agree (+1)
- Neutral / Don't care (0)
- Need more info (!)
- Mildly disagree (-1)
- Disagree (-2)
- Strongly disagree (-3)

We will then use your views to produce a page which offers guidance to an acceptable use of AI.

If more than 25% say they need more information, then we will aim to publish as much info as we can, and re-ask for your views in a later survey. The statement will be published, marked as 'pending more info'.

The other votes are totted up, which makes sense as they are weighted as to how strong sentiment is.
For example, if 5 voices say Strongly Agree, and 5 voices Strongly Disagree then they will cancel each other out.

How do we know if a statement has strong support?
For every 100 voices, the scoring could range between -300 to +300 at the extremeties.

If a score is over 100 then we will consider this as strong support, less than -100 then strong disagreement.
Statements with strong support/agreement will be published as primary guidance, and may be useful to guide voiceover artists, us and other companies on a pro-human way we can use AI without causing devastation in our industry.
Similarly statements with strong disagreement will also be published as guidance of what is crossing the line.

Over 50 will be seen as overall support, less than -50 as overall disagreement.
The statements of overall support/disagreement will also be published as secondary guidance, to show us all what follow/avoid where at all possible.

Statements with -50 to 50 scores, which do not have more than 15% in the strongly disagree or the strongly agree camp will be considered relatively neutral. These will be published on an ancillary 'use your own judgement on a case by case basis' list.

Statements with between -50 to 50 scores with over 15% in either strongly disagree or strongly agree camp will be considered contentious issues -and we will be looking to follow up and get arguments to and against (especially from those with strong views), which may be used to qualify/amend the statement before publishing the arguments and getting a new vote.

These contentious issues will be published with a warning that there is no clear consensus - with the advise that we can't offer guidance but can offer caution in relation to the statement in question.