Why I Still Don’t Use AI Voices

Why I Still Don’t Use AI Voices

There is a small joke hiding in this being Backstage 007.

After all, if there was ever a number that suggested technology, gadgets and increasingly implausible ways of getting things done, this is probably it.

But for now, 007 remains strictly analogue.

At least when it comes to the voice you hear on EnTrance.

That probably needs explaining, because this is not an anti-AI article.

In fact, it is almost the opposite.

We Have Used AI Voices Before

Back in 2020, we started experimenting with something called EnTrance Global.

At that point, we already had hundreds of hypnosis sessions recorded in English. The obvious question was whether some of that work could travel further.

We selected and reduced the catalogue, concentrating on some of the most popular sessions, and experimented with automated translation and AI voices across around 30 languages.

The project went online in 2021.

It was an experiment. A way of testing overseas markets and seeing how the material performed, particularly on YouTube.

And broadly speaking, it worked.

There were a few complaints about the AI voices, but not many. For listeners hearing the material in languages I don't speak myself, the results were acceptable enough to test whether there was an audience.

That was five years ago.

Since then, the technology has moved forward enormously.

Translation has improved. AI voices have improved. Everything has become more convincing.

If I were producing those translated versions today, I suspect I would be reasonably happy with what modern technology could produce.

But there is an important difference.

I don't listen to English voices in the same way.

The Problem Isn't That AI Can't Speak

I spend a great deal of my time listening.

Not casually.

Listening for the small things.

The change in emphasis.

The slight emotional movement in a phrase.

A hesitation.

An intention.

The difference between a word simply being spoken and a word being delivered for a particular reason.

When I listen to AI voices in English, those are the things I tend to notice.

That does not necessarily mean an AI voice cannot be made convincing.

Given enough time, I suspect you could adjust pauses, emphasis, rhythm and delivery until much of it was more than acceptable.

But there is a practical question hiding inside that.

How long would it take?

Because if I am spending substantial amounts of time correcting and refining a synthetic performance, there comes a point where it may simply be quicker for me to walk into the studio and record the thing myself.

And even when AI sounds good, there can sometimes be something slightly loose about the performance.

Not wrong enough to immediately stop you listening.

Just a little... sloppy.

There is that phrase again: AI slop.

I don't particularly like the term, but occasionally it fits.

A Hypnosis Voice Isn't Just a Voice

A person can have an excellent voice and still be wrong for a particular job.

An audiobook is performed differently from a television advert.

An advert is performed differently from a reality television voiceover.

A narration is performed differently from hypnosis.

The words may be perfectly understandable in all of them.

But understanding the words is not the same thing as understanding how they need to be performed.

That matters particularly in hypnosis.

The performance carries information alongside the language.

Where the emphasis falls.

When something is relaxed.

When something is allowed to delier.

When a pause creates anticipation.

When it creates resolution.

And where a hypnotic cue needs to be heard.

Some of those decisions are conscious.

When writing or preparing a script, I am already looking for places where the delivery needs to change. Where an inflection needs to be enhanced. Where something needs additional emphasis. Where a particular technique requires a particular approach.

But not everything is planned in advance.

Some of it happens while recording.

Riding the Terrain

The closest comparison I can think of is riding a motorbike.

You see a bump in the road ahead and, almost automatically, you prepare for it.

You adjust your weight.

You lift slightly.

Your body becomes part of the suspension.

Cars and lorries do much of that work mechanically, which means we barely notice the bumps.

But when recording a hypnosis script, I think the voice is doing something similar.

You can see the bumps coming.

The peaks.

The troughs.

The change in the terrain of the language.

And while you may have prepared for some of them, others are felt in the flow of the performance.

The intensity changes.

The pace shifts slightly.

The weight of a phrase changes before you reach it.

You are not simply reading a sequence of sentences.

You are moving through the terrain.

Could AI eventually do that?

Perhaps.

I honestly don't know.

I haven't spent the last year testing the latest generation of AI voice systems. I've been rather busy rebuilding, organising and refining EnTrance itself.

And I have learned enough about technology over the years to avoid confidently predicting what it will not be able to do next.

AI music has moved forward remarkably quickly.

AI voices are improving constantly.

There may come a point where I am fooled myself.

I may already have heard an AI voice somewhere and assumed it was human.

But that is not really the decision I am making here.

I Use AI. Quite a Lot.

The idea that this is about refusing to adopt technology would be completely wrong.

I use AI because it can speed up work.

Not because I want to produce more things simply for the sake of getting more things out of the door.

Because EnTrance is large enough that small changes become large jobs.

There are nearly 200 sessions in the store.

So if something needs changing across the catalogue, it isn't one change.

It can be 200.

Take product descriptions as an example.

I would not ask AI to rewrite the entire catalogue and then blindly put the results live.

But I can take an existing description, feed it into a system and ask for a different approach or a refresh.

That can save time.

It can help me move through a large amount of work more efficiently.

But I am still checking it.

Still deciding.

Still responsible for what goes out.

That distinction matters.

Because my experience of AI so far is that it can sometimes behave like your worst employee.

It starts confidently.

Tells you what it can do.

Explains what it is going to do.

And then, somewhere along the way, it can make a spectacular mistake, lose consistency, misunderstand the task or simply produce something that needs considerable repair.

I experienced this while using AI during work on the EnTrance store.

Numbering and cross-referencing went wrong.

Not once.

Repeatedly.

Fortunately, I am fairly diligent about checking work, and the errors were caught before anything went live.

But it is a useful reminder.

Confidence is not the same thing as accuracy.

And a convincing result is not necessarily a correct one.

The Real Risk Is Switching Off Your Own Judgement

AI can be extremely useful.

But I think one of the dangers is that people can become less critical because the output looks convincing.

I have seen systems confidently invent backstories.

I have seen information presented as fact when I knew firsthand that it was wrong because it concerned me or people I know.

And when challenged, AI can occasionally seem determined to defend the mistake.

That doesn't make the technology useless.

It means it needs supervising.

The same applies to creative work.

The question isn't necessarily whether AI can produce something.

The question is whether you have actually checked what it has produced—and whether you are prepared to put your name to the result.

So Why Draw the Line at the Voice?

This is where I need to be brutally honest.

The decision isn't entirely about proving that a human voice is technically superior.

I believe something simpler.

I think EnTrance listeners expect a human voice.

And that is what I want to give them.

The EnTrance catalogue has an established identity.

The sessions are written, developed and recorded as human performances.

I don't particularly want a situation where customers are browsing the store and some sessions are performed by human beings while others are generated or cloned, with that distinction becoming increasingly blurred.

So the line is clear.

The EnTrance store remains human voice only.

For now, that means my real voice.

An analogue human voice.

And yes, there may be another reason underneath that decision too.

I have worked with hundreds of voice artists over the years.

I know how much skill sits inside a performance that many people simply hear as somebody reading.

And I am genuinely sad to see work disappearing for people whose abilities I respect because technology can now produce something that is good enough for certain applications.

I don't know whether that makes me more resistant to AI voices.

Perhaps it does.

I wouldn't pretend otherwise.

But I also don't think that means I have to reject the technology entirely.

There Are Two Different Paths

The public EnTrance catalogue is one thing.

Working with an individual client may eventually be another.

Imagine an intensive process where a client's sessions are being adjusted regularly over several days.

Producing seven one-hour recordings in a week is not necessarily practical.

But what if a cloned version of my own voice could allow a personalised session to be updated quickly when the circumstances genuinely justified it?

What if that meant the client received a more responsive service?

That is a completely different question.

And if a client understood the choice and said they were comfortable receiving either a recording from me, a cloned version of my voice or another AI-generated solution, I would consider what technology might make possible.

That doesn't mean it is happening.

It doesn't mean the system is ready.

And it doesn't mean I have decided how that future would work.

It simply means I am watching.

And listening.

Because that is how I have always approached technology.

I try things.

I test things.

I look for ways they can genuinely improve the work.

And sometimes the answer is yes.

Sometimes the answer is not yet.

The Boundary Is Deliberate

So this isn't EnTrance versus AI.

It isn't an argument that AI will never be capable of a convincing performance.

It isn't a prediction that I will always be able to tell the difference.

And it certainly isn't a promise that my thinking about technology will never change.

Technology changes.

I change my mind when the evidence changes.

But I don't see any reason to automate something simply because automation has become possible.

For EnTrance, the public catalogue has a clear boundary.

The work in the store remains human.

Outside that boundary, there may eventually be other possibilities worth exploring.

The technology can keep moving.

I'll keep watching it.

But for now, when you press play on an EnTrance session, there is still a person on the other side of the microphone.

And for Backstage 007, that seems an appropriate place to remain strictly analogue.


Tags: #VoiceAI #AIVoice #VoiceCloning #HumanVoice #HypnosisAudio #AudioProduction #Voiceover #HumanPerformance #CreativeTechnology #SoundDesign

Back to blog

Leave a comment

Please note, comments need to be approved before they are published.