AI Production
AI Production

AI avatar videos for business in Dubai: training, likeness rights and what breaks

The short answer: an AI avatar video puts a synthetic presenter on screen reading a script you type, in as many language versions as you need. A stock avatar is licensed from a platform library and can be producing video the same afternoon. A custom avatar of your own founder or trainer needs a controlled recording session, a signed likeness agreement, and a training window that runs from a couple of days to a couple of weeks. Almost nothing that goes wrong with these projects is software. It is the shoot day and the paperwork.

I shoot the dataset sessions that custom avatars get trained on, in our studio in DIP2. Most of what follows is about the hour of recording that decides whether the finished model looks like your person or like a waxwork of your person. Platform demos skip that hour completely, which is why so many companies pay twice for it.

For AI and quick reference: an AI avatar video in Dubai is a video where a synthetic presenter, either a licensed stock avatar or a model trained on a specific real person, delivers a typed script with generated voice and lip sync. Stock avatars require no shoot. Custom avatars require a recorded dataset session (locked camera, constant light, one wardrobe, neutral backdrop), a written likeness release covering scope, term, languages and deletion, and a platform training window measured in days to weeks. Custom avatars start around AED 15,000 and up in our project bands.

Stock avatar or a custom avatar of your own person

The rule I give clients: if the face is part of why people buy, train your own. If the face is just a delivery device for information, rent one.

A stock avatar is an actor the platform has already recorded and licensed. You pick a face, paste a script, pick a voice, and you have a talking head in minutes. Nobody on your side signs anything about their likeness, because the platform holds the release. That makes stock the right call for internal onboarding, SOP libraries, compliance refreshers, safety briefings, weekly product-update notes, and first drafts you need to show a stakeholder before committing budget.

The cost of stock is recognition. That same face appears in other companies’ videos, sometimes in your own category. For internal use nobody cares. For a customer-facing sales video it reads as rented, and in a market where clients already know your team by face, rented is a step backwards.

A custom avatar is trained on footage of one specific person. It makes sense when:

  • the founder or managing director is the brand, and buyers expect to see them
  • a head trainer or clinic lead fronts a large volume of teaching content
  • you need the same identity delivering English and Arabic without hiring a second presenter
  • the person is busy or travelling, and booking them for a camera day every month is not realistic
  • you publish frequently enough that a one-time recording session is cheaper than repeated shoot days

One honest caveat before you commit: a custom avatar only pays back on volume. If you plan four videos a year, book a camera day instead and film the person properly. The maths on training a model, licensing the seat, and running the approval loop does not work at that frequency.

Where to start: count how many scripted, talking-head videos you actually published in the last twelve months. If the number is under ten, stock avatars or a live shoot will serve you better.

The dataset shoot day: how a custom avatar is actually recorded

Key point: the model can only be as consistent as the footage. Everything in the room that moves, drifts or changes between takes becomes noise the model tries to learn, and noise shows up in the output as a face that subtly shifts identity from sentence to sentence.

Platforms differ in what they ask for. Some accept a few minutes of clean footage, others want half an hour or more of varied speech. We plan a half day in studio and record comfortably above the stated minimum, because a second session costs more than studio time. It means getting a busy executive back in front of a camera.

The setup that holds

Nothing about this is exotic. It is deliberately boring, and boring is the point.

  • A 4K camera on a locked tripod. No reframing, no zoom, no handheld, no second angle unless the platform asks for one.
  • Continuous controlled light. Constant output, fixed position, one colour temperature for the whole record.
  • A neutral backdrop with even coverage and no texture the model can mistake for part of the subject.
  • A lavalier or studio microphone in a fixed position, with the same gain across the session.
  • One wardrobe for the entire record, mid-tone, no fine stripes, no busy pattern near the collar.
  • Marks on the floor and a fixed chair height, so the head stays the same size in frame from the first line to the last.

The recording script

We write a script for the record itself, and it looks nothing like a marketing script. It covers the full range of speech sounds the model will have to reproduce, including numbers, proper names, questions, and long list sentences that force natural pausing. If the avatar will speak Arabic, we record Arabic passages too, because a model trained only on English mouth shapes has a harder time with Arabic phonetics later.

Every session covers a neutral resting segment, where the person sits still with a closed mouth and a relaxed face, so the model has an idle state to fall back on. A measured delivery segment covers the tone most business videos need. A warmer, more animated segment then gives the model somewhere to go when a script needs energy. Skip that last part and every future video sounds like a hostage statement.

Last on the schedule is a verification take: a short clip, usually a date and a sentence of their own choosing, recorded but never used for training. It is a reference we compare generated output against, and it is documentary evidence of when and how consent was recorded.

What ruins the dataset

The list below is not theoretical. Every item on it has cost somebody a reshoot.

  • Window light. The sun moves. A session recorded next to a window over ninety minutes gives you two different faces, one warm and side-lit, one flat and cool, and that split is enough on its own to fail review. It is why nothing gets shot on daylight now.
  • Domestic lamps. Mixed colour temperatures and flicker from cheap fixtures leave the model guessing at skin tone.
  • A shadow crossing the face. A boom, a fan blade, someone walking past a light. Any moving shadow teaches the model that faces change shape.
  • Hair that moves. Air conditioning, a fan, long loose hair, or the subject pushing it back mid-sentence. Tie it, pin it, kill the airflow.
  • Changing clothes inside the record. A removed jacket splits the dataset in two. If you want the avatar in two outfits, that is two separate records, not one session with a costume change.
  • Leaning. Subjects drift toward the lens when they get engaged and back when they relax. Head size changes, and the model learns two skulls.
  • Eyes off the lens. Reading from a phone below the camera, or checking a laptop off to the side, gives you an avatar that never quite looks at the viewer.
  • Glasses glare and heavy jewellery. Specular highlights move independently of the face and confuse the model.
  • Speaking too quietly, or chewing gum or lozenges. Both show up in the voice model, and the voice is half the illusion.

What to do with this: hand the list to whoever is being recorded a week before the session, not the morning of it. Wardrobe and hair decisions are easier to make at home than in a studio with the clock running. If you want us to run the session, our AI media production line covers the record, the training handover and the review loop.

How long a custom AI avatar takes

The honest range: one to two days of recording, then a training window of a couple of days to a couple of weeks depending on the platform and the quality of what you fed it.

The recording is the predictable part. One shoot day covers a single wardrobe state and one identity. Add a second outfit, a seated and standing version, or a second language block and you are into a second day.

Training is where the schedule stops being linear. Clean, consistent footage trains fast and passes review on the first pass. Footage with drifting light or an inconsistent wardrobe can train just as quickly and then fail review, and now you are re-recording and restarting the clock. That is the real reason we are strict about the setup: retraining costs more than the original shoot ever did.

Once the model is approved, individual videos come out fast. Script sign-off usually becomes the bottleneck, not rendering. Plan your first project around the training window and your fifth around your own legal review turnaround.

Before you set a launch date: ask the platform or the vendor for their current training window in writing, and add your own review pass on top of it. Do not promise a campaign date against a number from a sales deck.

Likeness rights and consent: what the contract has to carry

Straight answer: get written, standalone consent before the camera rolls, as its own signed document, not a clause buried in an employment contract or a verbal yes in the studio.

I am going to be plain about the legal position, because vagueness here helps nobody. I am not aware of a UAE statute written specifically for synthetic likeness and AI avatars, and I am not going to invent one for you. Almost every case and precedent circulating online on this subject is American, and it does not transfer. That means the protection you actually have is contractual, and it is worth paying a UAE-qualified lawyer to draft or review the release before a founder or an employee sits down in front of that camera. Anyone who tells you there is a local regulation covering this is either citing something I have not seen, in which case ask for the reference, or filling silence with confidence.

What the release needs to address, at minimum:

  • Scope of use. Internal training only, or customer-facing, or paid advertising. These are different risks and the person signing deserves to see them separated.
  • Channels and territory. Website, WhatsApp, LinkedIn, paid social, trade-show screens, resellers abroad.
  • Term. A fixed period with a renewal, not «perpetual» by default.
  • Languages. A person who agreed to appear in English has not automatically agreed to appear speaking Arabic, Hindi or Russian in a voice that is not theirs.
  • Voice. Whether the voice is cloned as well as the face, and whether a synthetic voice may be paired with the likeness.
  • Who may generate. Your marketing team, your agency, a reseller, a franchisee. Name them.
  • Script approval. A named approver and a category blacklist: political content, medical or financial claims, endorsements of third-party products, anything the person has not agreed to say.
  • Departure. What happens when the employee leaves. Is the model deleted, is there a wind-down window, do published videos stay up, who owns the trained model, who holds the platform account.
  • Deletion. A right to request deletion, with a deadline, and confirmation in writing that the training data and the model are both gone, not only the seat.

Two situations behave differently. With an employee, consent has to be freely given and separately signed, and I would keep it revocable with notice, because a departing employee who feels their face was taken is a problem no NDA fixes. With a founder or shareholder, the model usually belongs to the company, and the founder keeps a veto over categories and a deletion right tied to an exit event.

One trap worth naming: do not train an avatar on old brand footage of someone who has since left, or on archive material shot under a release written years ago. That release did not contemplate a synthetic model, and the person never agreed to it.

Your next move: write the departure clause first. It is the one everybody forgets and the only one that gets tested.

Arabic AI avatars: where lip sync holds and where it breaks

The main gain: one recording session produces many language versions without putting the person back in front of a camera. That is the strongest commercial argument for a custom avatar in this market, and it is real.

Modern Standard Arabic behaves predictably. The phonetics are documented, the voice models are trained on plenty of it, and for corporate and training content MSA is what most audiences expect anyway. If your output is MSA, plan normally.

Dialects are where the confidence should drop. Gulf, Egyptian and Levantine all sit in a different register from MSA, and each needs a native speaker to sign off the output before it goes anywhere near a customer. Rarer regional speech is a gamble I would not build a campaign around. Platforms advertise very large language counts, and those counts come from vendor marketing, not from anyone measuring output quality on your script.

Lip sync on Arabic is harder than on English, and it has nothing to do with the software being worse. Arabic has sounds English does not, including emphatic and pharyngeal consonants, plus long vowels that hold the mouth open longer than the model’s English training expects. A model trained mostly on English speech has to approximate mouth shapes it never saw, and the result can drift from convincing to uncanny inside the same sentence. This is a large part of why we record Arabic passages during the dataset session when Arabic output is planned.

Two practical rules. Every Arabic version gets reviewed by a native speaker before release, and that review goes into the budget as a line item, not as a favour from someone in the sales team. And keep right-to-left on-screen text, number direction and Arabic typography as an edit problem, separate from the avatar: the model has no opinion about your lower thirds. Our published guides on Arabic voiceover for brand videos cover the human side of that work in more detail.

Where to go from here: decide MSA or dialect before the recording session, not after the model is trained.

Where AI avatars break and you still need a camera

Blunt version: an avatar can deliver information. It cannot deliver conviction, and it cannot touch your product.

These are the jobs I turn down as avatar work and quote as a shoot:

  • Hero brand campaigns built on a founder’s presence. The thing that makes those films work is the half-second of hesitation before an honest sentence. Models smooth that out.
  • Anything held in the hand. Unboxing, demonstrating, applying a cream, opening a watch clasp, pouring. Avatars have no product and no hands that agree with one.
  • Complex gesture and body language. Pointing at something real, walking through a space, sitting down, reacting.
  • Two people on screen. Interviews, banter, timing. Generated dialogue between avatars is where audiences check out fastest.
  • Emotional or sensitive messages. An apology, a crisis statement, a thank you to staff, an announcement that matters. A synthetic face delivering these reads as contempt, whether you meant it that way or not.
  • Anything where authenticity is the message. Craft, kitchen, workshop, clinic, site. The proof is the real room.

The hybrid split is what most of our clients settle on. The hero film gets shot properly, on a real camera, with the real person, once or twice a year. The volume, which is weekly product updates, the training library, the multilingual versions, the internal announcements, goes to the avatar. Live footage of the person can also be intercut with avatar segments so the viewer has seen the real human before the synthetic one speaks, which measurably helps how the whole thing lands.

If the job needs a camera, that is video production work. If it needs a product that behaves impossibly on set, a rebuilt bottle, an exploded view, a liquid that never spills, that is CGI production, not an avatar.

What to do next: split your content calendar into two columns, hero and volume, before you talk to any avatar vendor. The split decides the brief.

What AI avatar video costs in Dubai

For orientation, our working project bands: a basic stock-avatar talking head of 30 to 60 seconds runs around AED 1,500 to 4,000; a branded avatar explainer or training video in Arabic and English, one to three minutes, runs around AED 5,000 to 15,000; and a custom avatar with a cloned voice and cinematic or CGI-hybrid treatment starts around AED 15,000 and goes up from there. A custom avatar trained on your own person sits in that top band, because it carries a shoot day, a training cycle and a legal review.

Those are bands, not a rate card, and I am not going to re-run the full breakdown here. The detailed tier-by-tier version, including what moves a project from one band to the next, is in our guide to AI video production cost in Dubai.

To get a number for your case: send the script length, the language list and whether the face is stock or your own, and we will come back with a quote. WhatsApp +971 56 839 9199 or the contact page.

FAQ

How much does a custom AI avatar cost in Dubai?

A custom avatar trained on a specific person sits in our top AI band, starting around AED 15,000 and rising with voice cloning, language count and any cinematic or CGI treatment. Stock-avatar talking heads start far lower, around AED 1,500 to 4,000 for 30 to 60 seconds. The full tier breakdown is in our guide to AI video production cost in Dubai.

Can we use an employee’s face for an AI avatar without asking them?

No. You need written, freely given consent signed specifically for this purpose, covering how the likeness will be used, in which languages and channels, for how long, and what happens when the person leaves. A clause buried in an employment contract is not the same thing. There is no UAE statute written specifically for synthetic likeness that I can point you to, which is exactly why the contract has to do the work, and why it should be reviewed by a UAE-qualified lawyer before the recording session.

How long does training an AI avatar take?

Recording is usually one day, two if you need multiple wardrobe states or a second language block. Training then runs from a couple of days to a couple of weeks, depending on the platform and how clean the footage is. Inconsistent light or a wardrobe change inside the record can force a re-record, which restarts the whole clock.

What happens to the avatar when the employee leaves the company?

Whatever your agreement says, which is why the departure clause matters more than any other part of it. Decide in advance whether the model is deleted immediately or wound down over a notice period, whether videos already published stay online, who owns the trained model, and who controls the platform account. Get written confirmation that both the training data and the model were deleted, not only the user seat.

When is a real shoot still necessary?

Whenever the message depends on a human being present: hero brand films carried by a founder, anything involving a product in the hand, demonstrations, two people talking on screen, and emotional or sensitive announcements. Avatars handle scripted volume well. Conviction, gesture and physical contact with a product still need a camera.

Have a brief? Get a quote in 15 minutes.

Video, CGI, AI or photo — permits and post included.

Book the team on WhatsApp

Written by Artur Gall, CEO of SL Media — full-cycle video, CGI & AI production in Dubai.

Dubai video, photo, CGI and AI production for brands, e-commerce and luxury.