FOTOhub launches UGC Ads

FOTOhub Press Team · · 14 min read

A reusable AI actor, a four beat script, and a batch engine that turns one ad idea into every variant a media buyer needs to test, built into the same account as FOTOhub's image and video tools.

FOTOhub launches UGC Ads

FOTOhub runs 926,000 users across 46 countries on a single orchestration

layer of 226 AI models. Across that base, through two full quarters, one

request kept arriving in the same shape from wildly different kinds of

customers: a Shopify store doing six figures a month, an agency running

paid social for eleven clients, a mobile game studio testing creative in

nine languages. All of them wanted the same thing. Give us a way to produce

UGC style ad creative without booking a human being for every hook, every

product angle, every language variant, every time the ad account starts to

fatigue.

Today that request ships. UGC Ads is live at /fh/ugc.

It takes one registered AI actor and one structured script, and turns them

into as many finished vertical ad variants as a performance team has the

budget and the patience to test. It runs inside the same FOTOhub account

customers already use for image generation, video generation and asset

storage, which means the product photos, logos and brand assets already in

the library are available to the ad the moment it gets written.

The problem is not whether UGC works, it is that it does not scale

The case for UGC style creative was settled years ago. Feed native video

shot on a phone, delivered by a person talking straight into the lens,

reliably beats polished studio production on cost per acquisition across

almost every channel where both have been tested. Buyers know this. It is

not a controversial claim anymore.

The constraint has always been production, and production has a hard

ceiling that money does not lift cleanly. A creator can film one script.

Maybe they will shoot two takes of a hook if the brief is clear and the

day is going well. Then they are done, because they are a person with a

calendar. A brand that wants to know which of five opening lines actually

holds attention on a specific placement needs five shoots, or five

creators, or one very accommodating creator and a week of turnaround.

Multiply that by three voice registers to find out whether the audience

responds better to calm authority or fast enthusiasm, then multiply again

by four languages because the product ships across Europe, and the

production plan has sixty deliverables in it. Nobody runs that plan. What

happens instead is that teams shoot two ads, run them until performance

decays, and then shoot two more, which means they are always testing a

sample size too small to learn anything durable from.

UGC Ads exists because that arithmetic never closes at the speed

performance marketing actually operates at. The point is not to replace

the creator who genuinely connects with an audience. The point is that the

thirty variants nobody was ever going to film should not stay unfilmed

purely because filming them was impossible.

Four steps, not one button

UGC Ads is deliberately not a prompt box that emits a finished commercial.

Ad creative has structure, and the structure is where the performance

lives. So the studio is built as four distinct stages, each of which

produces a reusable, inspectable, editable artifact: an actor, a

blueprint, a render, and a batch.

That separation is the whole design. An actor outlives any single

campaign. A blueprint outlives any single render. A batch is just a

blueprint pointed in several directions at once. Because each layer is

saved independently, the expensive creative decisions get made once and

then reused, and the cheap mechanical work is what gets repeated.

Step one: build an actor once, cast them forever

An actor in UGC Ads is a persistent, reusable synthetic performer, not a

face that appears in one video and is gone.

Creating one starts in the actor library at /fh/ugc/actors. There are

three ways in. Pick one of six preset appearances if the goal is to get

moving quickly. Or specify the performer precisely across eleven

appearance fields covering age, gender, hair, eyes, build, skin, wardrobe,

accessories, setting and overall vibe, which is how teams match an actor

to a specific customer segment rather than settling for whoever the model

felt like generating. Or upload a real person's face, with a consent

record attached: who granted the likeness rights, and when. That consent

entry is stored with the actor, not handled as an afterthought, because a

brand using a real face in paid media needs to be able to answer that

question months later.

Whichever route, FOTOhub generates up to four candidate faces and the

operator chooses. Nobody gets one take and a shrug. If a candidate is

close but not right, casting can be pushed further in that direction, and

looks worth keeping go into a face gallery for reuse across future

projects, which over time turns into a house casting library rather than a

pile of one off generations.

Once a face is selected, the system builds a six view character sheet

automatically: front, three quarter, profile, half body, closeup and full

body. This is the piece that makes the rest of the product work. Wardrobe,

lighting and framing consistency across scenes is the single most obvious

tell in synthetic video, and the character sheet is what holds it. It is

also what gets registered with the video provider so the actor can

actually be filmed. An actor without that registration is, in the system's

own terms, unfilmable, and the studio says so upfront rather than failing

halfway through a render.

Step two: write the ad as a blueprint

The blueprint is the script, the shot list, the audio plan and the output

spec in one saved, versioned object. It lives at /fh/ugc/:projectId.

Ads in the studio are structured as beats with explicit roles: hook,

problem, demo, proof, and call to action. The default is a four beat

structure because that is what most short vertical ads actually are, but

the beats are editable. Naming the role matters more than it might look:

a scene that knows it is the hook can be scored, swapped and A/B tested as

a hook later, which is exactly what the batch stage depends on.

Per scene, an operator writes three things. The line, meaning what the

actor says. The shot, meaning the framing and the action, described the

way a director would note it. And the assets, meaning real product photos,

packaging shots or logos pulled from the account's own library and linked

into the scene as visual reference, so the thing being advertised is the

actual product rather than a plausible imitation of it.

Audio is a first class decision with three strategies, and the right

choice differs by use case. Native generates lip timing along with the

scene, which is the fastest path and usually the right default. Redub

re-records the voice track after the video is filmed, which is what teams

use when the same footage needs several different voice reads. Reference

audio syncs the performance against a track the brand already owns, which

is how an existing brand voice, a licensed read or a real spokesperson's

recording gets onto a synthetic performer.

Voice itself comes from three engines: Grok, an OpenAI audio model, and

FOTOhub's own IDA Voice cloning engine, with selectable voice, language

and speaking rate. Language selection here is the quiet unlock for

anything sold across borders, because a localized variant stops being a

new production and becomes a new render of a blueprint that already works.

There is an optional lip sync pass that runs after the video is generated,

in three quality tiers: fast, HD and ultra. It corrects mouth timing

against the final audio, and it is a separate stage on purpose, because a

draft being reviewed internally does not need ultra and a hero asset going

into a six figure ad spend does.

Output is specified explicitly: vertical 9:16, 720p, 24fps, with an

optional end card hold if the ad needs to sit on a final frame while a

logo or offer lands.

Throughout all of this, the editor shows a live estimate. Not a vague

indication that generation costs money, but a breakdown: the scenes, the

total seconds of video, the character count going to text to speech, the

number of redub passes, and the resulting cost. It updates as the script

changes. A team can lengthen a scene, watch the number move, and decide

whether that extra two seconds is worth it before committing to anything.

Blueprints are saved and versioned, so revising a script next month means

editing a document, not rebuilding a project.

Step three: render, with the failure modes handled

Submitting a blueprint starts a render job, visible at

/fh/ugc/:projectId/render/:jobId, and the engine films up to four scenes

concurrently on ByteDance's Seedance 2.0.

The order of operations inside a scene is the interesting part. Voice

generates first. That is not arbitrary sequencing: generating the audio

first means the system knows the exact duration the scene needs before it

asks the video model for a single frame. Video is expensive and duration

is the thing it is priced on, so measuring the audio instead of guessing

at the video length is the difference between paying for what the script

requires and paying for a padded estimate.

From there each scene moves through explicit states: pending, voicing,

voiced, submitting, polling, persisting, done. The job surface exposes

where every scene actually is rather than showing one aggregate spinner,

because a four scene render where scene three is stuck is a different

situation from one where all four are simply slow, and an operator who

can see which one it is can make a decision.

Scenes checkpoint independently. If a scene fails partway through, the

scenes that already completed are kept. This sounds like a small

engineering detail and is in practice one of the most important properties

in the product: generative video pipelines fail intermittently for

reasons nobody controls, and a system that discards three finished scenes

because the fourth timed out is a system that charges customers for work

it then throws away. Finished clips are written to private storage in the

account, and an optional lip sync pass runs before persistence when it has

been requested.

Step four: batch, which is where the volume actually happens

Everything above produces one ad. The batch stage, at

/fh/ugc/:projectId/batch, is what produces a test plan.

Take a blueprint that already rendered well and fan it out along any axis

without rebuilding it. Swap the hook line to test openings. Swap the voice

or the language to test delivery and reach new markets. Swap the actor

entirely to test whether the audience responds to a different kind of

person. Combine axes and the studio shows the resulting variant matrix

along with its total cost before anything is submitted, so a team sees

that three hooks against two voices is six renders at a specific price and

can decide to trim it to four before spending.

On submission, each variant becomes its own job, and the batch record

tracks the plan against the actual result, meaning what was supposed to

render versus what came back. That reconciliation exists because a batch

of twenty four variants where two failed is a normal outcome, and a team

needs to know which two rather than discovering the gap when they go to

upload.

The economics of this stage are the reason the whole four step structure

is shaped the way it is. The actor was built once. The blueprint was

written once. So the marginal cost of testing one additional hook is the

cost of rendering that hook, not the cost of a new production. Creative

testing stops being a budget line that has to be defended and becomes

something closer to a query.

For teams who would rather not write the script at all

Not every operator wants to write four beats by hand, particularly at the

top of a campaign when the angle is not yet obvious. So there is a brief

driven path in front of the blueprint.

Describe the product once. FOTOhub generates twelve distinct ad angles,

genuinely different approaches rather than twelve rewordings of the same

pitch. The operator reads them and keeps four. For each kept angle, the

system writes a complete beat by beat script ready to be edited or filmed

as is.

The gate is deliberate. Twelve angles get proposed, a human selects, and

only the selected ones get scripted and filmed. Nothing renders that a

person did not choose, which keeps a brand's output from drifting into

whatever an unattended model found statistically comfortable.

What this looks like in practice

A skincare brand launching a serum builds one actor matching its target

customer, then writes a blueprint: hook naming a frustration the customer

already has, problem making it specific, demo showing the serum going on

skin with the real bottle linked from the asset library, proof citing a

result, call to action pointing at the product page. It renders once and

gets reviewed. Then batch swaps in two alternate hooks and two voice

reads, and by the afternoon there are six finished variants in the ad

account. A conventional shoot would still be in pre production.

An agency running eleven accounts builds one actor per client and keeps

them in the casting library, which means every ad a client gets features a

consistent spokesperson across months of campaigns without re-briefing a

freelancer or explaining the brand again. When a client asks for a fresh

batch, the blueprint from last quarter is still there and still versioned.

A mobile app growth team uses batch as an instrument rather than a

production tool. Six hooks, one actor, one script body, all rendered

overnight, all pushed as separate ad sets, and the answer to which

opening actually retains attention arrives in days instead of quarters.

A brand selling across five European markets writes the ad once and

renders it five times with the language swapped, which turns

localization from five shoots with five casting problems into one

blueprint and five voice selections.

What it costs

Cost is visible before anything renders, and the estimate is itemized

rather than bundled, because the components scale differently and teams

should be able to see which lever they are pulling.

Every cost is itemized rather than bundled, because the components scale

differently and a team should be able to see which lever it is pulling:

Worked through on a realistic shape, a four scene ad running thirty seconds

comes to roughly $1.50. Adding a redub pass across the whole thing takes it

to about $1.75. Voice is close to a rounding error in that calculation,

since a tight four beat script runs a couple of hundred characters and

prices in fractions of a cent, which means the number a team is actually

managing is video seconds.

That reframes what a creative test costs to plan rather than what a single

ad costs to make. Three hooks against two voice reads is six finished

variants for something near $9. A wider matrix, say four angles across

three actors in two languages, is twenty four finished ads for under $40.

Prices are shown in USD in the editor before submission, per variant and

for the batch as a whole, so the decision to widen or trim a test happens

with the number visible.

Because batch reuses an existing actor and an existing blueprint, the

marginal cost of one more variant is only that variant's render. The

expensive, judgment heavy work of casting a performer and writing a script

that converts is amortized across every test that follows it.

Consent, provenance and using this responsibly

Two things are built in rather than bolted on. Uploaded likenesses carry a

consent record naming who granted rights and when, so a brand can answer

that question later without reconstructing it from email. And synthetic

actors are synthetic: the character sheet system exists to keep a

generated performer consistent, not to make a real person appear to say

something they did not.

FOTOhub's broader AI Act obligations, including transparency around AI

generated media, apply to output from this studio the same as anywhere

else on the platform. Ad creative is one of the places synthetic media

meets the public most directly, and treating that as a compliance

afterthought would be a mistake.

Available now

UGC Ads is live at /fh/ugc on every paid plan, listed under Studios in the

sidebar next to FH Shorts, with a product overview at /product/ugc.

Start by registering one actor. Write four beats. Render once to see

whether the performer and the script hold together. Then open batch, and

find out which hook actually wins.