FOTOhub launches UGC Ads
A reusable AI actor, a four beat script, and a batch engine that turns one ad idea into every variant a media buyer needs to test, built into the same account as FOTOhub's image and video tools.
FOTOhub runs 926,000 users across 46 countries on a single orchestration
layer of 226 AI models. Across that base, through two full quarters, one
request kept arriving in the same shape from wildly different kinds of
customers: a Shopify store doing six figures a month, an agency running
paid social for eleven clients, a mobile game studio testing creative in
nine languages. All of them wanted the same thing. Give us a way to produce
UGC style ad creative without booking a human being for every hook, every
product angle, every language variant, every time the ad account starts to
fatigue.
Today that request ships. UGC Ads is live at /fh/ugc.
It takes one registered AI actor and one structured script, and turns them
into as many finished vertical ad variants as a performance team has the
budget and the patience to test. It runs inside the same FOTOhub account
customers already use for image generation, video generation and asset
storage, which means the product photos, logos and brand assets already in
the library are available to the ad the moment it gets written.
The problem is not whether UGC works, it is that it does not scale
The case for UGC style creative was settled years ago. Feed native video
shot on a phone, delivered by a person talking straight into the lens,
reliably beats polished studio production on cost per acquisition across
almost every channel where both have been tested. Buyers know this. It is
not a controversial claim anymore.
The constraint has always been production, and production has a hard
ceiling that money does not lift cleanly. A creator can film one script.
Maybe they will shoot two takes of a hook if the brief is clear and the
day is going well. Then they are done, because they are a person with a
calendar. A brand that wants to know which of five opening lines actually
holds attention on a specific placement needs five shoots, or five
creators, or one very accommodating creator and a week of turnaround.
Multiply that by three voice registers to find out whether the audience
responds better to calm authority or fast enthusiasm, then multiply again
by four languages because the product ships across Europe, and the
production plan has sixty deliverables in it. Nobody runs that plan. What
happens instead is that teams shoot two ads, run them until performance
decays, and then shoot two more, which means they are always testing a
sample size too small to learn anything durable from.
UGC Ads exists because that arithmetic never closes at the speed
performance marketing actually operates at. The point is not to replace
the creator who genuinely connects with an audience. The point is that the
thirty variants nobody was ever going to film should not stay unfilmed
purely because filming them was impossible.
Four steps, not one button
UGC Ads is deliberately not a prompt box that emits a finished commercial.
Ad creative has structure, and the structure is where the performance
lives. So the studio is built as four distinct stages, each of which
produces a reusable, inspectable, editable artifact: an actor, a
blueprint, a render, and a batch.
That separation is the whole design. An actor outlives any single
campaign. A blueprint outlives any single render. A batch is just a
blueprint pointed in several directions at once. Because each layer is
saved independently, the expensive creative decisions get made once and
then reused, and the cheap mechanical work is what gets repeated.
Step one: build an actor once, cast them forever
An actor in UGC Ads is a persistent, reusable synthetic performer, not a
face that appears in one video and is gone.
Creating one starts in the actor library at /fh/ugc/actors. There are
three ways in. Pick one of six preset appearances if the goal is to get
moving quickly. Or specify the performer precisely across eleven
appearance fields covering age, gender, hair, eyes, build, skin, wardrobe,
accessories, setting and overall vibe, which is how teams match an actor
to a specific customer segment rather than settling for whoever the model
felt like generating. Or upload a real person's face, with a consent
record attached: who granted the likeness rights, and when. That consent
entry is stored with the actor, not handled as an afterthought, because a
brand using a real face in paid media needs to be able to answer that
question months later.
Whichever route, FOTOhub generates up to four candidate faces and the
operator chooses. Nobody gets one take and a shrug. If a candidate is
close but not right, casting can be pushed further in that direction, and
looks worth keeping go into a face gallery for reuse across future
projects, which over time turns into a house casting library rather than a
pile of one off generations.
Once a face is selected, the system builds a six view character sheet
automatically: front, three quarter, profile, half body, closeup and full
body. This is the piece that makes the rest of the product work. Wardrobe,
lighting and framing consistency across scenes is the single most obvious
tell in synthetic video, and the character sheet is what holds it. It is
also what gets registered with the video provider so the actor can
actually be filmed. An actor without that registration is, in the system's
own terms, unfilmable, and the studio says so upfront rather than failing
halfway through a render.
Step two: write the ad as a blueprint
The blueprint is the script, the shot list, the audio plan and the output
spec in one saved, versioned object. It lives at /fh/ugc/:projectId.
Ads in the studio are structured as beats with explicit roles: hook,
problem, demo, proof, and call to action. The default is a four beat
structure because that is what most short vertical ads actually are, but
the beats are editable. Naming the role matters more than it might look:
a scene that knows it is the hook can be scored, swapped and A/B tested as
a hook later, which is exactly what the batch stage depends on.
Per scene, an operator writes three things. The line, meaning what the
actor says. The shot, meaning the framing and the action, described the
way a director would note it. And the assets, meaning real product photos,
packaging shots or logos pulled from the account's own library and linked
into the scene as visual reference, so the thing being advertised is the
actual product rather than a plausible imitation of it.
Audio is a first class decision with three strategies, and the right
choice differs by use case. Native generates lip timing along with the
scene, which is the fastest path and usually the right default. Redub
re-records the voice track after the video is filmed, which is what teams
use when the same footage needs several different voice reads. Reference
audio syncs the performance against a track the brand already owns, which
is how an existing brand voice, a licensed read or a real spokesperson's
recording gets onto a synthetic performer.
Voice itself comes from three engines: Grok, an OpenAI audio model, and
FOTOhub's own IDA Voice cloning engine, with selectable voice, language
and speaking rate. Language selection here is the quiet unlock for
anything sold across borders, because a localized variant stops being a
new production and becomes a new render of a blueprint that already works.
There is an optional lip sync pass that runs after the video is generated,
in three quality tiers: fast, HD and ultra. It corrects mouth timing
against the final audio, and it is a separate stage on purpose, because a
draft being reviewed internally does not need ultra and a hero asset going
into a six figure ad spend does.
Output is specified explicitly: vertical 9:16, 720p, 24fps, with an
optional end card hold if the ad needs to sit on a final frame while a
logo or offer lands.
Throughout all of this, the editor shows a live estimate. Not a vague
indication that generation costs money, but a breakdown: the scenes, the
total seconds of video, the character count going to text to speech, the
number of redub passes, and the resulting cost. It updates as the script
changes. A team can lengthen a scene, watch the number move, and decide
whether that extra two seconds is worth it before committing to anything.
Blueprints are saved and versioned, so revising a script next month means
editing a document, not rebuilding a project.
Step three: render, with the failure modes handled
Submitting a blueprint starts a render job, visible at
/fh/ugc/:projectId/render/:jobId, and the engine films up to four scenes
concurrently on ByteDance's Seedance 2.0.
The order of operations inside a scene is the interesting part. Voice
generates first. That is not arbitrary sequencing: generating the audio
first means the system knows the exact duration the scene needs before it
asks the video model for a single frame. Video is expensive and duration
is the thing it is priced on, so measuring the audio instead of guessing
at the video length is the difference between paying for what the script
requires and paying for a padded estimate.
From there each scene moves through explicit states: pending, voicing,
voiced, submitting, polling, persisting, done. The job surface exposes
where every scene actually is rather than showing one aggregate spinner,
because a four scene render where scene three is stuck is a different
situation from one where all four are simply slow, and an operator who
can see which one it is can make a decision.
Scenes checkpoint independently. If a scene fails partway through, the
scenes that already completed are kept. This sounds like a small
engineering detail and is in practice one of the most important properties
in the product: generative video pipelines fail intermittently for
reasons nobody controls, and a system that discards three finished scenes
because the fourth timed out is a system that charges customers for work
it then throws away. Finished clips are written to private storage in the
account, and an optional lip sync pass runs before persistence when it has
been requested.
Step four: batch, which is where the volume actually happens
Everything above produces one ad. The batch stage, at
/fh/ugc/:projectId/batch, is what produces a test plan.
Take a blueprint that already rendered well and fan it out along any axis
without rebuilding it. Swap the hook line to test openings. Swap the voice
or the language to test delivery and reach new markets. Swap the actor
entirely to test whether the audience responds to a different kind of
person. Combine axes and the studio shows the resulting variant matrix
along with its total cost before anything is submitted, so a team sees
that three hooks against two voices is six renders at a specific price and
can decide to trim it to four before spending.
On submission, each variant becomes its own job, and the batch record
tracks the plan against the actual result, meaning what was supposed to
render versus what came back. That reconciliation exists because a batch
of twenty four variants where two failed is a normal outcome, and a team
needs to know which two rather than discovering the gap when they go to
upload.
The economics of this stage are the reason the whole four step structure
is shaped the way it is. The actor was built once. The blueprint was
written once. So the marginal cost of testing one additional hook is the
cost of rendering that hook, not the cost of a new production. Creative
testing stops being a budget line that has to be defended and becomes
something closer to a query.
For teams who would rather not write the script at all
Not every operator wants to write four beats by hand, particularly at the
top of a campaign when the angle is not yet obvious. So there is a brief
driven path in front of the blueprint.
Describe the product once. FOTOhub generates twelve distinct ad angles,
genuinely different approaches rather than twelve rewordings of the same
pitch. The operator reads them and keeps four. For each kept angle, the
system writes a complete beat by beat script ready to be edited or filmed
as is.
The gate is deliberate. Twelve angles get proposed, a human selects, and
only the selected ones get scripted and filmed. Nothing renders that a
person did not choose, which keeps a brand's output from drifting into
whatever an unattended model found statistically comfortable.
What this looks like in practice
A skincare brand launching a serum builds one actor matching its target
customer, then writes a blueprint: hook naming a frustration the customer
already has, problem making it specific, demo showing the serum going on
skin with the real bottle linked from the asset library, proof citing a
result, call to action pointing at the product page. It renders once and
gets reviewed. Then batch swaps in two alternate hooks and two voice
reads, and by the afternoon there are six finished variants in the ad
account. A conventional shoot would still be in pre production.
An agency running eleven accounts builds one actor per client and keeps
them in the casting library, which means every ad a client gets features a
consistent spokesperson across months of campaigns without re-briefing a
freelancer or explaining the brand again. When a client asks for a fresh
batch, the blueprint from last quarter is still there and still versioned.
A mobile app growth team uses batch as an instrument rather than a
production tool. Six hooks, one actor, one script body, all rendered
overnight, all pushed as separate ad sets, and the answer to which
opening actually retains attention arrives in days instead of quarters.
A brand selling across five European markets writes the ad once and
renders it five times with the language swapped, which turns
localization from five shoots with five casting problems into one
blueprint and five voice selections.
What it costs
Cost is visible before anything renders, and the estimate is itemized
rather than bundled, because the components scale differently and teams
should be able to see which lever they are pulling.
Every cost is itemized rather than bundled, because the components scale
differently and a team should be able to see which lever it is pulling:
- Video generation: $0.05 per second of finished footage
- Voice: $0.015 per 1,000 characters of script
- Optional redub pass: $0.008 per second of video
- Lip sync pass: tracked in GPU seconds
Worked through on a realistic shape, a four scene ad running thirty seconds
comes to roughly $1.50. Adding a redub pass across the whole thing takes it
to about $1.75. Voice is close to a rounding error in that calculation,
since a tight four beat script runs a couple of hundred characters and
prices in fractions of a cent, which means the number a team is actually
managing is video seconds.
That reframes what a creative test costs to plan rather than what a single
ad costs to make. Three hooks against two voice reads is six finished
variants for something near $9. A wider matrix, say four angles across
three actors in two languages, is twenty four finished ads for under $40.
Prices are shown in USD in the editor before submission, per variant and
for the batch as a whole, so the decision to widen or trim a test happens
with the number visible.
Because batch reuses an existing actor and an existing blueprint, the
marginal cost of one more variant is only that variant's render. The
expensive, judgment heavy work of casting a performer and writing a script
that converts is amortized across every test that follows it.
Consent, provenance and using this responsibly
Two things are built in rather than bolted on. Uploaded likenesses carry a
consent record naming who granted rights and when, so a brand can answer
that question later without reconstructing it from email. And synthetic
actors are synthetic: the character sheet system exists to keep a
generated performer consistent, not to make a real person appear to say
something they did not.
FOTOhub's broader AI Act obligations, including transparency around AI
generated media, apply to output from this studio the same as anywhere
else on the platform. Ad creative is one of the places synthetic media
meets the public most directly, and treating that as a compliance
afterthought would be a mistake.
Available now
UGC Ads is live at /fh/ugc on every paid plan, listed under Studios in the
sidebar next to FH Shorts, with a product overview at /product/ugc.
Start by registering one actor. Write four beats. Render once to see
whether the performer and the script hold together. Then open batch, and
find out which hook actually wins.