Field notes · updated Aug 2026

We made 10 YouTube Shorts with Gathos

Ten vertical shorts, one batch, one flat-price account. Every image and every voice in them came from Gathos — scripts, AI images, TTS narration, upbeat music, captions, thumbnails. Here's the exact workflow we used, and the two mistakes we fixed on the second pass.

The same stack you'd need to reproduce this: Gathos Pro $18/mo (image + voice)

The pipeline

From brief to finished Short in six steps

Step 1 · Brief

Pick the pain points

Each Short targets one searchable problem — "why is ElevenLabs so expensive", "why AI images can't spell", "meter anxiety". One pain point, one 55-second video, one clear call to action.

Step 2 · Scripts

Five scenes, hook first

HOOK (0–5s) → PAIN → REVEAL → SOLUTION → CTA, at ~145 words per minute. The hook is the first three seconds — no intros, no setup, straight in.

Step 3 · Setup

One batch, zero API calls

A setup script creates all runs locally — scripts, style cards, upload descriptions. Nothing hits the API until you tell it to render.

Step 4 · Render

Images → voice → assembly

Gathos generates the visuals and narration; FFmpeg assembles with Ken Burns motion, baked captions and upbeat background music. Roughly 12–16 API calls per Short.

Step 5 · Verify

Check, don't assume

Every output gets checked: 1088×1920 vertical, audio track present, thumbnail generated, the right voice on the right video.

Step 6 · Upload

Metadata done for you

Each video ships with a title, description, tags and thumbnail — the description carries the link to the matching review page and the free trial.

Lesson one

Voice selection is half the quality

Gathos ships six preset voices: Josh, Koko, Pixxy, Prof, Rochie and Spraky — plus zero-shot cloning of your own voice. Our first pass used the warm preset for two videos and it came out too slow and too high-pitched for punchy Shorts. The fix wasn't more editing — it was choosing the right voices and steering the delivery.

VoiceCharacterBest for
JoshEnergetic male presenterMoney, tool and creator topics — our default
ProfAuthoritative malePricing numbers, comparisons, technical fixes
KokoConfident female, authoritativeSavings and workaround topics needing authority
PixxyLively female, fast and punchyHigh-energy, anti-friction messages
SprakyBright female, fastVariety and A/B tests
RochieWarm female— (too slow/high for Shorts pacing)

One more thing: you can pass a short narrator description to steer delivery — "fast-paced, energetic, authoritative" changed the perceived pace dramatically without touching the script. Same words, very different video.

Lesson two

Text in images: short, exact, at the end

Our first pass asked the image API to render URLs and sentences like "full breakdown at gathosreview.com". AI image models garble that. The rule we now follow: no URLs, no sentences — only short exact labels, placed at the end of the visual description.

Don't ask forAsk for
bold text: full breakdown at gathosreview.coma clean smartphone with a single button labelled Start Free
a price tag labelled voice cloninga padlock icon and a small plain price tag
labelled eighteen dollars a montha small label that reads $18
the total of one hundred and thirty four dollars circledthe total marked with the number $134 circled in red
crisp readable headline (with no text specified)a crisp readable headline that reads TODAY

Why it matters: the URL and the call to action live in the narration and the video description — where they're always spelled right. The image just needs to look clean and trustworthy.

The honest part

What queue congestion actually looks like

Gathos runs on a 600-submission / 6-hour fair-use window with queue controls. On a busy evening, image jobs sat in the queue for 5–15 minutes and occasionally timed out. Our batch of 10 shorts (~140 API calls total) handled it with automatic retries — a self-healing loop re-runs any unfinished video until its MP4 exists. It turned a ~40-minute theoretical render into a few patient hours, with zero manual intervention. One honest limit to plan around, and the retry path is already built into the workflow.

Tool stackTypical monthly costThis batch used
Image generator (Nano Banana-style)$134–$240 per 1,000 images~140 calls, flat-rate account ($18 Pro for image + voice)
Voice + cloning (ElevenLabs-style)$5–$990 metered by character
Video / assembly toolingUsage-based or extra subscription

Full cost comparison on the pricing page and the hands-on review.

Reuse it

The same workflow powers any batch

The batch pattern is reusable: change the pain points, re-run the setup, render. Faceless channel content, affiliate promos, course ads, multilingual narration — it's the same pipeline. We now run it through an agent skill that carries the voice and text rules forward, so the first render comes out right instead of the third.

🎙️ Voice rules, saved

Which presets work for Shorts pacing, which to avoid, and the narrator-style steering that fixes delivery speed.

🖼️ Text rules, saved

Short exact labels only, label at the end of the visual description, never render URLs or sentences in images.

Your Shorts stack, one flat price.

Images, voices and video from one platform. 7-day trial, no credit card.

Start your free 7-day trial →

No credit card · cancel anytime