Field notes · updated Aug 2026
We made 10 YouTube Shorts with Gathos
Ten vertical shorts, one batch, one flat-price account. Every image and every voice in them came from Gathos — scripts, AI images, TTS narration, upbeat music, captions, thumbnails. Here's the exact workflow we used, and the two mistakes we fixed on the second pass.
The same stack you'd need to reproduce this: Gathos Pro $18/mo (image + voice)
The pipeline
From brief to finished Short in six steps
Step 1 · Brief
Pick the pain points
Each Short targets one searchable problem — "why is ElevenLabs so expensive", "why AI images can't spell", "meter anxiety". One pain point, one 55-second video, one clear call to action.
Step 2 · Scripts
Five scenes, hook first
HOOK (0–5s) → PAIN → REVEAL → SOLUTION → CTA, at ~145 words per minute. The hook is the first three seconds — no intros, no setup, straight in.
Step 3 · Setup
One batch, zero API calls
A setup script creates all runs locally — scripts, style cards, upload descriptions. Nothing hits the API until you tell it to render.
Step 4 · Render
Images → voice → assembly
Gathos generates the visuals and narration; FFmpeg assembles with Ken Burns motion, baked captions and upbeat background music. Roughly 12–16 API calls per Short.
Step 5 · Verify
Check, don't assume
Every output gets checked: 1088×1920 vertical, audio track present, thumbnail generated, the right voice on the right video.
Step 6 · Upload
Metadata done for you
Each video ships with a title, description, tags and thumbnail — the description carries the link to the matching review page and the free trial.
Lesson one
Voice selection is half the quality
Gathos ships six preset voices: Josh, Koko, Pixxy, Prof, Rochie and Spraky — plus zero-shot cloning of your own voice. Our first pass used the warm preset for two videos and it came out too slow and too high-pitched for punchy Shorts. The fix wasn't more editing — it was choosing the right voices and steering the delivery.
| Voice | Character | Best for |
|---|---|---|
| Josh | Energetic male presenter | Money, tool and creator topics — our default |
| Prof | Authoritative male | Pricing numbers, comparisons, technical fixes |
| Koko | Confident female, authoritative | Savings and workaround topics needing authority |
| Pixxy | Lively female, fast and punchy | High-energy, anti-friction messages |
| Spraky | Bright female, fast | Variety and A/B tests |
| Rochie | Warm female | — (too slow/high for Shorts pacing) |
One more thing: you can pass a short narrator description to steer delivery — "fast-paced, energetic, authoritative" changed the perceived pace dramatically without touching the script. Same words, very different video.
Lesson two
Text in images: short, exact, at the end
Our first pass asked the image API to render URLs and sentences like "full breakdown at gathosreview.com". AI image models garble that. The rule we now follow: no URLs, no sentences — only short exact labels, placed at the end of the visual description.
| Don't ask for | Ask for |
|---|---|
| bold text: full breakdown at gathosreview.com | a clean smartphone with a single button labelled Start Free |
| a price tag labelled voice cloning | a padlock icon and a small plain price tag |
| labelled eighteen dollars a month | a small label that reads $18 |
| the total of one hundred and thirty four dollars circled | the total marked with the number $134 circled in red |
| crisp readable headline (with no text specified) | a crisp readable headline that reads TODAY |
Why it matters: the URL and the call to action live in the narration and the video description — where they're always spelled right. The image just needs to look clean and trustworthy.
The honest part
What queue congestion actually looks like
Gathos runs on a 600-submission / 6-hour fair-use window with queue controls. On a busy evening, image jobs sat in the queue for 5–15 minutes and occasionally timed out. Our batch of 10 shorts (~140 API calls total) handled it with automatic retries — a self-healing loop re-runs any unfinished video until its MP4 exists. It turned a ~40-minute theoretical render into a few patient hours, with zero manual intervention. One honest limit to plan around, and the retry path is already built into the workflow.
| Tool stack | Typical monthly cost | This batch used |
|---|---|---|
| Image generator (Nano Banana-style) | $134–$240 per 1,000 images | ~140 calls, flat-rate account ($18 Pro for image + voice) |
| Voice + cloning (ElevenLabs-style) | $5–$990 metered by character | |
| Video / assembly tooling | Usage-based or extra subscription |
Full cost comparison on the pricing page and the hands-on review.
Reuse it
The same workflow powers any batch
The batch pattern is reusable: change the pain points, re-run the setup, render. Faceless channel content, affiliate promos, course ads, multilingual narration — it's the same pipeline. We now run it through an agent skill that carries the voice and text rules forward, so the first render comes out right instead of the third.
🎙️ Voice rules, saved
Which presets work for Shorts pacing, which to avoid, and the narrator-style steering that fixes delivery speed.
🖼️ Text rules, saved
Short exact labels only, label at the end of the visual description, never render URLs or sentences in images.
Your Shorts stack, one flat price.
Images, voices and video from one platform. 7-day trial, no credit card.
Start your free 7-day trial →No credit card · cancel anytime