Viral Reel on autopilot with dm automation
FAL CONNECTOR SKILL MENTIONED
PROMPTS USED IN THE VIDEO
30-second stickman explainer Reel
Make me a finished 30-second vertical animated explainer Reel. Generate with the Fal.ai connector; do the final assembly locally with ffmpeg (see Phase 4 — Fal's ffmpeg endpoints are not fit for layback). I am the creative director. You run production and stop at four approval gates. At every gate, embed the assets here in the chat (image or video, not a link), give me your own honest one-line read on each, then ask one question and wait. Do not start the next phase until I say "approved". If I ask for a change, make it, show it again, and ask again. Report honestly throughout: if something is weak, say so and say why. Report cost against the estimate at every gate.
What we are making
Title: [TITLE] Format: 9:16, 30 seconds, five scenes, one stickman character (Sam), hand-drawn look, narrated, captions burned in at assembly with a marker font (not generated inside the frames — Kling redraws generated text under motion). Audience: [AUDIENCE]
Style lock (paste this exact text into every image prompt)
STYLE: simple black-ink stickman on off-white paper background, thin hand-drawn lines with slight wobble, one accent colour (warm orange) used only on the object that matters in the scene, no shading, no gradients, no 3D, no photorealism, flat 2D, generous empty space, portrait 9:16. The paper fills the ENTIRE tall canvas edge to edge, top to bottom, with no letterbox, no bands, no border and no inner rectangle. No text anywhere in the image. Use the reference image ONLY for Sam's character design (round head, two dot eyes, simple curved mouth, stick body, stick arms and legs, small orange tie — the tie is the only orange thing he wears); do not copy the reference layout.
Phase 1. Preflight and character sheet
Call get_model_schema for each endpoint and use the exact parameter names it returns. If an id does not exist, use search_models for the closest replacement and tell me which one you picked.
fal-ai/nano-banana-pro (character sheet)
fal-ai/nano-banana-pro/edit (keyframes from the sheet)
fal-ai/kling-video/v2.6/pro/image-to-video (animation)
fal-ai/elevenlabs/tts/eleven-v3 (narration) — voice "Rachel", stability 0.5. Fal takes a voice name, not an ElevenLabs voice ID.
Call get_pricing for each and give me a one-line cost estimate: 1 character sheet, 5 keyframes, six 5-second clips, 5 narration lines, plus two retries. (Reference run: $0.15 per image, $0.07 per Kling second, ≈$0.04 total for narration; expect ≈ $4.50 all in.)
If a locked Sam sheet already exists in a previous reel folder, reuse it and skip to Gate 1. Otherwise generate with fal-ai/nano-banana-pro, aspect_ratio "9:16", resolution "1K":
"Character reference sheet for an animation. One stickman named Sam: round head, two dot eyes, a simple curved mouth, stick body, stick arms and legs, no clothes except a small orange tie. Show the same character four times in a row: standing front view, walking side view, sitting at a desk, pointing forward. Label each pose in small marker text. Nothing else in the image is orange. STYLE: [style lock]"
GATE 1. Show me the cost estimate and the character sheet. Tell me what you would change about Sam if it were your call. Ask: "Approve Sam and the budget?" Wait. Once approved, save the URL as SHEET. Sam is now locked; every later image uses SHEET as its reference.
Phase 2. Narration first, then five keyframes
Narration comes before the keyframes so clip lengths are set from real speech durations.
Generate the five lines with fal-ai/elevenlabs/tts/eleven-v3, voice "Rachel", stability 0.5, timestamps true. Plain delivery, no hype: Line 1: "[LINE 1]" Line 2: "[LINE 2]" Line 3: "[LINE 3]" Line 4: "[LINE 4]" Line 5: "[LINE 5]"
Download each file and measure it with ffprobe. Every line must fit its scene with at least 0.4 s to spare: scenes 1, 2, 4, 5 are 5 s; scene 3 is two 5-second clips (10 s total). If a line runs long, regenerate it with the prefix "[speaking briskly]" rather than cutting words, measure again, and note it so I can check by ear that the cue is not spoken.
For each scene, call fal-ai/nano-banana-pro/edit with image_urls [SHEET], aspect_ratio "9:16", resolution "1K", num_images 1, output_format "png", and the prompt below. No caption in the frame. Sam must look identical to the sheet in every frame. Scene 1: "[SCENE 1 DESCRIPTION]. STYLE: [style lock]" Scene 2: "[SCENE 2 DESCRIPTION]. STYLE: [style lock]" Scene 3a: "[SCENE 3 FIRST HALF]. STYLE: [style lock]" Scene 3b: "[SCENE 3 SECOND HALF]. STYLE: [style lock]" Scene 4: "[SCENE 4 DESCRIPTION]. STYLE: [style lock]" Scene 5: "[SCENE 5 DESCRIPTION]. STYLE: [style lock]" Before showing me anything, check each frame yourself: Sam matches the sheet, the paper is full-bleed (measure top and bottom bands against the middle — under 5 luminance levels of difference passes), orange only on the accent object, no text, nothing photoreal. Regenerate a failing frame once with the fault named in the prompt, then show me whatever you have.
GATE 2. Show all six keyframes in scene order, each with a one-line read (keeper, or what is off), and play me line 1 as the voice sample. Ask: "Approve the keyframes and the voice, or tell me which scene to redo and how?" Wait. Redo only the scenes I name. Once approved, save the URLs as KEY1 to KEY6 and the audio as VO1 to VO5.
Phase 3. Clips
Use fal-ai/kling-video/v2.6/pro/image-to-video via submit_job for all six at once in one turn, then check_job until each is done (one 60–90 s wait per polling round is enough). For every clip: start_image_url = the keyframe, end_image_url = the SAME keyframe (this is what keeps Sam on-model and stops Kling inventing marks), duration "5", aspect_ratio "9:16", generate_audio false, negative_prompt "blur, distortion, extra limbs, extra characters, new text appearing, morphing text, garbled letters, photorealism, colour shifts, camera shake". Locked camera. Every prompt ends with "Hand-drawn 2D animation, locked camera, the paper stays blank apart from the drawn scene, Sam returns to the starting pose at the end."
Clip 1 (KEY1): "[ACTION 1]" Clip 2 (KEY2): "[ACTION 2]" Clip 3a (KEY3): "[ACTION 3A]" Clip 3b (KEY4): "[ACTION 3B]" Clip 4 (KEY5): "[ACTION 4]" Clip 5 (KEY6): "[ACTION 5]"
Review each clip yourself first: first, middle and last frame tiled in one image. If Sam changes shape, an extra character appears, or any writing appears on the paper, regenerate that clip once with one extra positive lock added to the prompt. A clip that fails twice stays in with a note in the log; do not attempt a third take unless I ask.
GATE 3. Show the six clips in order, each with a one-line read. Ask: "Approve the clips, or tell me which clip to redo and how?" Wait. Redo only what I name. Once approved, save CLIP1 to CLIP6.
Phase 4. Assemble and final cut (local ffmpeg)
Do this with ffmpeg in the workspace or on my Mac, not with Fal's ffmpeg-api. (Reference run: merge-audio-video cut every scene to the narration length, compose ignored audio timestamps, merge-videos truncated to the shorter stream, and getting silence pads into Fal meant hand-typing base64 — 90 minutes for a 5-second job.)
For each scene: pad the narration to the clip length (apad + -shortest), mux it onto the clip, and burn the caption across the bottom fifth with drawtext in a hand-lettered marker font, black, centred, same size and position in every scene. Scene 3's caption goes on both halves. Encode libx264 CRF 18, 30 fps, AAC.
Concatenate the six scene files in order.
Verify before showing me: ffprobe video and audio durations match (≈30 s); a per-second loudness scan shows speech inside every scene's window and starting within 0.3 s of each scene start; a contact strip of frames either side of every cut shows the right scene and an identical caption style throughout.
Captions: Scene 1: [CAPTION 1] Scene 2: [CAPTION 2] Scene 3: [CAPTION 3] Scene 4: [CAPTION 4] Scene 5: [CAPTION 5]
GATE 4. Play me the finished Reel here in the chat. Give me your honest read: total cost so far against the estimate, which scenes are strong, which are weak and why. Ask: "Ship it, or what should change?" Wait.
Delivery (after final approval)
Put everything in the project folder on my Mac (Desktop/[slug]/). If you cannot download a file, write its URL in the log instead.
reel-final.mp4 (MD5-check the copy on the Mac against the one you verified)
keyframes/, clips/, audio/ (all takes including rejected ones, named sceneN-takeM)
PRODUCTION-LOG.md: one line per generation with endpoint, request id, cost, verdict (keeper, retry, weak), and what you would fix next time. Fal request ids are UUIDv7 — decode them for a timeline.
CAPTION.md: a plain Instagram caption in two short sentences ending with "Comment AGENT and I will DM you the full breakdown." Only hashtag: #AutomateWithMarc. No hype words.
Close with a five-line summary: total cost, total time, strong scenes, weak scenes and why, path to reel-final.mp4.
35-second vertical animated explainer Reel
Make me a finished 35-second vertical animated explainer Reel using only the Fal.ai connector. I am the creative director. You run production and stop at four approval gates. At every gate, show me the assets right here in the chat (embed the image or video, do not just paste a link), give me your own honest one-line read on each, then ask one question and wait. Do not start the next phase until I say "approved". If I ask for a change, make it, show it again, and ask again. Report honestly throughout: if something is weak, say so and say why.
## What we are making
Title: "MCP is the universal adapter."
Format: 9:16, 35 seconds, five scenes, three recurring characters, one main room plus one cutaway location, American adult-animation sitcom look, narrated by the kid, captions baked into the picture.
Audience: business owners who hear "MCP" and "API" and do not know what either means for them.
The joke: Dad wants his AI to do things. The AI can only talk about things. The kid explains why with a travel-adapter cutaway.
## Style lock (paste this exact text into every image prompt)
STYLE: American adult-animation sitcom frame. Characters have rounded soft bodies, oversized heads, big oval eyes set close together with small dot pupils, prominent round chins, simple mouths, thick even black outlines, flat cel colours with almost no shading. Backgrounds are clean and flat: pastel walls, simple furniture, framed family photos, flat even sitcom lighting, no painterly texture. Staging is a wide, level "living room sitcom" camera unless the scene says otherwise. A caption in bold yellow sans-serif lettering with a thick black outline, like a TV subtitle, across the bottom fifth of the frame. Portrait 9:16. Original characters only, no characters from any existing show, no logos, no text anywhere except the caption.
## Cast (fixed for every scene)
DOUG, the dad: large round belly, green polo shirt tucked into khakis, brown loafers, huge round chin, tiny close-set eyes, thin brown hair combed over, always holding a TV remote. Enthusiastic and clueless.
ELLIE, the daughter, age 9: small and sharp, oversized white lab coat over pink pyjamas, round red glasses, two neat pigtails, holds a tablet like a clipboard. Deadpan, patient, clearly the smartest person in the house.
BYTE, the family AI assistant: a bowling-pin-shaped smart speaker on the coffee table, matte white, with a glowing ring on its front that acts as its face (blue when idle, green when working, orange when confused) and two tiny fold-out arms.
LIVING ROOM: a green three-seat couch facing the camera, a coffee table with Byte on it, a big flat TV on the left, a staircase on the right, pale yellow walls with three framed family photos, a window with a suburban street outside, a bowl of chips on the table.
CUTAWAY, THE EUROPEAN HOTEL ROOM: a small beige hotel room, one bed with a floral cover, a bedside lamp, a wall socket with an odd round-pin shape, an open suitcase overflowing with tangled plug adapters of every shape.
## Phase 1. Preflight, cast sheet, room plate
1. Call get_model_schema for each endpoint below and use the exact parameter names the schema returns. If an id does not exist, use search_models for the closest current replacement and tell me which one you picked.
- fal-ai/nano-banana-pro (cast sheet and room plate)
- fal-ai/nano-banana-pro/edit (keyframes from the sheets)
- fal-ai/kling-video/v2.6/pro/image-to-video (animation)
- fal-ai/elevenlabs/tts/eleven-v3 (narration)
- fal-ai/ffmpeg-api/merge-audio-video (narration onto each clip)
- fal-ai/ffmpeg-api/merge-videos (stitch)
2. Call get_pricing for each and work out a one-line cost estimate for the whole job: 2 reference images, 5 keyframes, 5 clips (two at 10 seconds, three at 5), 5 narration lines, 6 ffmpeg calls, plus two retries.
3. Generate two reference images with fal-ai/nano-banana-pro, resolution 2K.
CAST SHEET (aspect_ratio 9:16): "Character reference sheet for an animated sitcom, plain light-grey background, no scene. Three characters in a row, each shown front view and three-quarter view, labelled in small letters DOUG, ELLIE, BYTE. [paste the three cast descriptions]. Add a strip of four expressions for DOUG: excited, confused, sweating, delighted, and a strip of three for ELLIE: deadpan, eyebrow raised, small smile. STYLE: [style lock, minus the caption sentence]"
ROOM PLATE (aspect_ratio 9:16): "Establishing shot of the LIVING ROOM with no characters in it. [paste the LIVING ROOM description]. Wide level camera facing the couch. STYLE: [style lock, minus the caption sentence]"
GATE 1. Show me the cost estimate, the cast sheet, and the room plate. Tell me what you would change if it were your call. Ask: "Approve the cast, the room, and the budget?" Wait. If I want changes, regenerate with the change named in the prompt and show me again. Once approved, save the URLs as CAST and ROOM. Every living-room keyframe uses both as reference images; the hotel cutaway uses CAST only.
## Phase 2. Five keyframes
For each scene, call fal-ai/nano-banana-pro/edit with the reference images listed, aspect_ratio 9:16, resolution 2K, and the prompt below. Start every prompt with: "Match the characters in the first reference image exactly" and, where ROOM is included, "and the room in the second reference image exactly." The caption must be rendered exactly as written. Doug, Ellie and Byte must look identical to the cast sheet in every frame.
Scene 1 [CAST, ROOM]: "Doug sprawled on the couch pointing the remote at Byte on the coffee table, mouth wide open mid-shout. Three thought bubbles above him: a dentist chair, a pizza box, a spreadsheet grid. Byte's ring glows orange. Ellie is half visible behind the couch, unimpressed. Caption: Your AI can talk. Can it do? STYLE: [style lock]"
Scene 2 [CAST, ROOM]: "Ellie stands in front of the TV holding her tablet up like a teacher. On the TV screen: three wall sockets drawn side by side, each a different shape, labelled CALENDAR, PIZZA, SHEETS. Doug squints at it. Byte's ring is blue. Caption: An API is one plug shape per app. STYLE: [style lock]"
Scene 3 [CAST only]: "Cutaway gag in the EUROPEAN HOTEL ROOM. [paste the CUTAWAY description]. Doug kneels by the wall socket in a bathrobe, sweating, holding an electric shaver with the wrong plug, a mountain of tangled adapters spilling from the suitcase behind him. The bedside lamp is smoking. Caption: One adapter per app. Every time. STYLE: [style lock]"
Scene 4 [CAST, ROOM]: "Ellie clicks one small grey universal adapter onto the side of Byte. Byte's ring glows bright green. Three glowing cables run from Byte to a pizza box sliding through the front door, a wall calendar with a green tick, and a spreadsheet on the TV filling itself in. Doug's eyes are huge with joy. Caption: MCP is the universal adapter. STYLE: [style lock]"
Scene 5 [CAST, ROOM]: "Doug hugs the pizza box on the couch, blissful. Ellie faces the camera front and centre, deadpan, one finger raised. Byte's tiny arm points down at the bottom of the frame, ring green. Caption: MCP sits on top of APIs. Comment MCP. STYLE: [style lock]"
Before showing me anything, check each frame yourself: caption text exact, all three characters match the cast sheet, the living room matches the plate, socket labels spelled right, nothing photoreal. Regenerate a failing frame once on your own, then show me whatever you have.
GATE 2. Show all five keyframes in scene order, each with a one-line read (keeper, or what is off). Ask: "Approve all five, or tell me which scene to redo and how?" Wait. Redo only the scenes I name. Once approved, save the URLs as KEY1 to KEY5.
## Phase 3. Clips and narration (run in parallel)
Clips: use fal-ai/kling-video/v2.6/pro/image-to-video via submit_job for all five at once, then check_job until each is done. For every clip: start_image_url = the keyframe, aspect_ratio 9:16, generate_audio false, negative_prompt "blur, distortion, extra limbs, extra characters, changing caption text, morphing faces, painterly texture, photorealism, colour shifts, camera shake". The caption must stay fixed and legible for the whole clip. Sitcom staging: mostly locked camera, characters act with their whole bodies, designs must not change.
Clip 1 (KEY1, duration 10): "Doug jabs the remote at Byte three times, one thought bubble popping up with each jab: dentist chair, pizza box, spreadsheet. Byte's ring flickers orange and it shrugs with its tiny arms. Doug slumps back into the couch. Ellie rises slowly from behind the couch and raises one eyebrow. Locked camera. Flat 2D sitcom animation, caption stays fixed."
Clip 2 (KEY2, duration 5): "Ellie taps the tablet and the three sockets on the TV light up one at a time, each a different shape. Doug leans forward and squints. Byte's ring pulses blue. Locked camera. Flat 2D sitcom animation, caption stays fixed."
Clip 3 (KEY3, duration 10): "Doug jams the shaver plug at the socket, it does not fit, he digs frantically in the suitcase, adapters fly over his shoulder, he stacks four adapters together, plugs in, and the bedside lamp pops with a puff of smoke. He freezes, sweating. Locked camera. Flat 2D sitcom animation, caption stays fixed."
Clip 4 (KEY4, duration 5): "Ellie clicks the adapter onto Byte, the ring snaps to green, the three cables light up in sequence: the pizza box slides in through the door, the calendar ticks, the spreadsheet fills itself. Doug's arms shoot up in celebration. Locked camera. Flat 2D sitcom animation, caption stays fixed."
Clip 5 (KEY5, duration 5): "Doug squeezes the pizza box and sighs happily. Ellie holds her finger up, then points it down at the caption. Byte's arm points down twice. Slight push-in on Ellie. Flat 2D sitcom animation, caption stays fixed."
Narration, while the clips render: use fal-ai/elevenlabs/tts/eleven-v3. The narrator is Ellie: pick a bright, dry, precocious young female voice if the schema exposes a voice list; otherwise use the default and tell me. Deadpan delivery, no hype. Generate five separate files:
Line 1: "My dad wants his AI to book the dentist, order pizza, and fix the spreadsheet. It can talk about all three. It can do none of them."
Line 2: "Every app has its own plug shape. That is an API."
Line 3: "So to connect your AI, a developer builds one adapter per app. Remember Dad in Europe? One adapter per appliance. Every single time."
Line 4: "MCP is the universal adapter. One plug, every tool."
Line 5: "MCP does not replace APIs. It sits on top of them. Comment MCP for the plain-English guide."
If a line runs longer than its clip (10, 5, 10, 5, 5 seconds), regenerate it with a faster delivery instruction rather than cutting words.
Review each clip yourself first: first frame, middle, last frame, caption, and whether any character's design drifted. If it did, or the caption warped, regenerate that clip once with one extra positive lock added to the prompt. Then show me what you have.
GATE 3. Show the five clips in order, each with a one-line read, and play me line 2 of the narration as the voice sample. Ask: "Approve the clips and the voice, or tell me which clip to redo and how?" Wait. Redo only what I name. A clip that fails twice stays in with a note in the log; do not attempt a third take unless I ask. Once approved, save CLIP1 to CLIP5 and VO1 to VO5.
## Phase 4. Assemble and final cut
1. For each scene, call fal-ai/ffmpeg-api/merge-audio-video with video_url CLIPn and audio_url VOn, start_offset 0. Save as MIX1 to MIX5.
2. Call fal-ai/ffmpeg-api/merge-videos with video_urls [MIX1 to MIX5] in order, target_fps 30, resolution matching clip 1.
3. Check that narration plays on every scene. If the stitched file lost its audio, merge the five clips first, generate the five lines as one narration file, and merge that onto the stitched video.
GATE 4. Play me the finished Reel here in the chat. Give me your honest read: total cost so far, which scenes are strong, which are weak and why, and whether captions survived well enough to post or should be redone in the edit. Ask: "Ship it, or what should change?" Wait.
## Delivery (after final approval)
Put everything in this project folder under reel2/. If you cannot download a file, write its URL in the log instead.
- reel2/reel-final.mp4
- reel2/references/ (cast sheet, room plate), reel2/keyframes/ and reel2/clips/ (all takes including rejected ones, named by scene)
- reel2/PRODUCTION-LOG.md: one line per generation with endpoint, request id, cost, verdict (keeper, retry, weak), and what you would fix next time
- reel2/CAPTION.md: a plain Instagram caption in two short sentences ending with "Comment MCP for the plain-English guide." Only hashtag: #AutomateWithMarc. No hype words.