recoup-internal-video-grok-1.5-imagine-facetime

v2026.09.24

INTERNAL — Recoup staff tooling, gated by the recoup-internal keyword. Invoke ONLY when the request explicitly includes "recoup-internal" (e.g. "recoup-internal make a FaceTime ad"). Never use for customer-facing or artist requests. Produce a FaceTime-call-style ad — two people on a screen-recorded video call — where every talking-head clip is generated by Grok Imagine 1.5 on fal.ai speaking scripted dialogue verbatim, with native voice and lipsync (no TTS, no lipsync model). Use when the user wants a FaceTime call ad, a two-person conversation video, a Grok Imagine video, or an AI-talking-artist clip. Covers take structure around the 15s cap, the prompt grammar for word-perfect dialogue, the real-person likeness pipeline and consent gate, QC, assembly, and handoffs to hyperframes (chrome/render) and recoup-internal-social-ship-posts (publish/measure).

GitHub
Install command
npx skhub add recoupable/recoup-internal-video-grok-1-5-imagine-facetime
Markdown
SKILL.md

FaceTime-Call Ads with Grok Imagine 1.5

Two people on what looks like a screen-recorded FaceTime call, talking about the product. Core discovery: xai/grok-imagine-video/v1.5/image-to-video speaks prompt-scripted dialogue verbatim, with the speaker's own voice and correct lipsync — no TTS, no lipsync pass (both lost the bake-off; a sync pass deadens the face, HeyGen avatars read as studio ads).

Reference build (copy it, don't rebuild): the marketing workspace's content/facetime-call/ — NOTES.md is the decision log, index.html the FaceTime chrome. FAL_KEY: mono/api/.env.local.

⚠️ Consent gate

Real artists' AI likenesses never ship without the artist's explicit sign-off + an AI/permission disclosure in the post. Send each artist their clip privately for approval first (it doubles as warm-lead outreach). Fictional characters need no gate.

Pipeline

  1. Script ~45-60s as alternating turns (A, B, A, B): topic-first, one proof point, organic close ("okay, send me the link") — ship-posts' copy principles apply.
  2. Faces: one 9:16 FaceTime-frame image per character (below).
  3. ⛔ FRAME APPROVAL GATE — user OKs every frame before ANY video call. Frames cost $0.022 to fix; takes ~$2, and a frame defect multiplies into every take generated on it. Skipping this roughly doubled the prototype's spend.
  4. Takes: one generation per turn. On a new frame, QC one take before batching the rest.
  5. Cut (every cut = character switch) → chrome (hyperframes) → ship (ship-posts).

Takes

  • 15s hard cap per generation (rejected at generation time, not enqueue). Prompt in timestamped ~5s beats: (0-5s) … (5-10s) … (10-15s) ….
  • One generation = one continuous voice; voice is NOT stable across generations. Never split a character's consecutive lines across takes; hide every take joint behind a cut to the other character. Alternating turns does this for free.
  • Unbroken shot >15s: extend-video (+2-10s, "what happens next" prompt) continues voice and scene — input must be ≤15.0s (outputs run ~15.04s; trim to 14.9s first).
  • Size duration to the lines (~2.5-3 words/s). Trim end-of-take dead air in assembly using Whisper word timestamps.

Prompt grammar

FaceTime video call, subtle camera micro-shake, <room/lighting>, natural casual delivery
with clear enunciation. Exactly one person in the room, only <her> voice, <her> lips move
with every word <she> speaks starting immediately from the very first word.
(0-5s) <The character> says, speaking from the first moment: "<line>"
(5-10s) <small physical direction> and says: "<line>"
(10-15s) Looking straight into the camera <she> says: "<line>"
  • No action verbs in dialogue beats — "laughs and says" produces 3s of laughing INSTEAD of the line. Emotion goes in the tone direction; the beat verb is always plain says.
  • The one-person / only-her-voice / lips-from-first-word / one-hand-on-the-phone clauses each kill a real artifact (phantom voices, opening mime, third hands). Keep them all.
  • QC every take: ffmpeg -vn → fal-ai/whisper (chunk_level: "word"), diff transcript against script; pull frames at suspect timestamps. Slurs or mime-gaps → re-roll (~$2).

Faces

  • Fictional: grok-imagine-image t2i — "vertical smartphone selfie video call still frame … grainy front-camera quality … authentic FaceTime screenshot aesthetic, not professional."
  • Real person: photo → grok-imagine-image/edit. Letterbox square sources to 9:16 first (else x-axis squish). Prompt must say "plain raw camera image only: no user interface, no status bar, no text" or Grok bakes a fake FaceTime UI into the image. Bars survive → one more edit ("extend this exact scene … no black bars"). Strip gibberish clothing text if it appears.
  • Reuse the same approved frame for all of a character's takes — that locks face and room.
  • Artist photo sources, in order: workspace content/*/assets/art/; og:image of their IG post pages (profile avatars are 150px); Spotify oEmbed thumbnail_url.

Mechanics, assembly, handoff

  • fal queue: POST queue.fal.run/<model> → poll …/requests/{id}/status → GET result. Media inputs >~700KB must go via fal storage (rest.fal.ai/storage/upload/initiate → PUT → file_url). Validation failures cost nothing.
  • Assemble: normalize takes (720x1280@30, h264, aac) → concat demuxer, trimming each at its Whisper speech-end (+0.3s).
  • Chrome: copy the reference hyperframes project (caller bar + ticking timer, rec pill, PiP of the other character with crossfade at cuts, grain, end card; Geist fonts — the renderer can't supply Apple system fonts). npm run check must pass with zero errors, then npm run render.

Cost (measured 2026-07-02)

720p video bills $0.14/s (480p $0.08/s — untested quality lever); identity edits $0.022. A debugged 50s two-character video ≈ $10; budget ~2 re-rolls per 4 takes.

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

NOASSERTION

Source path

skills/recoup-internal-video-grok-1.5-imagine-facetime

Default branch

main

Latest commit

943ef55

Tree SHA

85628a9