FaceTime-Call Ads with Grok Imagine 1.5
Two people on what looks like a screen-recorded FaceTime call, talking about the product.
Core discovery: xai/grok-imagine-video/v1.5/image-to-video speaks prompt-scripted dialogue
verbatim, with the speaker's own voice and correct lipsync — no TTS, no lipsync pass (both
lost the bake-off; a sync pass deadens the face, HeyGen avatars read as studio ads).
Reference build (copy it, don't rebuild): the marketing workspace's content/facetime-call/
— NOTES.md is the decision log, index.html the FaceTime chrome. FAL_KEY: mono/api/.env.local.
⚠️ Consent gate
Real artists' AI likenesses never ship without the artist's explicit sign-off + an AI/permission disclosure in the post. Send each artist their clip privately for approval first (it doubles as warm-lead outreach). Fictional characters need no gate.
Pipeline
- Script ~45-60s as alternating turns (A, B, A, B): topic-first, one proof point, organic close ("okay, send me the link") — ship-posts' copy principles apply.
- Faces: one 9:16 FaceTime-frame image per character (below).
- ⛔ FRAME APPROVAL GATE — user OKs every frame before ANY video call. Frames cost $0.022 to fix; takes ~$2, and a frame defect multiplies into every take generated on it. Skipping this roughly doubled the prototype's spend.
- Takes: one generation per turn. On a new frame, QC one take before batching the rest.
- Cut (every cut = character switch) → chrome (hyperframes) → ship (ship-posts).
Takes
- 15s hard cap per generation (rejected at generation time, not enqueue). Prompt in
timestamped ~5s beats:
(0-5s) … (5-10s) … (10-15s) …. - One generation = one continuous voice; voice is NOT stable across generations. Never split a character's consecutive lines across takes; hide every take joint behind a cut to the other character. Alternating turns does this for free.
- Unbroken shot >15s:
extend-video(+2-10s, "what happens next" prompt) continues voice and scene — input must be ≤15.0s (outputs run ~15.04s; trim to 14.9s first). - Size duration to the lines (~2.5-3 words/s). Trim end-of-take dead air in assembly using Whisper word timestamps.
Prompt grammar
FaceTime video call, subtle camera micro-shake, <room/lighting>, natural casual delivery
with clear enunciation. Exactly one person in the room, only <her> voice, <her> lips move
with every word <she> speaks starting immediately from the very first word.
(0-5s) <The character> says, speaking from the first moment: "<line>"
(5-10s) <small physical direction> and says: "<line>"
(10-15s) Looking straight into the camera <she> says: "<line>"
- No action verbs in dialogue beats — "laughs and says" produces 3s of laughing INSTEAD of
the line. Emotion goes in the tone direction; the beat verb is always plain
says. - The one-person / only-her-voice / lips-from-first-word / one-hand-on-the-phone clauses each kill a real artifact (phantom voices, opening mime, third hands). Keep them all.
- QC every take:
ffmpeg -vn→fal-ai/whisper(chunk_level: "word"), diff transcript against script; pull frames at suspect timestamps. Slurs or mime-gaps → re-roll (~$2).
Faces
- Fictional:
grok-imagine-imaget2i — "vertical smartphone selfie video call still frame … grainy front-camera quality … authentic FaceTime screenshot aesthetic, not professional." - Real person: photo →
grok-imagine-image/edit. Letterbox square sources to 9:16 first (else x-axis squish). Prompt must say "plain raw camera image only: no user interface, no status bar, no text" or Grok bakes a fake FaceTime UI into the image. Bars survive → one more edit ("extend this exact scene … no black bars"). Strip gibberish clothing text if it appears. - Reuse the same approved frame for all of a character's takes — that locks face and room.
- Artist photo sources, in order: workspace
content/*/assets/art/; og:image of their IG post pages (profile avatars are 150px); Spotify oEmbedthumbnail_url.
Mechanics, assembly, handoff
- fal queue:
POST queue.fal.run/<model>→ poll…/requests/{id}/status→ GET result. Media inputs >~700KB must go via fal storage (rest.fal.ai/storage/upload/initiate→ PUT →file_url). Validation failures cost nothing. - Assemble: normalize takes (720x1280@30, h264, aac) → concat demuxer, trimming each at its Whisper speech-end (+0.3s).
- Chrome: copy the reference hyperframes project (caller bar + ticking timer, rec pill, PiP of
the other character with crossfade at cuts, grain, end card; Geist fonts — the renderer can't
supply Apple system fonts).
npm run checkmust pass with zero errors, thennpm run render.
Cost (measured 2026-07-02)
720p video bills $0.14/s (480p $0.08/s — untested quality lever); identity edits $0.022. A debugged 50s two-character video ≈ $10; budget ~2 re-rolls per 4 takes.