For the target video, at 0.00 seconds into the target video, Picture 1 (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] At 00:00.000, the video begins exactly from the provided image. A stylized 3D anime girl with long brown twin-tail hair sits at a restaurant table, resting her cheek against one hand. Preserve her exact hairstyle, facial design, black-and-yellow outfit, jewelry, restaurant background, warm lighting, camera angle, and composition from the reference image. She looks directly toward the camera with a soft affectionate expression. Over the first second, her expression gradually becomes more flustered: her cheeks develop a visible blush, her eyes become slightly shy, and she briefly looks away before looking back toward the camera. [Shot 2] At 00:01.000, she lowers her gaze slightly, smiles nervously, and gently touches her cheek with her fingertips. Her blush becomes stronger. She speaks softly and affectionately, with a slightly embarrassed delivery, <d>[English] Aww, Narukami-kun, I can't stop thinking about you.</d> Keep her face clearly visible throughout the entire dialogue for accurate lip synchronization. Her mouth movements should follow every word naturally. [Shot 3] At 00:03.800, she finishes the sentence and becomes visibly more flustered. She gives a shy little smile, her cheeks remain flushed, and she briefly looks away before glancing back toward the camera. Hold on her embarrassed, affectionate expression for the final moment. Camera movement & shots: preserve the starting composition from the reference image. Use only a very subtle slow push-in toward her face during the dialogue. No camera cuts, no large camera movement, no change of location, and no wide shot. Keep the framing primarily as a medium close-up focused on her face and upper body. Visual style: polished stylized 3D anime character rendering with realistic materials, expressive facial animation, detailed hair strands, soft warm restaurant lighting, natural eye movement, subtle blinking, and delicate facial micro-expressions. Preserve the character's identity and clothing exactly from the first frame. overall_soundscape: Warm restaurant ambience, faint conversation in the background, subtle room tone, soft clothing movement, and gentle breathing. Her voice is clear, intimate, and slightly flustered. No overlapping speech. non_diegetic_music: Very soft romantic background music with gentle piano, kept subtle beneath the dialogue.
Parameters used to generate this content
The node graph used to generate this content — pan and zoom to explore, or download it to run in ComfyUI.
AI models used to generate this content