subject_definitions: <Subject 1> is the creature shown in <Picture 1>. Preserve its exact visual identity, facial structure, horns, scales, coloration, body proportions, surface details and distinctive characteristics. <Audio 1> is the voice reference for <Subject 1>, used only for vocal identity, timbre, pitch characteristics, rhythm and delivery style. <Subject 2> is the creature shown in <Picture 2>. Preserve its exact visual identity, facial structure, horns, scales, coloration, body proportions, surface details and distinctive characteristics. <Audio 0> is the voice reference for <Subject 2>, used only for vocal identity, timbre, pitch characteristics and singing voice characteristics. summary: A continuous approximately 10-second cinematic audiovisual scene featuring two visually distinct creatures. <Subject 1> delivers a rhythmic spoken rap expressing determination despite having no wings, followed by <Subject 2> singing a short emotional melodic passage about longing to see and experience the open sky. Preserve both visual identities exactly and keep their voices clearly distinct according to their respective audio references. timing: [0.0–5.5s] <Subject 1> performs a rhythmic spoken rap: "No wings, no flight, but keep your head high. The sky is not a dream if you dare to try." The delivery is rhythmic, confident and expressive but clearly spoken rather than sung. [5.5–10.0s] <Subject 2> begins singing a soft emotional melodic passage: "I'm a creature without wings, dreaming of the endless sky. One day I'll see the clouds with my own eyes." The singing is sincere, wistful and melodic, with natural breathing and expressive vocal phrasing. detailed_description: From the beginning, <Subject 1> remains visually consistent with <Picture 1>, with subtle breathing, natural eye movement, restrained head movement and believable mouth, jaw and throat articulation synchronized with the spoken rap. Its voice follows the vocal identity of <Audio 1>. At approximately 5.5 seconds, attention transitions naturally to <Subject 2>, which remains visually consistent with <Picture 2>. <Subject 2> begins singing rather than speaking, using the vocal identity and singing characteristics of <Audio 0>. Its mouth, jaw, throat, breathing and facial movements synchronize naturally with the sung melody. The transition between the two performers is cinematic and continuous, with coherent camera movement, lighting and environmental continuity. Both subjects remain anatomically consistent with their respective reference images throughout the sequence. No wings are present on either subject. Do not introduce wings, flight, anatomical redesign, identity mixing, voice swapping, additional dialogue, additional vocals or unrelated characters. overall_soundscape: Natural environmental ambience appropriate to the reference environment, subtle wind and restrained physical environmental sounds. The spoken rap and singing remain clearly audible and synchronized. No background music. No additional voices.
Parameters used to generate this content
AI models used to generate this content