prompt.txt

subject_definitions:
<Subject 1> is the young swordswoman in <Picture 1>, with long wavy black hair, wearing a sleeveless long purple vest with gold piping over a white short-sleeved top, a yellow-orange sash, purple wrist guards, white trousers and purple boots, carrying a pinkish-red long whip.
<Picture 1>: The reference sheet shows the SAME single person from three angles (front, side, back) — the three figures are ONE person, not three different people.
<Picture 2> is the environment and lighting reference for [Shot 1].

summary:
[reference generation] A 5-second vertical 9:16 cinematic shot for a Chinese Paladin music video. <Subject 1> supply the identities of the characters, and the remaining reference images supply the environment and lighting.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - face, hairstyle and costume kept exactly as in <Picture 1>.
<Picture 1> ([Shot 1] identity anchor): partially_preserved - the three views are the same single character; identity and costume only, never a group of people.
<Picture 2> ([Shot 1] environment): partially_preserved - used for setting, architecture and lighting mood only.

detailed_description:
[Shot 1] An empty rain-wet Jiangnan street at dawn, wooden towers with red lanterns on both sides, thick mist, a distant gate. The swordswoman walks alone along the flagstones away from camera, seen from behind; her sash and hair drift in the cold morning air while mist rolls slowly. Slow tracking shot following her. Small motion. The shot lasts 5 seconds, vertical 9:16. The referenced characters appear from 0.00 seconds onward. Style: semi-realistic Chinese fantasy cinematic CGI, volumetric light, glowing mist, soft bloom, cobalt blue and teal palette with honey-gold accents. No text, no watermark, no extra people.

overall_soundscape:
Soft wind, faint dripping water, distant temple bell.

non_diegetic_music:
A soft restrained orchestral pad in a minor key, very low in the mix, leaving room for the song.
下载此文件