「いい感じにして」ボタンは存在しない
まずは、きらきらした幻想をそっと終わらせておきます。プロンプトを入れたら完成作品が出てくる、そんな魔法のボタンはありません。もしあるなら、少なくとも私たちは招待されていません。
いま見た映像は1分37秒。制作には8時間34分、費用は$85.45かかりました。キャラクター設計、衣装、シーン作り、スタジオカット、音声、リップシンク、編集まで含めた金額です。ちなみに編集とは、楽観がいちばん丁寧に打ち砕かれる場所でもあります。
この記事では、実際の流れをすべて見せます。カットごとのプロンプト、失敗、修正、費用まで。各工程にはそれぞれ独自の壊れ方があります。面倒ですが、それこそがこの制作の本体です。
Milaというキャラクター
Milaは実在の人物ではなく、ゼロから作った架空のキャラクターです。27歳、身長は約178cm。肩より下まであるシャンパンブロンドのセンターパート、グレーブルーのアーモンドアイ、卵型に近いハート形の輪郭、しっかりしたブロンドの眉、右目の下に小さなほくろがあります。
目指したのは、知的で落ち着いていて、観察力があり、現代的な人物像。いかにもな“映えるインフルエンサー顔”にはしたくありませんでした。あの方向に寄せると、照明だけが上手で個性はホテルのロビーくらい薄い、なめらかな同じ顔に溶けていきがちです。
設定はシンプルです。Milaは2030年のファッションエディターで、自分のルックをスタイリングして投稿している。特徴を細かく決めるのは重要です。どこにでもいそうな美人顔は、カットをまたいで同一人物か確認しにくいからです。そばかす、眉、ほくろのような要素が、見る側にもモデル側にも「これはまだ彼女だ」と示す手がかりになります。
Step 1 — リファレンスシート
ポートレート1枚では、まだキャラクターとは言えません。必要なのはリファレンスシートです。今回は、正面の基本ポートレート、30°の振り向き、90°の横顔、太もも中ほどまでのショット、全身、30°の全身ショットの6枚を作りました。
コツはとても単純です。良い画像ができたら、その生成済み画像を次のリファレンスとして戻すこと。そして、顔と体の構造を変えないよう短く指示すること。長いプロンプトを延々と書き足すより、こちらのほうがうまくいきました。長文化しすぎると、片頭痛持ちの似顔絵捜査官みたいな文章になります。
意図的に、ほぼひとつのエンジン内で進めました。そうすることで同一性のズレを減らせますし、5つのツールを同時に責める疲れた旅行中の親のようにならず、そのエンジンがどこで崩れるのかを学べます。






Milaのリファレンスシート。以降のカットはすべてここから生成。
シートを作った5つのプロンプト
各カットには実際のプロンプトを添え、なぜその指示を入れたのかも書いています。真似すべきなのは型です。同じ身体的特徴を毎回繰り返し、変える要素をひとつだけ指定し、承認済み画像を順にリファレンスとして加えていく。そこが要点です。
どのプロンプトでも同じ顔の説明を入れ直し、新しい角度をひとつだけ足しています。主役はアップロードしたポートレート。文章は「どちらを向くか」を伝える補助です。

プロンプトを表示
Use the uploaded portrait of Mila as the primary identity reference. Create a RAW photorealistic studio portrait of the same real-looking adult woman. Keep her exact face: fair skin with natural texture, oval heart-shaped face, defined cheekbones, soft refined jawline, large expressive grey-blue almond eyes, straight narrow nose, natural medium lips, distinctive full blonde eyebrows with a soft lifted arch, and a small beauty mark below her right eye. Champagne-blonde hair below the shoulders, centre part, soft waves. Head-and-shoulders view, face turned 30 degrees to her left, eyes looking into the camera, head upright and level. Calm, intelligent, slightly reserved expression, closed lips. Minimal makeup, graphite-grey crew-neck top, no jewellery. Light-grey seamless studio background, soft frontal light, 85mm lens look, realistic colour, unretouched human photo. No illustration, CGI, doll face, glamour filter, text or watermark.
厳密な横顔は、キャラクターを保つのが最も難しい角度です。眉とほくろを明記することで、なんとか同一人物としてつなぎ止めます。

プロンプトを表示
Use the uploaded portrait of Mila as the primary identity reference. Create a RAW photorealistic studio image of the same real-looking adult woman. She is tall and slender, with fair skin, natural texture, visible pores, an oval heart-shaped face, defined cheekbones, a soft refined jawline, large expressive grey-blue almond eyes, a straight narrow nose, natural medium lips, and a small beauty mark below her right eye. Her eyebrows are distinctive: light blonde, full, textured, elegant, with a soft lifted arch. Her hair is champagne blonde, below the shoulders, centre part, soft waves. Show a strict 90-degree right-facing side profile, head level, calm neutral expression, closed lips, hair tucked behind the visible ear. Graphite-grey crew-neck top, no jewellery. Light-grey seamless studio background, soft diffused light, true human photo. No cartoon, CGI, doll face, beauty filter, text or watermark.
「顔を作り直すな、解釈し直すな、美化するな、差し替えるな」は小言のように見えますが、この小言がズレを止めます。

プロンプトを表示
Use the uploaded portrait as the only identity reference. Create a RAW photorealistic studio photo of the exact same woman, framed from head to mid-thigh. Do not redesign, reinterpret, beautify or replace her face. Preserve her identity exactly: same facial geometry, proportions, eyes, eyebrows, nose, lips, jawline, skin tone, hairline, hairstyle and beauty mark. She is a clearly adult woman, tall and slim, with a defined waist and balanced feminine curves. She faces the camera in a relaxed neutral pose, arms slightly away from her torso. She wears a fitted graphite-grey bodysuit. Light-grey seamless background, soft even light, 50mm lens, realistic skin texture, true human photography. Keep the face sharp, detailed and clearly recognizable. No new face, face variation, cartoon, CGI, doll-like skin, distorted anatomy, text or watermark.
ここでリファレンスは2枚になります。ポートレートと、承認済みの太もも中ほどのショット。使える角度を戻すほど、彼女は安定していきます。

プロンプトを表示
Use the uploaded portrait and approved mid-thigh image as identity references. Create a RAW photorealistic full-body studio photo of the exact same woman. Do not redesign or replace her face. Preserve her identity exactly: same facial geometry, proportions, eyes, eyebrows, nose, lips, jawline, skin tone, hairline, hairstyle and beauty mark. She is a clearly adult woman, about 178 cm tall, with a slim elegant figure, defined waist, softly rounded hips, balanced feminine curves and long legs, suitable for dress fittings. She stands front-facing in a relaxed neutral pose, arms naturally at her sides, full body visible from head to bare feet. Fitted graphite-grey bodysuit, matte black leggings. Light-grey studio background, soft even light, 50mm lens, true human photo. No new face, face variation, cartoon, CGI, doll-like skin, distorted anatomy, cropped feet, text or watermark.
この段階まで来ると、プロンプトは短くできます。彼女はもう存在しているので、別角度の撮影を入れるだけです。

プロンプトを表示
Using the uploaded portrait as the face reference, create a photorealistic full-body studio photo of the exact same woman. Tall, slim figure with a defined waist and natural feminine curves. Body turned 30 degrees, head toward the camera, relaxed pose, full body and feet visible. Preserve her face and hair exactly. Simple fitted grey outfit, soft studio light, plain grey background, no cartoon, CGI or distorted anatomy.
1枚のポートレートを完整なリファレンスシートに広げた5つのプロンプト。
Step 2 — ワードローブを作る
ルックを決めたあと、画像モデルにドレス、バッグ、靴、サングラス、イヤリング、ブレスレットをそれぞれ単体の商品写真として切り出してもらいました。
このひと手間には価値があります。アイテムが単独のリファレンス画像として存在していれば、別のシーンに戻したときも同じものとして残りやすくなります。実際のルックブック作りと同じです。違うのは、ドアをふさぐ洋服ラックがないことくらいです。






各アイテムを単体の商品カット化。何度でも安定して再利用するため。
Step 3 — 話すカットを作る
トークシーンの構成は、正直かなり継ぎはぎです。静止画から話すクリップを作り、音声はElevenLabsで生成し、リップシンクはSync.soで行いました。
現時点では、音声とリップシンクはVynzoの外で処理しています。Kling 2.6ならVynzo内でネイティブにリップシンクできますし、GPT Image 2もまもなくVynzoに追加予定です。
ツール名より大事なのは実務上のコツです。リップシンクの各テイクは、ジェスチャーや表情が違う別々の元クリップから始めること。同じ元クリップを使い回すと、Milaが感じのいいカスタマーサービスのループに閉じ込められたような、ロボットっぽい映像になります。
トークシーン。言葉そのものより、表情と身ぶりが効きます。
プロンプトを表示
Use the uploaded reference image as the exact identity of the character. Preserve her face, hairstyle, skin tone, proportions, clothing and accessories. Do not redesign or stylize her. Medium close-up, eye-level camera. She looks into the lens and speaks naturally: "[DIALOGUE]". Precise lip sync. Expressions follow the speech: subtle eyebrow movement, realistic blinking, small smiles, brief pauses and natural breathing. She uses restrained hand gestures, small head nods and slight posture shifts. Movements are smooth, relaxed and conversational, never theatrical or repetitive. Keep hands anatomically correct. Negative prompt: identity drift, facial warping, frozen expressions, excessive blinking, random gestures, distorted hands, lip-sync errors, flicker, camera shake, cuts or background changes.
Step 4 — 自宅シーン
自宅シーンは、紙の上ではシンプルです。Milaが1クリップにつき1つの商品を紹介する。リファレンスセットから静止画を生成し、動かし、再生成し、いちばん良いテイクを残す。この繰り返しです。
各クリップには、同一性を保つこと、動きを自然で控えめにすることを同じように指示しました。本当の学びは各クリップ下のメモにあります。最初のバージョンこそ、モデルが新しい問題を丁寧に見せてくれる場所だからです。
全身で一回転させると、キャラクターの本当の強度が出ます。360°を通して本人でいられるか。このテイクはうまくいきましたが、プロンプトの大半が「してほしいこと」ではなく「しないでほしいこと」に費やされている点に注目です。
プロンプトを表示
Use the provided frame as a strict reference. Vertical 9:16, full-body, static camera. Keep the same woman, identity, face, body proportions, hairstyle, dress, bare feet, room, lighting and framing unchanged throughout. She starts already facing the camera, looking directly into the lens with a soft natural smile. From this exact starting pose, she gently holds the sides of the dress with both hands and makes one smooth natural full turn around herself at normal real-time speed, with small realistic steps. The movement should feel elegant and physically natural, with subtle fabric motion and slight hair movement. After the turn, she finishes facing the camera again. No slow motion, no strange grimaces, no exaggerated expressions, no face distortion, no identity drift, no hand glitches, no extra fingers, no foot deformation, no added heels or footwear, no camera movement, no flicker
いちばん好きな失敗例です。モデルがバッグを自分の腕にまっすぐ貫通させ続けました。幽霊のように。何度も試したあと、一部だけ使いました。手と物体の接触は、いまの動画モデルがまだ揺らぎやすいポイントです。
プロンプトを表示
A woman naturally presents a light-colored leather handbag in a minimalist room. She gently turns her upper body slightly toward the camera, carefully lifts the bag by its handle, and brings it a little closer to the lens to showcase its shape, leather texture, stitching, and hardware. She then subtly adjusts her grip and slowly rotates the handbag to reveal its side profile. Her expression remains calm and confident, with a soft, natural smile. Her hair and clothing move slightly with her body. The camera performs a slow, smooth push-in while keeping the handbag as the main focal point. Realistic premium fashion commercial, soft natural daylight, elegant and controlled movement, high detail, stable composition, smooth natural motion
見えないスマホと、意図的な手持ちの揺れを指定することで、本物のセルフィーらしく見えます。頭を回すとサングラスのつるが消えがちでした。そこは再生成しながら、その細部を名指しするしかありません。
プロンプトを表示
Create a photorealistic vertical 9:16 front-camera selfie video from the reference image. Preserve her identity, hair, sunglasses, mint athletic outfit, room, lighting, and framing. The phone is in her hand but never visible, so the camera should have subtle natural handheld motion throughout: tiny shakes, micro-sways, and slight distance shifts. She looks at the screen, slowly turns her head to one side and freezes, holding still for a moment to show the sunglasses. Then she slowly turns her head to the other side and freezes again. After that, she gently brings the camera closer to her face for a close-up of the sunglasses and holds it there briefly. Add natural blinking, breathing, tiny facial movements, and slight hair motion. No speaking, cuts, sudden motion, warped features, or background changes.
ファーストフレーム/ラストフレーム方式です。誰も教えてくれないルールは、両方のフレームでカメラ位置をそろえること。そうしないと、カットの中で物体が流れていきます。
プロンプトを表示
Use the first frame as the starting pose and the last frame as the ending pose. In the same bright minimalist room, the woman sits on the sofa and smoothly transitions from reaching toward the open shoebox on the floor to lifting one iridescent heel out of the box and examining it in her hands. At the start, both shoes are clearly inside the shoebox among the tissue paper. During the action, she leans forward, reaches into the box, takes only one shoe by the ankle strap, lifts it up, and supports it with her other hand while looking at it with soft curiosity and a subtle pleased smile. By the end, one shoe is in her hands and the second shoe remains inside the box. Maintain consistent appearance, outfit, room layout, shoebox position, and shoe design. Soft daylight, static camera, realistic motion
2回目の試行です。1回目では、手首が落ち着いた様子で360°回転しました。人間の手首にはできません。
プロンプトを表示
A cinematic product showcase of a futuristic silver iridescent high-heel sandal being elegantly held in one hand inside a bright, minimalist living room. The camera performs a slow, smooth push-in combined with a subtle left-to-right arc, creating gentle parallax between the shoe and the softly blurred background. The hand naturally rotates the shoe a few degrees to reveal the shimmering holographic panels, metallic finish, sculptural transparent heel, and ankle strap. Soft daylight from the window creates realistic reflections and rainbow highlights that glide across the surface as the camera moves. The background remains calm and out of focus, emphasizing the product. Premium luxury fashion commercial aesthetic, ultra-realistic materials, clean composition, shallow depth of field, stabilized camera, natural motion only, no abrupt movements, no object deformation, no flickering, no warping, no extra fingers or artifacts.
自宅シーン。1クリップにつき1商品。
ファーストフレーム/ラストフレームのコツ
靴を箱から取り出すように、始まりと終わりがはっきりした動きでは、最初のフレームと最後のフレームを両方モデルに渡します。どこから動き始め、どこに着地すべきかを知らせる必要があるからです。
見落とされがちなのはカメラ位置です。最初と最後のフレームは同じカメラ位置でなければいけません。ずれていると、モデルはカメラと物体を同時に動かそうとして、物が滑ったり変形したりします。そうして靴は呪われます。


最初のフレームと最後のフレーム。カメラ位置が同じでないと、物体が滑ります。
Step 5 — スタジオとランウェイ
スタジオシーンでは、Milaに完成したルックを着せ、サングラスあり・なしを作り、その後メイクを加えました。ここでは同一性に特に注意が必要です。メイクは顔を変えます。キャラクターを見分けるための小さな印や微細な不完全さを隠してしまうことがあるからです。
撮影現場の舞台裏風カットは、思ったより早く形になり、説得力もありました。ストロボのフラッシュ、スタジオファン、画面外のフォトグラファー。ランウェイのカットは、映像全体を成立させる要です。衣装、キャラクター、動きがひとつの明快な瞬間になります。
Nano Banana 2は、アクセサリーを常に同じ形で保つのが得意とは言えませんでした。この用途ではGPT Image 2のほうが安定しました。アクセサリーをアップロードして、カットに編集で入れる。シンプルで、魔法ではありません。どうやらこれが今回のテーマになってきました。
全体の説得力を背負う1カットです。ドレス、サングラス、イヤリング、ブレスレット、ヒール、バッグまで明示しています。名前を書かなかったものは、モデルが静かにデザインし直すかもしれないからです。
プロンプトを表示
The same blonde model from the reference walks confidently down a luxury fashion runway with a professional catwalk stride, maintaining the exact same face, hairstyle, white dress, sunglasses, earrings, bracelet, heels, and cream handbag. She reaches the end of the runway, gracefully stops, and performs three elegant high-fashion poses, subtly shifting her weight, rotating her body, lifting her chin, and naturally presenting the handbag. Camera flashes illuminate her as photographers capture every pose. She executes a smooth runway pivot, then confidently walks back. The dress flows naturally, the handbag swings realistically, and every movement is poised and refined. Cinematic fashion film, glossy runway, soft spotlights, shallow depth of field, ultra-realistic, 9:16, 10 seconds, 4K, 24 fps, preserve the exact appearance from the reference, no outfit or face changes.
スタジオシーン。完成ルックからランウェイへ。
スタジオ静止画
スタジオ静止画は、ワードローブ用に作った商品カットから完成ルックを組み、スタジオ撮影の形にしたものです。ここで商品カット作りの投資が効いてきます。


スタジオ静止画。ワードローブカットから組み上げた完成ルック。
壊れたもの全部
多くの記事が省きがちな部分です。おそらく、完成動画ほど華やかではなく、人前でスーツケースが爆発したことを認めるような話だからでしょう。でも失敗は例外ではありません。これが仕事です。
写真家が撮影枚数を見込むように、再生成の予算を見込んでください。いちばん実用的だった発見は、Seedanceがよくスローモーション気味に出力することです。ネガティブプロンプトでは直らなかったので、編集で35〜40%ほど速度を上げました。
| 壊れたこと | なぜ問題か | 実際に効いた対処 |
|---|---|---|
| バッグが腕を貫通した | 手と物体の接触は、現在の動画モデルで最も弱い部分 | 何度も再生成し、使える部分だけ採用 |
| 振り向きの途中でサングラスのつるが消えた | 頭部が回転すると、小さく硬いディテールが消えやすい | そのディテールを明記して再生成 |
| 手が360°回転した | 気づいて指摘しない限り、解剖学は守られない | 1回再生成 |
| すべてがスローモーションになった | Seedanceが短い動きをクリップ尺いっぱいに引き伸ばす | ネガティブプロンプトは効かなかったため、編集で35〜40%高速化 |
| メイクペンシルのシーン | どう言い換えても、モデルが実行できない動作がある | カット |
| 終盤で背中が不自然に反った | 最後のフレームはポーズが崩れやすい | 編集でトリミング |
| カットごとにアクセサリーが変わった | 画像モデルは小物でズレやすい | Nano Banana 2は苦戦。GPT Image 2は安定。アクセサリーをアップロードして編集 |
自分の写真で4つのエンジンを試せる、ひとつのスタジオ。
スタジオを開くかかった費用
完成映像1分37秒に対して、制作費は合計$85.45、作業時間は8時間34分でした。内訳は、画像・動画生成が$55.45、リップシンクが$19、音声が$11です。
では、AIインフルエンサーは安いのでしょうか。モデル、フォトグラファー、スタジオ、ロケ地を使う実写撮影と比べれば、安いです。ただし無料でも一瞬でもありません。私たちはファッションブロガーではなく、これは現時点のツールで何ができるかを示すデモです。テイク数を増やし、より良い編集者を入れれば、さらに良くなります。
| 項目 | 内容 | 費用 |
|---|---|---|
| Seedance 2.0 Pro · Nano Banana 2 · GPT Image 2 | すべての画像と動画クリップ | $55.45 |
| Sync.so | リップシンク(確保分の約20%のみ使用) | $19.00 |
| ElevenLabs | ナレーション | $11.00 |
| 合計 | 完成映像1分37秒 | $85.45 |
| 時間 | 編集込み | 8 h 34 min |
自分の写真で4つのエンジンを試せる、ひとつのスタジオ。
スタジオを開く始める前に伝えたいこと
まず、魔法のボタンはありません。おそらく今後もないでしょう。残念ながら。AIは照明と布の質感を覚えましたが、それでも必要なのは、センスと忍耐力を持ち、同じカットをもう一度再生成しても個人的な恨みにしない人間です。
次に、明確で面白い脚本がなければ、高価なノイズができあがります。きれいなキャラクターがきれいなカットを歩くだけでは物語になりません。それは請求書つきのスクリーンセーバーです。
最後に、キャラクターこそが商品です。何より先にリファレンスシートへ投資してください。そのうえで、再生成と編集者の予算を見込むこと。スローモーション、不自然な終わり方、成立しないショットは、編集で救うしかない場面が必ずあります。
