Text to video generation for dialog between two characters. Run locally using ComfyUI and MiniMax H3.

Dialog is both impressive and tricky to pull off. It happens rather easily but as I was sharing with the potato timelapse, there are a lot of variables with videos and dialog just adds another one. In rendering the clips, I encountered understandable things like the wrong character saying the line and weird things like pronouncing the word ‘donut’ consistently as ‘dough-nuck’.
MiniMax H3 refers to characters as
One trick is to render the video in a very low resolution (0.1MP) to work out any obvious timing or ordering problems. After seeing lots of samples, I also came to refine the characters to what I liked. The different shades of skin and hair gave the world a better sense of itself than brown hair on the colorful characters.
Making these clips (mini-episodes) gave me a good practice run with different problems to solve. The scripts were just little punchlines that I came up with. I really should have started with the bubblegum clip as it was simpler and shorter but instead the headphone one came to mind first with more back-and-forth and potential timing issues.
I was happy with the first rendering of the headphone clip but then after all the other tricks I picked up and a new workflow to use reference images, I tried it again and found it much improved. I really liked the consistency across the clips
You can see the clips below in the order I rendered them (quality increases to the end).
Headphones prompt:#
Style: Simplified 2D character design, expressive cartoon aesthetic, vibrant flat colors, clean vector lines. No 3D rendering, no gradients, flat lighting.
Setting: Minimalist interior of a modern house, 2D flat style.
Characters: S1 (Blue boy, left profile), S2 (Green boy, right profile).
Action Sequence (18 Seconds Total):
0-3s: S1 stands on the left. S2 walks into the frame from the right holding a white box with a prominent orange letter “t”.
3-6s: S1 speaks: <d>[English]Another Temu package?</d> S2 remains silent, looking at the box.
6-10s: S2 speaks: <d>[English]I got hearing cancelling headphones.</d> While speaking, S2 sets the box down and pulls the headphones from the box and puts them on his head. S1 remains silent, watching.
10-13s: S1 speaks with a questioning/confused face: <d>[English]Do you mean noise cancelling headphones?</d> S2 stands completely still with headphones on, mouth closed.
13-15s: S2 suddenly shouts: <d>[English]What?</d> S2 quickly pulls the headphones off his head and holds them in front of him.
15-17s: The Reveal: Stylized red blood drips from S2’s ears. S2 has a pained expression; S1 has a terrified expression.
17-18s: The camera focuses on the interior of the headphones, revealing sharp metal spikes.
Don’t Run prompt:#
Style: Simplified 2D character design, expressive cartoon aesthetic, vibrant flat colors, clean vector lines. No 3D rendering, no gradients, flat lighting.
Setting: Minimalist interior of a modern house, 2D flat style.
Subject_definitions:<S1> is a boy with dynamic green skin, thick flat horizontal dark green fringe (geometric block), right profile.<S2> is a boy with sky blue skin, thick flat horizontal dark blue fringe (geometric block), left profile.
Action Sequence (24 Seconds Total):
0-2s: Green Boy S1 is on the right, putting on sneakers with a determined face and looking at his shoes while muttering </d>.
2-5s: Blue Boy S2 walks in from the left with a curious look and says <d>[English] “Where are you going? </d>
5-8s: Green Boy S1 speaks: <d>[English] “I’m going to a protest against running.” </d>
8-10s: Blue Boy S2 with a shocked, wide-eyed expression: <d>[English] “Somebody’s organizing a rally against exercise?"</d>
10-12s: With a proud, heroic face, S1 speaks: <d>[English] “Yeah! I signed up for the ‘Don’t Run’ at the shopping center."</d>
12-14s: S2 waves a hand and speaks in a sarcastic tone: <d>[English] “I see. Good luck!"</d>. S1 walks off-screen to the right.
14-16s: Transition: The screen cuts to a solid black background with centered white text: “1 hour later”
16-18s: S2 standing on the left. S1 enters from the right, drenched in sweat, chest heaving, looking exhausted.
18-23s: With a haughty expression, Blue Boy S2 says: <d>[English] “So, it was the DONUT Run where you have to eat a dough nut every lap of a 5K around the shopping center?"</d>
23-24s: S1’s eyes roll back and he collapses forward, hitting the floor face-first.
24-24s: S2 looks at the camera with a deadpan expression.
Gum prompt:#
Style: Simplified 2D character design, expressive cartoon aesthetic, vibrant flat colors, clean vector lines. No 3D rendering, no gradients, flat lighting.
Setting: Minimalist exterior of a modern house, 2D flat style.
Subject_definitions:<S1> is a boy with sky blue skin, thick flat horizontal dark blue fringe (geometric block), left profile.<S2> is a boy with dynamic green skin, thick flat horizontal dark green fringe (geometric block), right profile.
Action Sequence (10 Seconds Total):
0-2s: Blue Boy S1 is standing on the sidewalk with a wad of pink bubblegum stuck to his shoe. The gum stretches as he pulls but it holds him in place.
2-4s: Green Boy S2 walks from the left to the right. He stops to look back at S1 when he reaches the right side of the screen.
4-6s: Blue Boy S1 speaks: <d>[English] “Is this your bubble gum?” </d>
6-9s: Green Boy S2 : <d>[English] “Not anymore. That is all yours!"</d>
9-10s: Blue Boy S1 lunges for Green Boy S2 but the gum holds him in place.
Hygiene prompt:#
(I used a new workflow in ComfyUI with MiniMax H3 to provide 2 reference images, which kept the characters consistent. I was even able to introduce a new character.)
Style: Simplified 2D character design, expressive cartoon aesthetic, vibrant flat colors, clean vector lines. No 3D rendering, no gradients, no chibi proportions. Characters must have pre-teen proportions (slightly shorter stature than a teenager, but with natural, balanced head-to-body ratios).
Setting: Interior of a school bus, 2D flat style.
subject_definitions:<S1> is <Picture 2> a teen boy with a lean build and dynamic green skin. He has a thick, solid dark green fringe that is a single geometric block sweeping forward with a distinct outward curve at the corner, creating a slight “wing” shape over his forehead. No hair is not behind or below the ears. He has a whiny/nasally voice. He wears a gray t-shirt with dark red shorts.
<S2> is <Picture 1> a teen boy with a lean build and sky blue skin. He has a thick, solid dark blue fringe that is a single geometric block sweeping forward and curving slightly downward at the edges to cover his forehead. No hair is not behind or below the ears. He has a normal boy voice. He wears a white t-shirt and gray shirt. His legs move normally, it is just raised in the reference image.
<S3> is a teen girl with a lean build and light red skin. She has a solid dark red shoulder-length haircut. She wears a yellow shirt and a brown skirt.
Action Sequence (5 Seconds Total):
0-1s: Blue Boy S1 and Green Boy S2 are sitting next to each other in a school bus seat, facing left.
1-3s: Red Girl S3 silently walks from left to right down the aisle of the bus with a smile, looking to (unseen) friends in seats further back. Green Boy S2 watches her walk by. Blue Boy S1 watches Green Boy S2 with a scowl.
3-4s: Green Boy S1 looks back at Blue Boy S2.
4-6s: Green Boy S1 says: <d>[English] “Do you think Jean likes me?” </d>
6-9s: Blue Boy S2 says: <d>[English] “No! Your hi Jean is basically saying… </d>
9-10s: Green Boy S2 looks forward and down with a sad look.
Gen Z prompt:#
Style: Simplified 2D character design, expressive cartoon aesthetic, vibrant flat colors, clean vector lines. No 3D rendering, no gradients, no chibi proportions. Characters must have pre-teen proportions (slightly shorter stature than a teenager, but with natural, balanced head-to-body ratios).
Setting: Exterior of an American school, 2D flat style.
subject_definitions:<S1> is <Picture 2> a teen boy with a lean build and dynamic green skin. He has a thick, solid dark green fringe that is a single geometric block sweeping forward with a distinct outward curve at the corner, creating a slight “wing” shape over his forehead. No hair is not behind or below the ears. He has a whiny/nasally voice. He wears a gray t-shirt with dark red shorts.
<S2> is <Picture 1> a teen boy with a lean build and sky blue skin. He has a thick, solid dark blue fringe that is a single geometric block sweeping forward and curving slightly downward at the edges to cover his forehead. No hair is not behind or below the ears. He has a normal boy voice. He wears a white t-shirt and gray shirt. His legs move normally, it is just raised in the reference image.
Action Sequence (18 Seconds Total):
0-2s: Blue Boy S2 and Green Boy S1 are standing facing each other in front of a school. S2 is on the left, facing right and S1 is on the right, facing left.
2-7s: Green Boy S1 says: <d>[English] “Rizz for days. You’re delulu. We are so cooked. This school is highkey mid. Bet! No cap.” </d>
7-8s: Blue Boy S2 silently dials 3 numbers on a cell phone and holds it to his ear without saying anything. The phone on the other end rings quietly.
8-10s: Green Boy S1 slowly says: <d>[English] “Sus…” </d>
10-13s: Blue Boy S1 says: <d>[English] “Hello, 911? Yes, my friend is having a stroke.” </d>
13-16s: <d>[English] “Don’t crash out, bro!” </d>
16-18s: A 2d ambulance drives in from the left and blocks the camera. The side of the ambulance says “Gen Z Quarantine”.
Headphones Prompt v2#
(Now that I had the new workflow with reference images, I recreated the headphones clip with the original action sequence and minimal timing changes for the dialog to land properly.)
Style: Simplified 2D character design, expressive cartoon aesthetic, vibrant flat colors, clean vector lines. No 3D rendering, no gradients, no chibi proportions. Characters must have pre-teen proportions (slightly shorter stature than a teenager, but with natural, balanced head-to-body ratios).
Setting: Interior of an American house, 2D flat style.
subject_definitions:<S2> is <Picture 2> a teen boy with a lean build and dynamic green skin. He has a thick, solid dark green fringe that is a single geometric block sweeping forward with a distinct outward curve at the corner, creating a slight “wing” shape over his forehead. No hair is not behind or below the ears. He has a whiny/nasally voice. He wears a gray t-shirt with dark red shorts.
<S1> is <Picture 1> a teen boy with a lean build and sky blue skin. He has a thick, solid dark blue fringe that is a single geometric block sweeping forward and curving slightly downward at the edges to cover his forehead. No hair is not behind or below the ears. He has a normal boy voice. He wears a white t-shirt and gray shirt. His legs move normally, it is just raised in the reference image.
1Action Sequence (17 Seconds Total):
0-2s: S1 stands silently on the left. S2 silently walks into the frame from the right holding a white box with a prominent orange letter “t”.
2-4s: S1 speaks: <d>[English] Another Temu package? </d> S2 remains silent, looking at the box.
4-6s: S2 speaks: <d>[English] “I got hearing cancelling headphones.” </d>
6-9s: S2 sets the box down and pulls the headphones from the box and puts them on his head. S1 remains silent, watching.
9-12s: S1 speaks with a questioning/confused face: <d>[English] “Do you mean noise cancelling headphones?” </d> S2 stands completely still with headphones on, mouth closed.
12-13s: S2 suddenly shouts: <d>[English] “What?” </d> S2 quickly pulls the headphones off his head and holds them in front of him.
13-16s: The Reveal: Stylized red blood drips from S2’s ears. S2 has a pained expression; S1 has a terrified expression.
16-17s: The camera focuses on the interior of the headphones, revealing sharp metal spikes pointing inward.


