Table 4: Effect of joint cross attention on depth condition learning. Joint Cross Attention is our proposed method, which trains without tracking condition.

Depth Condition Background Consistency ↑ Subject Consistency ↑ Dynamic ↑ Imaging Quality ↑ CLIP-I ↑ DINO-I ↑
0.934 0.942 0.480 0.675 0.698 0.563
Controlnet 0.918 0.880 0.333 0.558 0.720 0.393
Joint Cross Attention (ours) 0.935 0.945 0.480 0.682 0.720 0.575