Table 4: Effect of joint cross attention on depth condition learning.
Joint Cross Attention
is our proposed method, which trains without tracking condition.
Depth Condition
Background Consistency ↑
Subject Consistency ↑
Dynamic ↑
Imaging Quality ↑
CLIP-I ↑
DINO-I ↑
✗
0.934
0.942
0.480
0.675
0.698
0.563
Controlnet
0.918
0.880
0.333
0.558
0.720
0.393
Joint Cross Attention (ours)
0.935
0.945
0.480
0.682
0.720
0.575