Table 3: Per-sample quantitative content-fidelity results for the five selected cuts (Sample 1–5 with their mean), together with the mean over an expanded cut set taken in reading order for generalization (cut index ≤10, n=7; ≤20, n=14).

Metric (comparison) 5 selected cuts Expanded mean
Sample 1 Sample 2 Sample 3 Sample 4 Sample 5 Mean ≤10 ≤20
Storyboard VQAScore (per-cut text cut image) 0.847 0.872 0.707 0.917 0.721 0.813 0.798 0.805
Photoreal CLIP-I (webtoon cut photoreal image) 0.713 0.745 0.736 0.711 0.667 0.714 0.673 0.652
Video VQAScore (generated video cut description) 0.843 0.724 0.902 0.834 0.808 0.822 0.756 0.756
Video ViCLIP (generated video cut description) 0.349 0.224 0.241 0.249 0.257 0.264 0.173 0.175
Composite quality score S 0.801 0.780 0.782 0.821 0.732 0.783 0.742 0.738