2 , an interesting trend emerged: VideoLLaMA’s mean scores consistently fell between those of Researcher A and Researcher B, QwenVL tended to assign higher ratings than both experts, while InternVL produced consistently lower scores.
← all excerpts
Benchmark evaluation of video large language models in quality assessment of science popularization videos for dry eye.
3
—
—
The sentences
For VIQI II (information accuracy), both VideoLLaMA and InternVL presented significant agreement with both experts ( p < 0.05), whereas QwenVL again failed to achieve significance.
InternVL also showed statistical significance with Researcher B, while QwenVL did not reach statistical significance with either rater.