Abstract:
Computer vision benchmarks are built on the assumption that a model's score reflects the capability we intend to measure. In this talk, we examine invisible biases: factors correlated with a model's outcome that are not visible in the data itself, yet systematically skew what an evaluation measures. Across three recent studies, we show that such hidden factors are more common and consequential than expected. We first examine the invisible bias introduced by cameras themselves, through acquisition and processing parameters. We then explore data leakage, showing how evaluation images frequently reappear, identically or near-identically, in training data, inflating reported performance and undermining fair comparison across models. Finally, we discuss gender bias in vision-language models, where demographic priors distort predictions independent of the visual content itself. Overall, the talk presents a broader view of bias in evaluation practices, not only as a fairness concern, but as a general threat to the construct validity of benchmarks, models, and visual representations.
Public events of RIKEN Center for Advanced Intelligence Project (AIP)
Join community