Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS COCO) that do not showcase complex human interactions or behaviors, only a handful of non curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing...