상세 보기
초록
Often, the needs and visual abilities differ between the annotator group and the end user group. Generating detailed diagram descriptions for blind and low-vision (BLV) users is one such challenging domain. Sighted annotators could describe visuals with ease, but existing studies have shown that direct generations by them are costly, bias-prone, and somewhat lacking by BLV standards. In this study, we ask sighted individuals to assess-rather than produce-diagram descriptions generated by vision-language models (VLM) that have been guided with latent supervision via a multipass inference. The sighted assessments prove effective and useful to professional educators who are themselves BLV and teach visually impaired learners. We release SIGHTATION, a collection of diagram description datasets spanning 5k diagrams and 137k samples for completion, preference, retrieval, question answering, and reasoning training purposes and demonstrate their fine-tuning potential in various downstream tasks. © 2025 Elsevier B.V., All rights reserved.
- 제목
- Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
- 저자
- Kang, Wanju; Kim, Eunki; An, Namin; Kim, Sangryul; Choi, Haemin; Kwak, Ki-hoon; Thorne, James
- 발행일
- 2025
- 유형
- Conference paper
- 저널명
- Proceedings of the Annual Meeting of the Association for Computational Linguistics
- 권
- 1
- 페이지
- 27585 ~ 27621