Document detail
ID

oai:arXiv.org:2405.16234

Topic
Computer Science - Computer Vision...
Author
Xia, Shiyu Xiong, Junyu Dong, Haoyu Zhao, Jianbo Tian, Yuzhang Zhou, Mengyu He, Yeye Han, Shi Zhang, Dongmei
Category

Computer Science

Year

2024

listing date

10/2/2024

Keywords
challenges capabilities vision recognition spreadsheet
Metrics

Abstract

This paper explores capabilities of Vision Language Models on spreadsheet comprehension.

We propose three self-supervised challenges with corresponding evaluation metrics to comprehensively evaluate VLMs on Optical Character Recognition (OCR), spatial perception, and visual format recognition.

Additionally, we utilize the spreadsheet table detection task to assess the overall performance of VLMs by integrating these challenges.

To probe VLMs more finely, we propose three spreadsheet-to-image settings: column width adjustment, style change, and address augmentation.

We propose variants of prompts to address the above tasks in different settings.

Notably, to leverage the strengths of VLMs in understanding text rather than two-dimensional positioning, we propose to decode cell values on the four boundaries of the table in spreadsheet boundary detection.

Our findings reveal that VLMs demonstrate promising OCR capabilities but produce unsatisfactory results due to cell omission and misalignment, and they notably exhibit insufficient spatial and format recognition skills, motivating future work to enhance VLMs' spreadsheet data comprehension capabilities using our methods to generate extensive spreadsheet-image pairs in various settings.

Xia, Shiyu,Xiong, Junyu,Dong, Haoyu,Zhao, Jianbo,Tian, Yuzhang,Zhou, Mengyu,He, Yeye,Han, Shi,Zhang, Dongmei, 2024, Vision Language Models for Spreadsheet Understanding: Challenges and Opportunities

Document

Open

Share

Source

Articles recommended by ES/IODE AI

Lung cancer risk and exposure to air pollution: a multicenter North China case–control study involving 14604 subjects
lung cancer case–control air pollution never-smokers nomogram model controls lung-related 14604 subjects north polluted consistent smokers quit exposure lung cancer risk air people factor smoking pollution study history