IC-025LMMs exhibit poor fine-grained perception in locating individual characters on original oracle bones

Zijian Chen, tingzhu chen, Wenjun Zhang, Guangtao Zhai

SourceOBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

The study evaluates the ability of 23 LMMs to perform dense object localization of oracle bone characters on original oracle bone images (from the O2BR dataset) and inked rubbings (from the Yinqiwenyuandetection dataset). The 'where' question required the models to return bounding boxes for each detected character. Performance was measured using mean Intersection over Union (mIoU). All evaluated models performed poorly, with Gemini 1.5 Pro significantly outperforming others by two orders of magnitude, yet still far from public-level human performance.

Evidence
correlational
Key metric
mIoU: Gemini 1.5 Pro 0.1126 (O2BR) / 0.0586 (Yinqiwenyuandetection); GPT-4o 0.0038 (O2BR) / 0.0182 (Yinqiwenyuandetection)
Caveat
Gemini 1.5 Pro's superior performance is attributed to its exclusive coordinate return function, which is not a standard capability of other models. Performance is far from human-level, indicating the task remains a significant challenge.
Model
Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Gemini 1.5 Flash, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, GPT-4o, Qwen-VL Qwen-VL-Max, XGen-MM, mPLUG-Owl3, MiniCPM-V 2.6, Moondream2, InternVL2 InternVL2-Llama3-76B, InternVL2-40B, InternVL2-8B, GLM-4V GLM-4V-9B, CogVLM2 CogVLM2-Llama3-19B, LLaVA-NeXT / LLaVA 1.6, Idefics2 Idefics2-8B, DeepSeek-VL, InternLM-XComposer2-VL, LLaVA-1.5 / LLaVA-v1.5
Concepts
Failure mode
Datasets
O2BR [eval], Yinqiwenyuandetection [eval]
Extraction
automatic-extraction