IC-026LLMs can assist in OB rejoining by identifying rejoinable fragments with moderate accuracy, but are not yet truly usable

Zijian Chen, tingzhu chen, Wenjun Zhang, Guangtao Zhai

SourceOBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

The study evaluates LMMs on the OB rejoining task using a new dataset (OBI-Rejoin). Models were assessed on their ability to determine if two fragments are rejoinable, using both 'yes-or-no' queries (measuring accuracy) and 'how' queries (measuring rejoinable probability with acc@k). GPT-4o and Gemini 1.5 Pro achieved acc@10 scores over 76%, suggesting a potential for assisting scholars by narrowing down candidates. However, open-source models performed poorly, with average acc@1, acc@5, and acc@10 scores of 3.85%, 11.39%, and 18.65%, respectively.

Evidence
correlational
Key metric
acc@10: GPT-4o 78.47%, Gemini 1.5 Pro 76.67%; acc@1: GPT-4o 29.13%, Gemini 1.5 Pro 24.88%
Caveat
Performance is significantly better with 'how' queries (probabilistic output) than 'yes-or-no' queries, suggesting that the format of the answer greatly influences the results. Open-source models remain far from practical utility.
Model
Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Gemini 1.5 Flash, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, GPT-4o, Qwen-VL Qwen-VL-Max, GLM-4V, XGen-MM, mPLUG-Owl3, MiniCPM-V 2.6, InternVL2 InternVL2-Llama3-76B, InternVL2-40B, InternVL2-8B, LLaVA-NeXT / LLaVA 1.6, Idefics2 Idefics2-8B, DeepSeek-VL
Concepts
Failure mode
Datasets
OBI-Rejoin [eval]
Extraction
automatic-extraction