IC-027LMM performance in deciphering oracle bone inscriptions is comparable to untrained humans for common characters but declines for rarer and structurally complex characters

Zijian Chen, tingzhu chen, Wenjun Zhang, Guangtao Zhai

SourceOBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

The study evaluates the interpretive abilities of 23 LMMs on 140 oracle bone characters across three datasets. Performance was measured using BERTScore to compare the model's generated English descriptions against expert-labelled golden descriptions. For high-frequency characters, models like GPT-4o, GPT-4V, and Gemini 1.5 Pro approached or exceeded the performance of a public-level human baseline. However, performance decreased with lower character frequency, and was noticeably worse for ideograms and structurally similar radical variants.

Evidence
correlational
Key metric
Average BERTScore on HUST-OBS: GPT-4o 0.3687, Gemini 1.5 Pro 0.3730, Public human 0.3943. Average BERTScore on radical subset: 0.1636.
Caveat
Decipherment is subjective and BERTScore is an approximation of semantic similarity, not definitive accuracy. The evaluation is limited to 140 characters and only uses English descriptions, potentially missing nuances of interpretation in other languages or from historical context.
Model
Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Gemini 1.5 Flash, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, GPT-4o, Qwen-VL Qwen-VL-Max, XGen-MM, mPLUG-Owl3, MiniCPM-V 2.6, Moondream2, InternVL2 InternVL2-Llama3-76B, InternVL2-40B, InternVL2-8B, GLM-4V GLM-4V-9B, CogVLM2 CogVLM2-Llama3-19B, LLaVA-NeXT / LLaVA 1.6, Idefics2 Idefics2-8B, DeepSeek-VL, InternLM-XComposer2-VL, LLaVA-1.5 / LLaVA-v1.5
Concepts
Failure mode
Datasets
HUST-OBS [eval], EVOBC [eval], OBI Component 20 [eval]
Methods
BERTScore [eval]
Extraction
automatic-extraction