IC-027LMM performance in deciphering oracle bone inscriptions is comparable to untrained humans for common characters but declines for rarer and structurally complex characters
The study evaluates the interpretive abilities of 23 LMMs on 140 oracle bone characters across three datasets. Performance was measured using BERTScore to compare the model's generated English descriptions against expert-labelled golden descriptions. For high-frequency characters, models like GPT-4o, GPT-4V, and Gemini 1.5 Pro approached or exceeded the performance of a public-level human baseline. However, performance decreased with lower character frequency, and was noticeably worse for ideograms and structurally similar radical variants.
Evidence
correlational
Key metric
Average BERTScore on HUST-OBS: GPT-4o 0.3687, Gemini 1.5 Pro 0.3730, Public human 0.3943. Average BERTScore on radical subset: 0.1636.
Caveat
Decipherment is subjective and BERTScore is an approximation of semantic similarity, not definitive accuracy. The evaluation is limited to 140 characters and only uses English descriptions, potentially missing nuances of interpretation in other languages or from historical context.