IC-1610Llama-2-7b-chat underperforms on small molecule editing tasks due to limited domain-specific pretraining

Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, Chaowei Xiao

SourceConversational Drug Editing Using Retrieval and Domain Feedback

When used as the backbone in the ChatDrug framework for 28 small molecule editing tasks (16 single-objective, 12 multi-objective), Llama-2-7b-chat outperforms baseline methods on only 7 of the 28 tasks. The authors attribute this to the model not being well pretrained on small molecule domain datasets. In contrast, the same model performs adequately on peptide and protein editing tasks, indicating the limitation is domain-specific rather than a general capability gap.

Evidence
correlational
Key metric
outperforming on 7 out of 28 small molecule tasks; e.g. task 101 (more soluble in water, loose threshold): 61.87 ± 2.67 hit ratio vs. 94.13 ± 1.04 for GPT-3.5-turbo
Caveat
The authors state the reason 'may be' limited domain pretraining; this is a hypothesis, not a verified cause. Performance is measured within the ChatDrug framework, not in isolation.
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b
Concepts
Failure mode
Datasets
ZINC [source]
Methods
MoleculeSTM [compared-to], Principal component analysis [compared-to]
Related work
MoleculeSTM [compared-to]
Related findings
IC-1611, IC-1612
Extraction
automatic-extraction