MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915588957274112 |
|---|---|
| author | Zhang, Wei Guo, Zekun Xia, Yingce Jin, Peiran Xie, Shufang Qin, Tao Li, Xiang-Yang |
| author_facet | Zhang, Wei Guo, Zekun Xia, Yingce Jin, Peiran Xie, Shufang Qin, Tao Li, Xiang-Yang |
| contents | Structure-based drug design (SBDD), which maps target proteins to candidate molecular ligands, is a fundamental task in drug discovery. Effectively aligning protein structural representations with molecular representations, and ensuring alignment between generated drugs and their pharmacological properties, remains a critical challenge. To address these challenges, we propose MolChord, which integrates two key techniques: (1) to align protein and molecule structures with their textual descriptions and sequential representations (e.g., FASTA for proteins and SMILES for molecules), we leverage NatureLM, an autoregressive model unifying text, small molecules, and proteins, as the molecule generator, alongside a diffusion-based structure encoder; and (2) to guide molecules toward desired properties, we curate a property-aware dataset by integrating preference data and refine the alignment process using Direct Preference Optimization (DPO). Experimental results on CrossDocked2020 demonstrate that our approach achieves state-of-the-art performance on key evaluation metrics, highlighting its potential as a practical tool for SBDD. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_27671 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design Zhang, Wei Guo, Zekun Xia, Yingce Jin, Peiran Xie, Shufang Qin, Tao Li, Xiang-Yang Artificial Intelligence Machine Learning Structure-based drug design (SBDD), which maps target proteins to candidate molecular ligands, is a fundamental task in drug discovery. Effectively aligning protein structural representations with molecular representations, and ensuring alignment between generated drugs and their pharmacological properties, remains a critical challenge. To address these challenges, we propose MolChord, which integrates two key techniques: (1) to align protein and molecule structures with their textual descriptions and sequential representations (e.g., FASTA for proteins and SMILES for molecules), we leverage NatureLM, an autoregressive model unifying text, small molecules, and proteins, as the molecule generator, alongside a diffusion-based structure encoder; and (2) to guide molecules toward desired properties, we curate a property-aware dataset by integrating preference data and refine the alignment process using Direct Preference Optimization (DPO). Experimental results on CrossDocked2020 demonstrate that our approach achieves state-of-the-art performance on key evaluation metrics, highlighting its potential as a practical tool for SBDD. |
| title | MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design |
| topic | Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2510.27671 |