ShallowBench: Benchmarking Generative Drug Design Models on Shallow-Pocket Targets
Introduces ShallowBench, a benchmark for evaluating generative drug models on low-concavity targets using Alpha Shape volume differences.
Key Findings
Methodology
A two-step volumetric approach was employed: first, generating an Alpha Shape mesh to define a 'lid' volume (Vlid), then calculating the protein atom voxel volume (Vatom). The concavity was quantified as Vlid minus Vatom. Targets with Concavity<500 ų and surface area>50 Ų were selected, resulting in 5780 shallow-pocket targets. The dataset was split into training (4995) and testing (785) using 30% sequence identity clustering. Evaluation metrics included binding affinity via AutoDock Vina, chemical validity, QED, and shape complementarity scores, across models DiffSBDD, SimpleSBDD, and TargetDiff.
Key Results
- All models showed decreased binding affinity on shallow targets; for example, TargetDiff's average Vina score dropped from -7.33 kcal/mol (control) to -5.26 kcal/mol. Chemical validity also declined from 87.51% to 79.71%. Shape complementarity and QED scores indicated models struggled to generate well-optimized ligands for flat surfaces. DiffSBDD maintained high validity (98.05%) but weaker affinity (-4.75 kcal/mol). SimpleSBDD achieved better affinity (-6.47 kcal/mol) but lower chemical validity. These results reveal the challenge of low-concavity interfaces for current generative models.
Significance
This study pioneers a dedicated benchmark for shallow protein targets, addressing a critical gap in structure-based drug design. By systematically evaluating models on targets like KRAS and MYC, it exposes their limitations in low-concavity environments, which are common in 'undruggable' proteins. The benchmark facilitates the development of new architectures and loss functions tailored for such challenging interfaces, accelerating progress toward therapeutics for traditionally elusive targets. It also guides industry efforts in early-stage drug discovery, emphasizing the need for models that can handle flat, featureless protein surfaces, thus broadening the scope of computational drug design.
Technical Contribution
The core technical innovation is the volumetric difference method based on Alpha Shapes to quantify pocket concavity, enabling precise screening of shallow-pocket targets. The creation of a curated dataset of 5780 targets with strict structural and sequence diversity, combined with a rigorous train/test split based on 30% sequence identity, provides a robust evaluation platform. The comparative assessment of state-of-the-art generative models highlights their weaknesses in low-concavity contexts, offering insights for architectural improvements such as multi-scale graph networks and reinforcement learning strategies to better anchor ligand generation on flat surfaces.
Novelty
This is the first work to systematically quantify and benchmark generative models on shallow, low-concavity protein surfaces using a volumetric Alpha Shape difference metric. It fills a key gap left by existing datasets biased toward deep pockets, providing a dedicated platform for evaluating model performance on flat interfaces. The approach of combining geometric volume analysis with rigorous dataset curation and splitting represents a significant methodological advance, enabling more realistic assessments of model capabilities in challenging scenarios.
Limitations
- Limited by computational resources, the evaluation was performed on a single seed without statistical error bars, potentially affecting robustness. AutoDock Vina scores, while standard, do not perfectly correlate with experimental binding affinities, which may limit real-world applicability. The models still struggle with chemical validity and shape complementarity on shallow surfaces, indicating the need for further architectural and loss function innovations. Future work should incorporate multiple seeds, experimental validation, and advanced geometric constraints.
Future Work
Future directions include expanding the shallow-pocket dataset with more diverse targets, developing specialized loss functions that penalize ligand 'floating', and designing architectures capable of modeling interactions across both deep pockets and broad flat surfaces. Incorporating reinforcement learning and agent-based iterative refinement could improve ligand anchoring and binding affinity. Combining computational predictions with experimental validation will be crucial for translating these models into practical drug discovery pipelines, ultimately enabling effective targeting of previously 'undruggable' proteins.
AI Executive Summary
This research introduces ShallowBench, a novel benchmark designed to evaluate the performance of generative drug design models on low-concavity protein targets. Traditional structure-based models excel in deep-pocket scenarios but falter when faced with flat, featureless protein surfaces like KRAS and MYC, which are considered 'undruggable.' To address this, the authors developed a two-step volumetric method based on Alpha Shape meshes to quantify pocket concavity, selecting 5780 targets with minimal cavity depth but sufficient surface area.
The dataset was carefully split into training and testing sets using 30% sequence identity clustering, ensuring rigorous evaluation. Three state-of-the-art models—DiffSBDD, SimpleSBDD, and TargetDiff—were assessed on this benchmark. Results showed a consistent decline in predicted binding affinity, with TargetDiff’s average Vina score decreasing from -7.33 kcal/mol (control) to -5.26 kcal/mol on shallow targets. Chemical validity also dropped, highlighting the difficulty of generating effective ligands in these challenging environments.
These findings underscore the limitations of current models and emphasize the need for architectural innovations, such as multi-scale graph networks and reinforcement learning, to better handle flat, low-concavity protein surfaces. The benchmark provides a critical resource for future research, aiming to unlock drug discovery for 'undruggable' targets, and paves the way for more versatile and robust generative models in structural biology.
Deep Analysis
Background
结构药物设计(SBDD)经历了从传统的基于结构的配体筛选到深度学习驱动的生成模型的演变。早期依赖高质量的蛋白-配体复合物结构,代表性工作如AutoDock、GOLD等。近年来,变分自编码器(VAE)、生成对抗网络(GAN)和扩散模型(Diffusion Models)等深度学习方法显著提升了候选药物的多样性和药效预测能力。然而,这些模型主要在深腔蛋白靶点上表现优异,面对浅口或无口界面时性能明显下降。深腔提供丰富的空间约束,便于模型学习和生成,而浅口蛋白缺乏此类结构特征,导致生成的配体“漂浮”或结合不良,限制了药物研发的突破。
Core Problem
浅口蛋白靶点如KRAS、MYC等因缺乏深腔而被视为“难啃”目标,传统模型难以有效生成结合配体。现有数据集偏重深腔目标,导致模型在浅口目标上的泛化能力不足。浅口界面受 solvent competition、有限的接触面积和几何约束限制,增加了生成难度。缺乏专门的评估基准使得模型性能难以量化,阻碍了创新。解决这一问题需要新颖的筛选方法和模型架构,提升浅口目标的药物设计能力。
Innovation
本研究的核心创新包括:1)提出基于Alpha Shape的体积差异指标,有效筛选低凹度靶点,确保目标具有足够的表面积;2)建立涵盖5780个浅口靶点的专用数据集,结合严格的序列相似性划分,支持模型微调和性能评估;3)系统性评估多种前沿生成模型在浅口蛋白上的表现,揭示其在低凹度界面上的局限,为模型改进提供指导;4)强调浅口蛋白的特殊几何特性,推动药物设计从深腔向浅口甚至无口界面拓展。
Methodology
- �� 目标筛选:利用蛋白-配体复合体的质心,提取8 Å范围内的蛋白原子。• 凹度指标:通过建立蛋白原子体素体积(Vatom)和Alpha Shape“盖子”体积(Vlid),定义凹度为Vlid - Vatom。• 筛选条件:Concavity<500 ų,表面积>50 Ų。• 数据集划分:采用30%序列相似性阈值,将目标分为训练(4995)和测试(785)集。• 评估指标:结合能、化学有效性、QED、形状匹配Score。• 模型:DiffSBDD(扩散模型)、SimpleSBDD(简化模型)、TargetDiff(目标导向扩散模型)。
Experiments
- �� 数据集:基于CrossDocked2020筛选浅口目标,结合控制集进行对比。• 评估:在GPU集群上运行模型,采样每个目标一个配体。• 超参数:DiffSBDD采用1000步扩散,SimpleSBDD从ZINC库采样,TargetDiff使用预训练模型。• 统计:比较结合能、化学有效性、QED和形状匹配Score。• 重点:验证模型在浅口蛋白上的性能瓶颈和潜在改进方向。
Results
- �� 所有模型在浅口靶点上的结合能表现明显差于控制集,TargetDiff由-7.33降至-5.26 kcal/mol。• DiffSBDD保持最高化学有效性(98.05%),但结合能较差。• SimpleSBDD在结合能上表现较优,但化学合理性较低。• TargetDiff在形状匹配方面表现较好,但化学有效性下降,显示模型在低凹度界面上的适应性不足。• 结果强调浅口蛋白的结构特性对生成模型提出了更高要求。
Applications
- �� 立即应用:可用于药物筛选平台,辅助发现KRAS、MYC等“难啃”靶点的小分子候选。• 长远目标:推动模型在低凹度界面的优化,结合实验验证,实现从虚拟筛选到临床候选药物的转变,解决“难啃”靶点的药物研发瓶颈。
Limitations & Outlook
- �� 计算资源限制,未能多次随机种子评估,影响统计稳健性。• AutoDock Vina的评分与实际结合能相关性有限,可能影响结果的实际指导价值。•模型在浅口蛋白上的性能仍不足,尤其在化学合理性和几何匹配方面,需结合强化学习、多尺度建模进行优化。
Plain Language Accessible to non-experts
想象你在一个工厂里,生产各种各样的玩具。深口蛋白就像有很多隐藏的洞,可以放入不同的零件,工人们很容易把零件放进去,组装也很顺利。而浅口蛋白就像平坦的桌面,没有洞,零件要怎么放进去呢?这就像拼拼图,没有明显的凹槽,拼图块很容易“漂浮”在空中,难以拼出完整的图案。研究者用一种特殊的“体积差”方法,找到那些没有深洞但仍有空间的目标,就像用尺子测量桌面和拼图块的距离,确保它们可以拼在一起。这个新方法帮助我们找到那些“难啃”的目标,让药物设计变得更有挑战性,也更有意义。
ELI14 Explained like you're 14
假设你在玩拼图游戏,有些拼图块可以插到深深的洞里,很容易拼好,但有些拼图块只是在平坦的桌面上,没有洞,怎么拼都很难。这就像一些蛋白质表面很平,没有深槽,药物分子很难找到地方“卡住”。科学家们用一种特别的尺子,测量蛋白质表面是不是有深槽,或者只是平坦的空间。他们用这个方法筛选出那些没有深槽但仍有空间的蛋白质,然后用电脑模型试着设计药物分子。虽然这些药物设计比深槽的蛋白质更难,但这一步很重要,因为很多“难啃”的目标,比如KRAS和MYC,都属于这种平坦的蛋白质表面。这个研究帮助我们更好地理解和攻克这些“难啃”的蛋白质,为未来新药的开发打开了新思路。
Abstract
While generative AI models have demonstrated remarkable success in structure-based drug design, they predominantly rely on deep binding pockets and struggle to sample effective ligands for challenging low-pocketability targets, such as the historically "undruggable" oncology targets KRAS and MYC. To address this gap, we introduce ShallowBench, a strictly curated benchmark of 5,780 shallow-pocket targets extracted from CrossDocked2020. By computing the difference between an Alpha Shape "lid" volume and the underlying protein atom voxel volume, we successfully isolated targets with low concavity while ensuring sufficient surface area for binding. Evaluating various state-of-the-art generative models reveals weaker predicted binding affinity on these low-concavity interfaces. ShallowBench therefore provides a rigorous benchmark for generative biology models and highlights the necessity of new architectural innovations or loss functions capable of navigating these challenging targets.