DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery
Introduces LLM-guided symbolic regression (SR) framework for automated kinetic model discovery, reducing iteration count by up to 79.3%.
Key Findings
Methodology
This paper proposes DASyR-LLM, an innovative framework integrating large language models (LLMs) into the symbolic regression process for kinetic model discovery. The core involves: • Using ADoK-S as the SR backbone, combining model structure search, parameter estimation, and model selection via AIC; • At each iteration, the LLM qualitatively critiques the best candidate models, assessing their physical plausibility and identifying structural inconsistencies; • The LLM also generates new candidate expressions based on chemical knowledge and current models, guiding the search toward physically meaningful structures; • Synthetic datasets from four in silico case studies of increasing complexity—covering heterogeneous catalysis and bioprocesses—are used to validate the framework, comparing pure SR with LLM-guided approaches. Results demonstrate significant efficiency improvements, with the LLM reducing the number of iterations needed for model identification and proposing correct structures in over half of guided runs.
Key Results
- Across four simulated cases, the LLM-guided SR reduced the number of iterations required to identify the ground-truth model by an average of 41.7% to 79.3%, depending on case complexity. For example, in a complex catalytic reaction network, traditional SR took around 20 iterations, whereas the LLM-guided method required only 4 to 12 iterations. More than 50% of runs with LLM guidance successfully proposed models matching the true structure, confirming its effectiveness in structural inference.
- Predictive performance on independent validation datasets remained high, with R² exceeding 0.98 in all cases, indicating that the models generalize well despite the reduced search effort. Ablation studies revealed that both the size of the LLM and the SR component contributed to performance, with smaller LLMs still maintaining high efficiency, demonstrating robustness and scalability.
- Practically, this approach can substantially cut down experimental efforts, especially in settings where each iteration involves wet-lab experiments. The framework’s ability to incorporate domain knowledge via LLMs enhances model interpretability and physical plausibility, making it a promising tool for accelerating reaction mechanism discovery and process optimization in industry.
Significance
This work addresses a long-standing challenge in scientific modeling: how to efficiently discover physically consistent and interpretable kinetic models from noisy data. By embedding LLMs into the SR pipeline, the authors bridge data-driven and knowledge-based approaches, enabling models that are both accurate and scientifically plausible. The methodology significantly reduces the number of experimental iterations needed, lowering costs and time in process development. Moreover, it exemplifies a new paradigm where AI reasoning complements traditional modeling, paving the way for fully automated, domain-aware scientific discovery pipelines. This advancement holds transformative potential for chemical engineering, catalysis, systems biology, and beyond, fostering faster innovation cycles and deeper mechanistic insights.
Technical Contribution
The main technical innovation lies in integrating LLMs as both qualitative evaluators and generative partners within an iterative SR framework. Unlike prior approaches that rely solely on statistical metrics or predefined physical constraints, this method leverages LLMs’ reasoning capabilities to assess model plausibility based on embedded domain knowledge. The framework employs a feedback loop where the LLM critiques candidate models, proposes new expressions, and guides the search process, effectively narrowing the search space while maintaining flexibility. This hybrid approach combines the strengths of symbolic regression, model selection via AIC, and knowledge-guided hypothesis generation, resulting in a more efficient and physically meaningful discovery process. Additionally, the framework demonstrates robustness to LLM size reduction, indicating practical scalability.
Novelty
This study is the first to embed large language models directly into the symbolic regression process for reaction kinetics, enabling qualitative reasoning about model plausibility alongside candidate generation. Unlike existing methods that depend solely on statistical fit or predefined physical constraints, this approach uses LLMs to incorporate scientific reasoning dynamically, guiding the search toward physically consistent models. The integration of LLMs within an iterative model discovery and experimental design loop represents a novel paradigm shift, offering a more intelligent and autonomous pathway for scientific model building. This dual role of critique and proposal, combined with active experimental guidance, distinguishes this work from prior literature.
Limitations
- The effectiveness of LLMs depends heavily on their pretraining and embedded knowledge base, which may limit their ability to evaluate models involving novel or poorly documented chemistry, potentially biasing results.
- In highly complex reaction networks or noisy data scenarios, the stability and convergence speed of the framework might decrease, requiring further optimization or hybridization with other methods.
- The current validation relies on synthetic datasets; real experimental data often contain measurement errors, system non-idealities, and unmodeled phenomena that could challenge the model’s robustness and accuracy. Future work should focus on real-world applications and robustness enhancements.
Future Work
Future directions include expanding the chemical knowledge embedded in LLMs to handle more diverse reaction mechanisms, integrating multi-modal data sources such as spectral or imaging data, and developing adaptive prompting strategies for better model critique and proposal. Additionally, applying the framework to real experimental datasets will be crucial for assessing practical utility. Combining this approach with active learning and real-time experimental feedback could further accelerate model discovery. Exploring multi-task learning to incorporate thermodynamic and kinetic constraints simultaneously, as well as extending the methodology to other scientific domains like systems biology or materials science, are promising avenues for future research.
AI Executive Summary
Understanding complex chemical and biological reaction systems has long been a central challenge in chemical engineering. Traditional modeling approaches rely heavily on expert intuition, extensive experimentation, and iterative hypothesis testing, which are often time-consuming and costly. Symbolic regression (SR) has emerged as a promising data-driven technique capable of automatically deriving explicit, interpretable mathematical models directly from experimental data. However, conventional SR methods face significant limitations: they often explore vast, high-dimensional search spaces without incorporating domain knowledge, leading to physically implausible models, slow convergence, and high experimental costs.
Addressing these issues, the authors introduce DASyR-LLM, a novel framework that integrates large language models (LLMs) into the symbolic regression process. This hybrid approach leverages the reasoning and knowledge capabilities of LLMs to guide the search for physically meaningful kinetic models. The framework builds upon the ADoK-S algorithm, which combines symbolic regression with parameter estimation and model selection via Akaike Information Criterion (AIC). In each iteration, the LLM evaluates the best candidate models based on qualitative chemical and physical reasoning, identifying structural inconsistencies or implausible features. Simultaneously, it generates new candidate expressions informed by embedded chemical knowledge, effectively narrowing the search space.
The methodology involves a multi-step process: synthetic datasets are generated for four in silico case studies, ranging from heterogeneous catalysis to bioprocess systems. The SR algorithm first fits concentration trajectories and estimates derivatives, then searches for kinetic rate laws. The LLM critiques these models and proposes new candidates, which are evaluated via simulation and model selection metrics. This iterative loop continues until the true model is identified or the experimental budget is exhausted. Results show that the LLM-guided approach reduces the number of iterations by up to 79.3%, significantly accelerating the discovery process. Moreover, over half of guided runs successfully propose the correct model structure, demonstrating its potential for automating scientific discovery.
Predictive performance on independent validation datasets remains high, with R² exceeding 0.98 across all cases, confirming the models’ robustness. Ablation studies reveal that both the size of the LLM and the SR component influence efficiency, with smaller LLMs still maintaining substantial gains. The practical implications are profound: by embedding domain reasoning into the model search, the framework reduces the need for costly wet-lab experiments, making kinetic modeling more accessible and scalable. This work marks a significant step toward fully automated, domain-aware scientific modeling pipelines, with broad applications in catalysis, bioprocessing, and systems biology. Future research will focus on real-world data validation, knowledge base expansion, and integration with active experimental design, promising a new era of intelligent scientific discovery.
Deep Dive
Abstract
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.