AI is increasingly transforming drug discovery by enabling researchers to analyse complex biological data, identify hidden relationships and prioritise promising therapeutic candidates. One emerging area is the use of multimodal artificial intelligence to predict drug-drug and gene-gene synergies, potentially making combination-drug discovery more efficient, targeted and biologically informed.
Dr. Junzhou Huang, Jenkins Garrett Endowed Professor in the Department of Computer Science and Engineering at the University of Texas at Arlington, is working at the intersection of artificial intelligence, deep learning and biomedical research to advance this approach. His project brings together multimodal large language models, deep-learning techniques and Bayesian statistical modelling with drug-screening, multi-omics, chemical and gene-level data to identify and predict potential synergies.
A key focus of the research is addressing the complexity and incompleteness of biological data. By combining structured biological knowledge, genetic and sequence information and insights from biomedical literature, the proposed framework aims to develop richer representations of gene function and uncover relationships that may otherwise remain difficult to detect. Bayesian modelling is expected to complement deep learning by quantifying uncertainty and helping researchers distinguish robust predictions from context-dependent or less certain results.
In this interview with AI Spectrum, Dr. Huang discusses the AI architectures underpinning the research, the challenges of integrating heterogeneous biological datasets, and how Bayesian methods can improve the reliability and interpretability of synergy predictions. He also explains how high-throughput experimental screening could create an iterative feedback loop between AI predictions and laboratory validation, and explores the potential of AI-driven approaches to reduce the time and resources required for combination-drug discovery and eventually support more personalised therapeutic strategies.
AI Architecture: The project will use multimodal large language models and deep-learning approaches to predict gene-gene synergies. What specific AI architectures or techniques will be explored, and how will these models address the challenge of limited functional annotations for thousands of genes?
Answer: We will use transformer-based and BERT-style models within a multimodal deep-learning framework to improve gene function prediction and downstream synergy analysis. One of the major challenges is that functional annotations are still incomplete for many genes, so our approach combines structured biological knowledge, sequence information, and literature-derived information to build richer biological representations. By improving the completeness and quality of gene-function representations, we can better capture functional relationships among genes and use those relationships to support gene-gene and drug-drug synergy prediction. More broadly, the goal is to develop a knowledge-informed and biologically grounded AI framework that can learn from multiple complementary data sources rather than relying on any single annotation source.
Data Integration: The research aims to integrate drug-screening, multi-omics, chemical, and gene data. What are the key technical challenges in bringing together such heterogeneous datasets, and how will the proposed framework extract meaningful relationships across these data types?
Answer: One of the main challenges is that these datasets describe very different levels of the biological system. Drug-screening data reflect treatment response, multi-omics data capture the molecular state of cells, chemical data describe the compounds, and gene-level information provides functional context. They also differ in scale, completeness, and noise, which makes direct integration difficult. Our approach is to use a multimodal, biologically structured framework that first organizes these data according to their biological roles and then learns their relationships jointly. Rather than treating all features equally, the framework is designed to account for the fact that drug effects depend on both the properties of the compounds and the molecular context in which they are tested. This allows us to identify context-dependent patterns and biologically meaningful associations across data types, while avoiding simple feature aggregation.
Bayesian Modeling: How will Bayesian statistical modeling complement the deep-learning component of the project? Specifically, how will Bayesian methods help quantify uncertainty and improve the interpretability and reliability of predicted drug-drug synergies?
Answer: The deep-learning and Bayesian components play complementary roles in the project. Deep learning is used to capture complex biological patterns from multimodal data, while the Bayesian component provides a principled way to quantify uncertainty and account for variation across different biological contexts. The Bayesian model evaluates not only the expected synergy effect, but also how stable that effect is across different drugs, cell lines, and molecular backgrounds. This allows the framework to distinguish predictions that are consistently supported from those that remain uncertain. It also helps connect predicted synergy with interpretable biological factors, making the results more reliable and more useful for downstream experimental decision-making.
Explainability: Explainability is crucial when AI predictions influence drug-development decisions. How will the framework provide insights into the underlying biological mechanisms rather than simply producing a prediction?
Answer: For explainability, our goal is to make the predictions biologically interpretable rather than treating the model as a black box. The framework is designed to connect predicted drug synergy back to functional information about the target genes, relevant biological processes, and the molecular context in which the drugs are tested. In addition, the Bayesian component helps identify which biological factors are most associated with a predicted effect and whether that effect is broadly consistent or context-dependent. By combining these model-based explanations with knowledge extracted from the biomedical literature, we can move from simply predicting that a drug pair may be synergistic to generating biologically meaningful and experimentally testable hypotheses about why that synergy may occur.
Experimental Validation: The project includes high-throughput screening experiments to validate predictions. How will experimental results be fed back into the models, and will this create an iterative learning framework to progressively improve synergy predictions?
Answer: The experimental validation is designed to be more than a final check of the model. We will compare computational predictions with results from high-throughput screening and use both agreements and disagreements to better understand where the models work well and where they need improvement. Those experimental results can provide new evidence for refining both the gene-level and drug-level prediction components, creating an iterative feedback process between computation and experiment. The broader idea is that AI can guide which biological relationships are worth testing, while experimental results can in turn make future predictions more accurate and more reliable.
Future Impact: How could this AI-driven approach change the economics and timelines of combination-drug discovery, and what steps are required before these computational predictions can be translated into clinical development or precision medicine applications?
Answer: The main potential impact is to make combination-drug discovery more selective and more efficient. Instead of relying mainly on large experimental screens across a very large number of possible drug pairs, the computational framework can help prioritize combinations that have stronger predicted synergy and stronger biological support. That could reduce the number of low-value experiments and shorten the path from candidate generation to experimental testing. Moving from computational discovery toward clinical translation will then require several stages of validation, including independent experimental studies, testing across additional biological systems and disease settings, in vivo evaluation, and eventually clinical studies. The proposed framework can make candidate prioritization and experimental design more efficient, with the longer-term potential to support more personalized and evidence-driven combination therapies.

