← all papers · overview

Integrative Bioinformatics: Protein Structure Prediction Using Transformer-Based Multi-Omics Models

Abstract

The recent developments in both transformer designs and large protein language models have spawned a paradigm shift in computational structural biology, making it possible to predict protein three-dimensional structures with high fidelity, directly using sequence data and contextual biological signals. The proposed information in this paper is an integrative framework that builds on transformer-based protein language models and adds complementary multi-omics modalities (genomics, transcriptomics, proteomics, epigenomics and small-molecule interaction data) to the framework by introducing modality-aware attention and multi-modal fusion layers. The approach, which is proposed, makes use of pretrained protein transformers to encode evolutionary and physicochemical priors and multi-omics encoders inject cellular-context information which informs the conformational preference, post-translational modifications and interaction propensity. We introduce model architectures, training paradigms (self-supervised pretraining, cross-modal alignment and task-specific fine-tuning with auxiliary structure-aware losses (distance maps, torsion angle distributions, and interresidue contact likelihoods), and scalable inference policies of complex assemblies. Comparisons with current baselines (such as single-modality transformer predictors and end-to-end structure networks) prove to have accuracy improvement on a context-dependent conformation and complex assembly, especially isoforms and contextually modified proteins. Other topics that we talk about are calibration, interpretability (attention-based saliency, gradient-based attribution and motifconditioned activation maps), and the pathways of experimental validation. Ethical and reproducibility considerations and recommendations of data-sharing of multi-omics-augmented structural predictors are given. Such an integrative direction will bridge the sequence-function gap because structural prediction will now be placed in the more biological context that dictates protein behaviour.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).