Refine
Year of publication
Language
- English (99) (remove)
Is part of the Bibliography
- yes (99)
Keywords
- Arabidopsis thaliana (8)
- Network clustering (5)
- Protein complexes (3)
- Species comparison (3)
- respiration (3)
- Ascophyllum nodosum (2)
- Coherent partition (2)
- Graph partitions (2)
- GxE interaction (2)
- Metabolic networks (2)
Time hierarchies, arising as a result of interactions between system's components, represent a ubiquitous property of dynamical biological systems. In addition, biological systems have been attributed switch-like properties modulating the response to various stimuli across different organisms and environmental conditions. Therefore, establishing the interplay between these features of system dynamics renders itself a challenging question of practical interest in biology. Existing methods are suitable for systems with one stable steady state employed as a well-defined reference. In such systems, the characterization of the time hierarchies has already been used for determining the components that contribute to the dynamics of biological systems. However, the application of these methods to bistable nonlinear systems is impeded due to their inherent dependence on the reference state, which in this case is no longer unique. Here, we extend the applicability of the reference-state analysis by proposing, analyzing, and applying a novel method, which allows investigation of the time hierarchies in systems exhibiting bistability. The proposed method is in turn used in identifying the components, other than reactions, which determine the systemic dynamical properties. We demonstrate that in biological systems of varying levels of complexity and spanning different biological levels, the method can be effectively employed for model simplification while ensuring preservation of qualitative dynamical properties (i.e., bistability). Finally, by establishing a connection between techniques from nonlinear dynamics and multivariate statistics, the proposed approach provides the basis for extending reference-based analysis to bistable systems.
Cotton (Gossypium hirsutum) fibres consist of single cells that grow in a highly polarized manner, assumed to be controlled by the cytoskeleton(1-3). However, how the cytoskeletal organization and dynamics underpin fibre development remains unexplored. Moreover, it is unclear whether cotton fibres expand via tip growth or diffuse growth(2-4). We generated stable transgenic cotton plants expressing fluorescent markers of the actin and microtubule cytoskeleton. Live-cell imaging revealed that elongating cotton fibres assemble a cortical filamentous actin network that extends along the cell axis to finally form actin strands with closed loops in the tapered fibre tip. Analyses of F-actin network properties indicate that cotton fibres have a unique actin organization that blends features of both diffuse and tip growth modes. Interestingly, typical actin organization and endosomal vesicle aggregation found in tip-growing cell apices were not observed in fibre tips. Instead, endomembrane compartments were evenly distributed along the elongating fibre cells and moved bi-directionally along the fibre shank to the fibre tip. Moreover, plus-end tracked microtubules transversely encircled elongating fibre shanks, reminiscent of diffusely growing cells. Collectively, our findings indicate that cotton fibres elongate via a unique tip-biased diffuse growth mode.
Motivation:
Constraint-based modeling approaches allow the estimation of maximal in vivo enzyme catalytic rates that can serve as proxies for enzyme turnover numbers. Yet, genome-scale flux profiling remains a challenge in deploying these approaches to catalogue proxies for enzyme catalytic rates across organisms.
Results:
Here, we formulate a constraint-based approach, termed NIDLE-flux, to estimate fluxes at a genome-scale level by using the principle of efficient usage of expressed enzymes. Using proteomics data from Escherichia coli, we show that the fluxes estimated by NIDLE-flux and the existing approaches are in excellent qualitative agreement (Pearson correlation > 0.9). We also find that the maximal in vivo catalytic rates estimated by NIDLE-flux exhibits a Pearson correlation of 0.74 with in vitro enzyme turnover numbers. However, NIDLE-flux results in a 1.4-fold increase in the size of the estimated maximal in vivo catalytic rates in comparison to the contenders. Integration of the maximum in vivo catalytic rates with publically available proteomics and metabolomics data provide a better match to fluxes estimated by NIDLE-flux. Therefore, NIDLE-flux facilitates more effective usage of proteomics data to estimate proxies for kcatomes.
The unicellular green alga Chlamydomonas reinhardtii is a long-established model organism for studies on photosynthesis and carbon metabolism-related physiology. Under conditions of air-level carbon dioxide concentration [CO2], a carbon concentrating mechanism (CCM) is induced to facilitate cellular carbon uptake. CCM increases the availability of carbon dioxide at the site of cellular carbon fixation. To improve our understanding of the transcriptional control of the CCM, we employed FAIRE-seq (formaldehyde-assisted Isolation of Regulatory Elements, followed by deep sequencing) to determine nucleosome-depleted chromatin regions of algal cells subjected to carbon deprivation. Our FAIRE data recapitulated the positions of known regulatory elements in the promoter of the periplasmic carbonic anhydrase (Cah1) gene, which is upregulated during CCM induction, and revealed new candidate regulatory elements at a genome-wide scale. In addition, time series expression patterns of 130 transcription factor (TF) and transcription regulator (TR) genes were obtained for cells cultured under photoautotrophic condition and subjected to a shift from high to low [CO2]. Groups of co-expressed genes were identified and a putative directed gene-regulatory network underlying the CCM was reconstructed from the gene expression data using the recently developed IOTA (inner composition alignment) method. Among the candidate regulatory genes, two members of the MYB-related TF family, Lcr1 (Low-CO2 response regulator 1) and Lcr2 (Low-CO2 response regulator 2), may play an important role in down-regulating the expression of a particular set of TF and TR genes in response to low [CO2]. The results obtained provide new insights into the transcriptional control of the CCM and revealed more than 60 new candidate regulatory genes. Deep sequencing of nucleosome-depleted genomic regions indicated the presence of new, previously unknown regulatory elements in the C. reinhardtii genome. Our work can serve as a basis for future functional studies of transcriptional regulator genes and genomic regulatory elements in Chlamydomonas.
COMMIT
(2022)
Composition and functions of microbial communities affect important traits in diverse hosts, from crops to humans. Yet, mechanistic understanding of how metabolism of individual microbes is affected by the community composition and metabolite leakage is lacking. Here, we first show that the consensus of automatically generated metabolic reconstructions improves the quality of the draft reconstructions, measured by comparison to reference models. We then devise an approach for gap filling, termed COMMIT, that considers metabolites for secretion based on their permeability and the composition of the community. By applying COMMIT with two soil communities from the Arabidopsis thaliana culture collection, we could significantly reduce the gap-filling solution in comparison to filling gaps in individual reconstructions without affecting the genomic support. Inspection of the metabolic interactions in the soil communities allows us to identify microbes with community roles of helpers and beneficiaries. Therefore, COMMIT offers a versatile fully automated solution for large-scale modelling of microbial communities for diverse biotechnological applications. <br /> Author summaryMicrobial communities are important in ecology, human health, and crop productivity. However, detailed information on the interactions within natural microbial communities is hampered by the community size, lack of detailed information on the biochemistry of single organisms, and the complexity of interactions between community members. Metabolic models are comprised of biochemical reaction networks based on the genome annotation, and can provide mechanistic insights into community functions. Previous analyses of microbial community models have been performed with high-quality reference models or models generated using a single reconstruction pipeline. However, these models do not contain information on the composition of the community that determines the metabolites exchanged between the community members. In addition, the quality of metabolic models is affected by the reconstruction approach used, with direct consequences on the inferred interactions between community members. Here, we use fully automated consensus reconstructions from four approaches to arrive at functional models with improved genomic support while considering the community composition. We applied our pipeline to two soil communities from the Arabidopsis thaliana culture collection, providing only genome sequences. Finally, we show that the obtained models have 90% genomic support and demonstrate that the derived interactions are corroborated by independent computational predictions.
Understanding metabolic acclimation of plants to challenging environmental conditions is essential for dissecting the role of metabolic pathways in growth and survival. As stresses involve simultaneous physiological alterations across all levels of cellular organization, a comprehensive characterization of the role of metabolic pathways in acclimation necessitates integration of genome-scale models with high-throughput data. Here, we present an integrative optimization-based approach, which, by coupling a plant metabolic network model and transcriptomics data, can predict the metabolic pathways affected in a single, carefully controlled experiment. Moreover, we propose three optimization-based indices that characterize different aspects of metabolic pathway behavior in the context of the entire metabolic network. We demonstrate that the proposed approach and indices facilitate quantitative comparisons and characterization of the plant metabolic response under eight different light and/or temperature conditions. The predictions of the metabolic functions involved in metabolic acclimation of Arabidopsis thaliana to the changing conditions are in line with experimental evidence and result in a hypothesis about the role of homocysteine-to-Cys interconversion and Asn biosynthesis. The approach can also be used to reveal the role of particular metabolic pathways in other scenarios, while taking into consideration the entirety of characterized plant metabolism.
Highly efficient and accurate selection of elite genotypes can lead to dramatic shortening of the breeding cycle in major crops relevant for sustaining present demands for food, feed, and fuel. In contrast to classical approaches that emphasize the need for resource-intensive phenotyping at all stages of artificial selection, genomic selection dramatically reduces the need for phenotyping. Genomic selection relies on advances in machine learning and the availability of genotyping data to predict agronomically relevant phenotypic traits. Here we provide a systematic review of machine learning approaches applied for genomic selection of single and multiple traits in major crops in the past decade. We emphasize the need to gather data on intermediate phenotypes, e.g. metabolite, protein, and gene expression levels, along with developments of modeling techniques that can lead to further improvements of genomic selection. In addition, we provide a critical view of factors that affect genomic selection, with attention to transferability of models between different environments. Finally, we highlight the future aspects of integrating high-throughput molecular phenotypic data from omics technologies with biological networks for crop improvement.
Selection of high-performance lines with respect to traits of interest is a key step in plant breeding. Genomic prediction allows to determine the genomic estimated breeding values of unseen lines for trait of interest using genetic markers, e.g. single-nucleotide polymorphisms (SNPs), and machine learning approaches, which can therefore shorten breeding cycles, referring to genomic selection (GS). Here, we applied GS approaches in two populations of Solanaceous crops, i.e. tomato and pepper, to predict morphometric and colorimetric traits. The traits were measured by using scoring-based conventional descriptors (CDs) as well as by Tomato Analyzer (TA) tool using the longitudinally and latitudinally cut fruit images. The GS performance was assessed in cross-validations of classification-based and regression-based machine learning models for CD and TA traits, respectively. The results showed the usage of TA traits and tag SNPs provide a powerful combination to predict morphology and color-related traits of Solanaceous fruits. The highest predictability of 0.89 was achieved for fruit width in pepper, with an average predictability of 0.69 over all traits. The multi-trait GS models are of slightly better predictability than single-trait models for some colorimetric traits in pepper. While model validation performs poorly on wild tomato accessions, the usage as many as one accession per wild species in the training set can increase the transferability of models to unseen populations for some traits (e.g. fruit shape for which predictability in unseen scenario increased from zero to 0.6). Overall, GS approaches can assist the selection of high-performance Solanaceous fruits in crop breeding.
Genome-scale metabolic networks for model plants and crops in combination with approaches from the constraint-based modelling framework have been used to predict metabolic traits and design metabolic engineering strategies for their manipulation. With the advances in technologies to generate large-scale genotyping data from natural diversity panels and other populations, genome-wide association and genomic selection have emerged as statistical approaches to determine genetic variants associated with and predictive of traits. Here, we review recent advances in constraint-based approaches that integrate genetic variants in genome-scale metabolic models to characterize their effects on reaction fluxes. Since some of these approaches have been applied in organisms other than plants, we provide a critical assessment of their applicability particularly in crops. In addition, we further dissect the inferred effects of genetic variants with respect to reaction rate constants, abundances of enzymes, and concentrations of metabolites, as main determinants of reaction fluxes and relate them with their combined effects on complex traits, like growth. Through this systematic review, we also provide a roadmap for future research to increase the predictive power of statistical approaches by coupling them with mechanistic models of metabolism.
The current trends of crop yield improvements are not expected to meet the projected rise in demand. Genomic selection uses molecular markers and machine learning to identify superior genotypes with improved traits, such as growth. Plant growth directly depends on rates of metabolic reactions which transform nutrients into the building blocks of biomass. Here, we predict growth of Arabidopsis thaliana accessions by employing genomic prediction of reaction rates estimated from accession-specific metabolic models. We demonstrate that, comparing to classical genomic selection on the available data sets for 67 accessions, our approach improves the prediction accuracy for growth within and across nitrogen environments by 32.6% and 51.4%, respectively, and from optimal nitrogen to low carbon environment by 50.4%. Therefore, integration of molecular markers into metabolic models offers an approach to predict traits directly related to metabolism, and its usefulness in breeding can be examined by gathering matching datasets in crops. An increase in genomic selection (GS) accuracy can accelerate genetic gain by shortening the breeding cycles. Here, the authors introduce a network-based GS method that uses metabolic models and improves the prediction accuracy of Arabidopsis growth within and across environments.
Background: Biological systems adapt to changing environments by reorganizing their cellula r and physiological program with metabolites representing one important response level. Different stresses lead to both conserved and specific responses on the metabolite level which should be reflected in the underl ying metabolic network. Methodology/Principal Findings: Starting from experimental data obtained by a GC-MS based high-throughput metabolic profiling technology we here develop an approach that: (1) extracts network representations from metabolic conditiondependent data by using pairwise correlations, (2) determines the sets of stable and condition-dependent correlations based on a combination of statistical significance and homogeneity tests, and (3) can identify metabolites related to the stress response, which goes beyond simple ob servation s about the changes of metabolic concentrations. The approach was tested with Escherichia colias a model organism observed under four different environmental stress conditions (cold stress, heat stress, oxidative stress, lactose diau xie) and control unperturbed conditions. By constructing the stable network component, which displays a scale free topology and small-world characteristics, we demonstrated that: (1) metabolite hubs in this reconstructed correlation networks are significantly enriched for those contained in biochemical networks such as EcoCyc, (2) particular components of the stable network are enriched for functionally related biochemical path ways, and (3) ind ependently of the response scale, based on their importance in the reorganization of the cor relation network a set of metabolites can be identified which represent hypothetical candidates for adjusting to a stress-specific response. Conclusions/Significance: Network-based tools allowed the identification of stress-dependent and general metabolic correlation networks. This correlation-network-ba sed approach does not rely on major changes in concentration to identify metabolites important for st ress adaptation, but rather on the changes in network properties with respect to metabolites. This should represent a useful complementary technique in addition to more classical approaches.
Natural genetic diversity provides a powerful tool to study the complex interrelationship between metabolism and growth. Profiling of metabolic traits combined with network-based and statistical analyses allow the comparison of conditions and identification of sets of traits that predict biomass. However, it often remains unclear why a particular set of metabolites is linked with biomass and to what extent the predictive model is applicable beyond a particular growth condition. A panel of 97 genetically diverse Arabidopsis (Arabidopsis thaliana) accessions was grown in near-optimal carbon and nitrogen supply, restricted carbon supply, and restricted nitrogen supply and analyzed for biomass and 54 metabolic traits. Correlation-based metabolic networks were generated from the genotype-dependent variation in each condition to reveal sets of metabolites that show coordinated changes across accessions. The networks were largely specific for a single growth condition. Partial least squares regression from metabolic traits allowed prediction of biomass within and, slightly more weakly, across conditions (cross-validated Pearson correlations in the range of 0.27-0.58 and 0.21-0.51 and P values in the range of <0.001-<0.13 and <0.001-<0.023, respectively). Metabolic traits that correlate with growth or have a high weighting in the partial least squares regression were mainly condition specific and often related to the resource that restricts growth under that condition. Linear mixed-model analysis using the combined metabolic traits from all growth conditions as an input indicated that inclusion of random effects for the conditions improves predictions of biomass. Thus, robust prediction of biomass across a range of conditions requires condition-specific measurement of metabolic traits to take account of environment-dependent changes of the underlying networks.
Light and gravity are two key determinants in orientating plant stems for proper growth and development. The organization and dynamics of the actin cytoskeleton are essential for cell biology and critically regulated by actin-binding proteins. However, the role of actin cytoskeleton in shoot negative gravitropism remains controversial. In this work, we report that the actin-binding protein Rice Morphology Determinant (RMD) promotes reorganization of the actin cytoskeleton in rice (Oryza sativa) shoots. The changes in actin organization are associated with the ability of the rice shoots to respond to negative gravitropism. Here, light-grown rmd mutant shoots exhibited agravitropic phenotypes. By contrast, etiolated rmd shoots displayed normal negative shoot gravitropism. Furthermore, we show that RMD maintains an actin configuration that promotes statolith mobility in gravisensing endodermal cells, and for proper auxin distribution in light-grown, but not dark-grown, shoots. RMD gene expression is diurnally controlled and directly repressed by the phytochrome-interacting factor-like protein OsPIL16. Consequently, overexpression of OsPIL16 led to gravisensing and actin patterning defects that phenocopied the rmd mutant. Our findings outline a mechanism that links light signaling and gravity perception for straight shoot growth in rice.
Diatoms outcompete other phytoplankton for nitrate, yet little is known about the mechanisms underpinning this ability. Genomes and genome-enabled studies have shown that diatoms possess unique features of nitrogen metabolism however, the implications for nutrient utilization and growth are poorly understood. Using a combination of transcriptomics, proteomics, metabolomics, fluxomics, and flux balance analysis to examine short-term shifts in nitrogen utilization in the model pennate diatom in Phaeodactylum tricornutum, we obtained a systems-level understanding of assimilation and intracellular distribution of nitrogen. Chloroplasts and mitochondria are energetically integrated at the critical intersection of carbon and nitrogen metabolism in diatoms. Pathways involved in this integration are organelle-localized GS-GOGAT cycles, aspartate and alanine systems for amino moiety exchange, and a split-organelle arginine biosynthesis pathway that clarifies the role of the diatom urea cycle. This unique configuration allows diatoms to efficiently adjust to changing nitrogen status, conferring an ecological advantage over other phytoplankton taxa.
Reaction lumping in metabolic networks for application with thermodynamic metabolic flux analysis
(2021)
Thermodynamic metabolic flux analysis (TMFA) can narrow down the space of steady-state flux distributions, but requires knowledge of the standard Gibbs free energy for the modelled reactions. The latter are often not available due to unknown Gibbs free energy change of formation ,Delta fG0, of metabolites. To optimize the usage of data on thermodynamics in constraining a model, reaction lumping has been proposed to eliminate metabolites with unknown Delta fG0. However, the lumping procedure has not been formalized nor implemented for systematic identification of lumped reactions. Here, we propose, implement, and test a combined procedure for reaction lumping, applicable to genome-scale metabolic models. It is based on identification of groups of metabolites with unknown Delta fG0 whose elimination can be conducted independently of the others via: (1) group implementation, aiming to eliminate an entire such group, and, if this is infeasible, (2) a sequential implementation to ensure that a maximal number of metabolites with unknown Delta fG0 are eliminated. Our comparative analysis with genome-scale metabolic models of Escherichia coli, Bacillus subtilis, and Homo sapiens shows that the combined procedure provides an efficient means for systematic identification of lumped reactions. We also demonstrate that TMFA applied to models with reactions lumped according to the proposed procedure lead to more precise predictions in comparison to the original models. The provided implementation thus ensures the reproducibility of the findings and their application with standard TMFA.
The availability of high-throughput data from transcriptomics and metabolomics technologies provides the opportunity to characterize the transcriptional effects on metabolism. Here we propose and evaluate two computational approaches rooted in data reduction techniques to identify and categorize transcriptional effects on metabolism by combining data on gene expression and metabolite levels. The approaches determine the partial correlation between two metabolite data profiles upon control of given principal components extracted from transcriptomics data profiles. Therefore, they allow us to investigate both data types with all features simultaneously without doing preselection of genes. The proposed approaches allow us to categorize the relation between pairs of metabolites as being under transcriptional or post-transcriptional regulation. The resulting classification is compared to existing literature and accumulated evidence about regulatory mechanism of reactions and pathways in the cases of Escherichia coil, Saccharomycies cerevisiae, and Arabidopsis thaliana.
Stoichiometric Correlation Analysis: Principles of Metabolic Functionality from Metabolomics Data
(2017)
Recent advances in metabolomics technologies have resulted in high-quality (time-resolved) metabolic profiles with an increasing coverage of metabolic pathways. These data profiles represent read-outs from often non-linear dynamics of metabolic networks. Yet, metabolic profiles have largely been explored with regression-based approaches that only capture linear relationships, rendering it difficult to determine the extent to which the data reflect the underlying reaction rates and their couplings. Here we propose an approach termed Stoichiometric Correlation Analysis (SCA) based on correlation between positive linear combinations of log-transformed metabolic profiles. The log-transformation is due to the evidence that metabolic networks can be modeled by mass action law and kinetics derived from it. Unlike the existing approaches which establish a relation between pairs of metabolites, SCA facilitates the discovery of higherorder dependence between more than two metabolites. By using a paradigmatic model of the tricarboxylic acid cycle we show that the higher-order dependence reflects the coupling of concentration of reactant complexes, capturing the subtle difference between the employed enzyme kinetics. Using time-resolved metabolic profiles from Arabidopsis thaliana and Escherichia coli, we show that SCA can be used to quantify the difference in coupling of reactant complexes, and hence, reaction rates, underlying the stringent response in these model organisms. By using SCA with data from natural variation of wild and domesticated wheat and tomato accession, we demonstrate that the domestication is accompanied by loss of such couplings, in these species. Therefore, application of SCA to metabolomics data from natural variation in wild and domesticated populations provides a mechanistic way to understanding domestication and its relation to metabolic networks.
Plant organs consist of multiple cell types that do not operate in isolation, but communicate with each other to maintain proper functions. Here, we extract models specific to three developmental stages of eight root cell types or tissue layers in Arabidopsis thaliana based on a state-of-the-art constraint-based modeling approach with all publicly available transcriptomics and metabolomics data from this system to date. We integrate these models into a multi-cell root model which we investigate with respect to network structure, distribution of fluxes, and concordance to transcriptomics and proteomics data. From a methodological point, we show that the coupling of tissue-specific models in a multi-tissue model yields a higher specificity of the interconnected models with respect to network structure and flux distributions. We use the extracted models to predict and investigate the flux of the growth hormone indole-3-actetate and its antagonist, trans-Zeatin, through the root. While some of predictions are in line with experimental evidence, constraints other than those coming from the metabolic level may be necessary to replicate the flow of indole-3-actetate from other simulation studies. Therefore, our work provides the means for data-driven multi-tissue metabolic model extraction of other Arabidopsis organs in the constraint-based modeling framework.
Large-scale co-expression approach to dissect secondary cell wall formation across plant species
(2011)
Plant cell walls are complex composites largely consisting of carbohydrate-based polymers, and are generally divided into primary and secondary walls based on content and characteristics. Cellulose microfibrils constitute a major component of both primary and secondary cell walls and are synthesized at the plasma membrane by cellulose synthase (CESA) complexes. Several studies in Arabidopsis have demonstrated the power of co-expression analyses to identify new genes associated with secondary wall cellulose biosynthesis. However, across-species comparative co-expression analyses remain largely unexplored. Here, we compared co-expressed gene vicinity networks of primary and secondary wall CESAsin Arabidopsis, barley, rice, poplar, soybean, Medicago, and wheat, and identified gene families that are consistently co-regulated with cellulose biosynthesis. In addition to the expected polysaccharide acting enzymes, we also found many gene families associated with cytoskeleton, signaling, transcriptional regulation, oxidation, and protein degradation. Based on these analyses, we selected and biochemically analyzed T-DNA insertion lines corresponding to approximately twenty genes from gene families that re-occur in the co-expressed gene vicinity networks of secondary wall CESAs across the seven species. We developed a statistical pipeline using principal component analysis and optimal clustering based on silhouette width to analyze sugar profiles. One of the mutants, corresponding to a pinoresinol reductase gene, displayed disturbed xylem morphology and held lower levels of lignin molecules. We propose that this type of large-scale co-expression approach, coupled with statistical analysis of the cell wall contents, will be useful to facilitate rapid knowledge transfer across plant species.
Whole-genome duplications (WGDs) or polyploidy events have been studied extensively in plants. In a now widely cited paper, Jiao et al. presented evidence for two ancient, ancestral plant WGDs predating the origin of flowering and seed plants, respectively. This finding was based primarily on a bimodal age distribution of gene duplication events obtained from molecular dating of almost 800 phylogenetic gene trees. We reanalyzed the phylogenomic data of Jiao et al. and found that the strong bimodality of the age distribution may be the result of technical and methodological issues and may hence not be a "true" signal of two WGD events. By using a state-of-the-art molecular dating algorithm, we demonstrate that the reported bimodal age distribution is not robust and should be interpreted with caution. Thus, there exists little evidence for two ancient WGDs in plants from phylogenomic dating.
Metabolism is a key determinant of plant growth and modulates plant adaptive responses. Increased metabolic variation due to heterozygosity may be beneficial for highly homozygous plants if their progeny is to respond to sudden changes in the habitat. Here, we investigate the extent to which heterozygosity contributes to the variation in metabolism and size of hybrids of Arabidopsis thaliana whose parents are from a single growth habitat. We created full diallel crosses among seven parents, originating from Southern Germany, and analysed the inheritance patterns in primary and secondary metabolism as well as in rosette size in situ. In comparison to primary metabolites, compounds from secondary metabolism were more variable and showed more pronounced non-additive inheritance patterns which could be attributed to epistasis. In addition, we showed that glucosinolates, among other secondary metabolites, were positively correlated with a proxy for plant size. Therefore, our study demonstrates that heterozygosity in local A. thaliana population generates metabolic variation and may impact several tasks directly linked to metabolism.
The integration of experimental data into genome-scale metabolic models can greatly improve flux predictions. This is achieved by restricting predictions to a more realistic context-specific domain, like a particular cell or tissue type. Several computational approaches to integrate data have been proposed D generally obtaining context-specific (sub) models or flux distributions. However, these approaches may lead to a multitude of equally valid but potentially different models or flux distributions, due to possible alternative optima in the underlying optimization problems. Although this issue introduces ambiguity in context-specific predictions, it has not been generally recognized, especially in the case of model reconstructions. In this study, we analyze the impact of alternative optima in four state-of-the-art context-specific data integration approaches, providing both flux distributions and/or metabolic models. To this end, we present three computational methods and apply them to two particular case studies: leaf-specific predictions from the integration of gene expression data in a metabolic model of Arabidopsis thaliana, and liver-specific reconstructions derived from a human model with various experimental data sources. The application of these methods allows us to obtain the following results: (i) we sample the space of alternative flux distributions in the leaf-and the liver-specific case and quantify the ambiguity of the predictions. In addition, we show how the inclusion of l(1)-regularization during data integration reduces the ambiguity in both cases. (ii) We generate sets of alternative leaf-and liver-specific models that are optimal to each one of the evaluated model reconstruction approaches. We demonstrate that alternative models of the same context contain a marked fraction of disparate reactions. Further, we show that a careful balance between model sparsity and metabolic functionality helps in reducing the discrepancies between alternative models. Finally, our findings indicate that alternative optima must be taken into account for rendering the context-specific metabolic model predictions less ambiguous.
Photosynthesis and water use efficiency, key factors affecting plant growth, are directly controlled by microscopic and adjustable pores in the leaf-the stomata. The size of the pores is modulated by the guard cells, which rely on molecular mechanisms to sense and respond to environmental changes. It has been shown that the physiology of mesophyll and guard cells differs substantially. However, the implications of these differences to metabolism at a genome-scale level remain unclear. Here, we used constraint-based modeling to predict the differences in metabolic fluxes between the mesophyll and guard cells of Arabidopsis thaliana by exploring the space of fluxes that are most concordant to cell-type-specific transcript profiles. An independent C-13-labeling experiment using isolated mesophyll and guard cells was conducted and provided support for our predictions about the role of the Calvin-Benson cycle in sucrose synthesis in guard cells. The combination of in silico with in vivo analyses indicated that guard cells have higher anaplerotic CO2 fixation via phosphoenolpyruvate carboxylase, which was demonstrated to be an important source of malate. Beyond highlighting the metabolic differences between mesophyll and guard cells, our findings can be used in future integrated modeling of multicellular plant systems and their engineering towards improved growth.
Characterisation of gene-regulatory network (GRN) interactions provides a stepping stone to understanding how genes affect cellular phenotypes. Yet, despite advances in profiling technologies, GRN reconstruction from gene expression data remains a pressing problem in systems biology. Here, we devise a supervised learning approach, GRADIS, which utilises support vector machine to reconstruct GRNs based on distance profiles obtained from a graph representation of transcriptomics data. By employing the data fromEscherichia coliandSaccharomyces cerevisiaeas well as synthetic networks from the DREAM4 and five network inference challenges, we demonstrate that our GRADIS approach outperforms the state-of-the-art supervised and unsupervided approaches. This holds when predictions about target genes for individual transcription factors as well as for the entire network are considered. We employ experimentally verified GRNs fromE. coliandS. cerevisiaeto validate the predictions and obtain further insights in the performance of the proposed approach. Our GRADIS approach offers the possibility for usage of other network-based representations of large-scale data, and can be readily extended to help the characterisation of other cellular networks, including protein-protein and protein-metabolite interactions.
GeneReg
(2020)
Motivation
Large-scale metabolic models are widely used to design metabolic engineering strategies for diverse biotechnological applications. However, the existing computational approaches focus on alteration of reaction fluxes and often neglect the manipulations of gene expression to implement these strategies.
Results
Here, we find that the association of genes with multiple reactions leads to infeasibility of engineering strategies at the flux level, since they require contradicting manipulations of gene expression. Moreover, we identify that all of the existing approaches to design gene knockout strategies do not ensure that the resulting design may also require other gene alterations, such as up- or downregulations, to match the desired flux distribution. To address these issues, we propose a constraint-based approach, termed GeneReg, that facilitates the design of feasible metabolic engineering strategies at the gene level and that is readily applicable to large-scale metabolic networks. We show that GeneReg can identify feasible strategies to overproduce ethanol in Escherichia coli and lactate in Saccharomyces cerevisiae, but overproduction of the TCA cycle intermediates is not feasible in five organisms used as cell factories under default growth conditions. Therefore, GeneReg points at the need to couple gene regulation and metabolism to design rational metabolic engineering strategies.
Identifying the entirety of gene regulatory interactions in a biological system offers the possibility to determine the key molecular factors that affect important traits on the level of cells, tissues, and whole organisms. Despite the development of experimental approaches and technologies for identification of direct binding of transcription factors (TFs) to promoter regions of downstream target genes, computational approaches that utilize large compendia of transcriptomics data are still the predominant methods used to predict direct downstream targets of TFs, and thus reconstruct genome-wide gene-regulatory networks (GRNs). These approaches can broadly be categorized into unsupervised and supervised, based on whether data about known, experimentally verified gene-regulatory interactions are used in the process of reconstructing the underlying GRN. Here, we first describe the generic steps of supervised approaches for GRN reconstruction, since they have been recently shown to result in improved accuracy of the resulting networks? We also illustrate how they can be used with data from model organisms to obtain more accurate prediction of gene regulatory interactions.
Ribosome biogenesis is tightly associated to plant metabolism due to the usage of ribosomes in the synthesis of proteins necessary to drive metabolic pathways. Given the central role of ribosome biogenesis in cell physiology, it is important to characterize the impact of different components involved in this process on plant metabolism. Double mutants of the Arabidopsis thaliana cytosolic 60S maturation factors REIL1 and REIL2 do not resume growth after shift to moderate 10 degrees C chilling conditions. To gain mechanistic insights into the metabolic effects of this ribosome biogenesis defect on metabolism, we developed TC-iReMet2, a constraint-based modelling approach that integrates relative metabolomics and transcriptomics time-course data to predict differential fluxes on a genome-scale level. We employed TC-iReMet2 with metabolomics and transcriptomics data from the Arabidopsis Columbia 0 wild type and the reil1-1 reil2-1 double mutant before and after cold shift. We identified reactions and pathways that are highly altered in a mutant relative to the wild type. These pathways include the Calvin-Benson cycle, photorespiration, gluconeogenesis, and glycolysis. Our findings also indicated differential NAD(P)/NAD(P)H ratios after cold shift. TC-iReMet2 allows for mechanistic hypothesis generation and interpretation of system biology experiments related to metabolic fluxes on a genome-scale level.
Plasticity in metabolism underpins local responses to nitrogen in Arabidopsis thaliana populations
(2019)
Nitrogen (N) is central for plant growth, and metabolic plasticity can provide a strategy to respond to changing N availability. We showed that two local A. thaliana populations exhibited differential plasticity in the compounds of photorespiratory and starch degradation pathways in response to three N conditions. Association of metabolite levels with growth-related and fitness traits indicated that controlled plasticity in these pathways could contribute to local adaptation and play a role in plant evolution.
Physically interacting proteins form macromolecule complexes that drive diverse cellular processes. Advances in experimental techniques that capture interactions between proteins provide us with protein-protein interaction (PPI) networks from several model organisms. These datasets have enabled the prediction and other computational analyses of protein complexes. Here we provide a systematic review of the state-of-the-art algorithms for protein complex prediction from PPI networks proposed in the past two decades. The existing approaches that solve this problem are categorized into three groups, including: cluster-quality-based, node affinity-based, and network embedding-based approaches, and we compare and contrast the advantages and disadvantages. We further include a comparative analysis by computing the performance of eighteen methods based on twelve well-established performance measures on four widely used benchmark protein-protein interaction networks. Finally, the limitations and drawbacks of both, current data and approaches, along with the potential solutions in this field are discussed, with emphasis on the points that pave the way for future research efforts in this field. (c) 2022 The Author(s). Published by Elsevier B.V. on behalf of Research Network of Computational and Structural Biotechnology. This is an open access article under the CC BY license (http://creativecommons. org/licenses/by/4.0/).
High-throughput proteomics approaches have resulted in large-scale protein–protein interaction (PPI) networks that have been employed for the prediction of protein complexes. However, PPI networks contain false-positive as well as false-negative PPIs that affect the protein complex prediction algorithms. To address this issue, here we propose an algorithm called CUBCO+ that: (1) employs GO semantic similarity to retain only biologically relevant interactions with a high similarity score, (2) based on link prediction approaches, scores the false-negative edges, and (3) incorporates the resulting scores to predict protein complexes. Through comprehensive analyses with PPIs from Escherichia coli, Saccharomyces cerevisiae, and Homo sapiens, we show that CUBCO+ performs as well as the approaches that predict protein complexes based on recently introduced graph partitions into biclique spanned subgraphs and outperforms the other state-of-the-art approaches. Moreover, we illustrate that in combination with GO semantic similarity, CUBCO+ enables us to predict more accurate protein complexes in 36% of the cases in comparison to CUBCO as its predecessor.
High-throughput proteomics approaches have resulted in large-scale protein–protein interaction (PPI) networks that have been employed for the prediction of protein complexes. However, PPI networks contain false-positive as well as false-negative PPIs that affect the protein complex prediction algorithms. To address this issue, here we propose an algorithm called CUBCO+ that: (1) employs GO semantic similarity to retain only biologically relevant interactions with a high similarity score, (2) based on link prediction approaches, scores the false-negative edges, and (3) incorporates the resulting scores to predict protein complexes. Through comprehensive analyses with PPIs from Escherichia coli, Saccharomyces cerevisiae, and Homo sapiens, we show that CUBCO+ performs as well as the approaches that predict protein complexes based on recently introduced graph partitions into biclique spanned subgraphs and outperforms the other state-of-the-art approaches. Moreover, we illustrate that in combination with GO semantic similarity, CUBCO+ enables us to predict more accurate protein complexes in 36% of the cases in comparison to CUBCO as its predecessor.
Identification of protein complexes from protein-protein interaction (PPI) networks is a key problem in PPI mining, solved by parameter-dependent approaches that suffer from small recall rates. Here we introduce GCC-v, a family of efficient, parameter-free algorithms to accurately predict protein complexes using the (weighted) clustering coefficient of proteins in PPI networks. Through comparative analyses with gold standards and PPI networks from Escherichia coli, Saccharomyces cerevisiae, and Homo sapiens, we demonstrate that GCC-v outperforms twelve state-of-the-art approaches for identification of protein complexes with respect to twelve performance measures in at least 85.71% of scenarios. We also show that GCC-v results in the exact recovery of similar to 35% of protein complexes in a pan-plant PPI network and discover 144 new protein complexes in Arabidopsis thaliana, with high support from GO semantic similarity. Our results indicate that findings from GCC-v are robust to network perturbations, which has direct implications to assess the impact of the PPI network quality on the predicted protein complexes. (C) 2021 The Author(s). Published by Elsevier B.V. on behalf of Research Network of Computational and Structural Biotechnology.
PC2P
(2021)
Motivation:
Prediction of protein complexes from protein-protein interaction (PPI) networks is an important problem in systems biology, as they control different cellular functions. The existing solutions employ algorithms for network community detection that identify dense subgraphs in PPI networks. However, gold standards in yeast and human indicate that protein complexes can also induce sparse subgraphs, introducing further challenges in protein complex prediction.
Results:
To address this issue, we formalize protein complexes as biclique spanned subgraphs, which include both sparse and dense subgraphs. We then cast the problem of protein complex prediction as a network partitioning into biclique spanned subgraphs with removal of minimum number of edges, called coherent partition. Since finding a coherent partition is a computationally intractable problem, we devise a parameter-free greedy approximation algorithm, termed Protein Complexes from Coherent Partition (PC2P), based on key properties of biclique spanned subgraphs. Through comparison with nine contenders, we demonstrate that PC2P: (i) successfully identifies modular structure in networks, as a prerequisite for protein complex prediction, (ii) outperforms the existing solutions with respect to a composite score of five performance measures on 75% and 100% of the analyzed PPI networks and gold standards in yeast and human, respectively, and (iii,iv) does not compromise GO semantic similarity and enrichment score of the predicted protein complexes. Therefore, our study demonstrates that clustering of networks in terms of biclique spanned subgraphs is a promising framework for detection of complexes in PPI networks.
Time-series data from multicomponent systems capture the dynamics of the ongoing processes and reflect the interactions between the components. The progression of processes in such systems usually involves check-points and events at which the relationships between the components are altered in response to stimuli. Detecting these events together with the implicated components can help understand the temporal aspects of complex biological systems. Here we propose a regularized regression-based approach for identifying breakpoints and corresponding segments from multivariate time-series data. In combination with techniques from clustering, the approach also allows estimating the significance of the determined breakpoints as well as the key components implicated in the emergence of the breakpoints. Comparative analysis with the existing alternatives demonstrates the power of the approach to identify biologically meaningful breakpoints in diverse time-resolved transcriptomics data sets from the yeast Saccharomyces cerevisiae and the diatom Thalassiosira pseudonana.
The levels of cellular organization, from gene transcription to translation to protein-protein interaction and metabolism, operate via tightly regulated mutual interactions, facilitating organismal adaptability and various stress responses. Characterizing the mutual interactions between genes, transcription factors, and proteins involved in signaling, termed crosstalk, is therefore crucial for understanding and controlling cells' functionality. We aim at using high-throughput transcriptomics data to discover previously unknown links between signaling networks. We propose and analyze a novel method for crosstalk identification which relies on transcriptomics data and overcomes the lack of complete information for signaling pathways in Arabidopsis thaliana. Our method first employs a network-based transformation of the results from the statistical analysis of differential gene expression in given groups of experiments under different signal-inducing conditions. The stationary distribution of a random walk (similar to the PageRank algorithm) on the constructed network is then used to determine the putative transcripts interrelating different signaling pathways. With the help of the proposed method, we analyze a transcriptomics data set including experiments from four different stresses/signals: nitrate, sulfur, iron, and hormones. We identified promising gene candidates, downstream of the transcription factors (TFs), associated to signaling crosstalk, which were validated through literature mining. In addition, we conduct a comparative analysis with the only other available method in this field which used a biclustering-based approach. Surprisingly, the biclustering-based approach fails to robustly identify any candidate genes involved in the crosstalk of the analyzed signals. We demonstrate that our proposed method is more robust in identifying gene candidates involved downstream of the signaling crosstalk for species for which large transcriptomics data sets, normalized with the same techniques, are available. Moreover, unlike approaches based on biclustering, our approach does not rely on any hidden parameters.
Molecular phenotyping technologies (e.g., transcriptomics, proteomics, and metabolomics) offer the possibility to simultaneously obtain multivariate time series (MTS) data from different levels of information processing and metabolic conversions in biological systems. As a result, MTS data capture the dynamics of biochemical processes and components whose couplings may involve different scales and exhibit temporal changes. Therefore, it is important to develop methods for determining the time segments in MTS data, which may correspond to critical biochemical events reflected in the coupling of the system's components. Here we provide a novel network-based formalization of the MTS segmentation problem based on temporal dependencies and the covariance structure of the data. We demonstrate that the problem of partitioning MTS data into k segments to maximize a distance function, operating on polynomially computable network properties, often used in analysis of biological network, can be efficiently solved. To enable biological interpretation, we also propose a breakpoint-penalty (BP-penalty) formulation for determining MTS segmentation which combines a distance function with the number/length of segments. Our empirical analyses of synthetic benchmark data as well as time-resolved transcriptomics data from the metabolic and cell cycles of Saccharomyces cerevisiae demonstrate that the proposed method accurately infers the phases in the temporal compartmentalization of biological processes. In addition, through comparison on the same data sets, we show that the results from the proposed formalization of the MTS segmentation problem match biological knowledge and provide more rigorous statistical support in comparison to the contending state-of-the-art methods.
Recent analyses have demonstrated that plant metabolic networks do not differ in their structural properties and that genes involved in basic metabolic processes show smaller coexpression than genes involved in specialized metabolism. By contrast, our analysis reveals differences in the structure of plant metabolic networks and patterns of coexpression for genes in (non)specialized metabolism. Here we caution that conclusions concerning the organization of plant metabolism based on network-driven analyses strongly depend on the computational approaches used.
Devising computational methods to accurately reconstruct gene regulatory networks given gene expression data is key to systems biology applications. Here we propose a method for reconstructing gene regulatory networks by simultaneous consideration of data sets from different perturbation experiments and corresponding controls. The method imposes three biologically meaningful constraints: (1) expression levels of each gene should be explained by the expression levels of a small number of transcription factor coding genes, (2) networks inferred from different data sets should be similar with respect to the type and number of regulatory interactions, and (3) relationships between genes which exhibit similar differential behavior over the considered perturbations should be favored. We demonstrate that these constraints can be transformed in a fused LASSO formulation for the proposed method. The comparative analysis on transcriptomics time-series data from prokaryotic species, Escherichia coli and Mycobacterium tuberculosis, as well as a eukaryotic species, mouse, demonstrated that the proposed method has the advantages of the most recent approaches for regulatory network inference, while obtaining better performance and assigning higher scores to the true regulatory links. The study indicates that the combination of sparse regression techniques with other biologically meaningful constraints is a promising framework for gene regulatory network reconstructions.
Abiotic stresses cause oxidative damage in plants. Here, we demonstrate that foliar application of an extract from the seaweed Ascophyllum nodosum, SuperFifty (SF), largely prevents paraquat (PQ)-induced oxidative stress in Arabidopsis thaliana. While PQ-stressed plants develop necrotic lesions, plants pre-treated with SF (i.e., primed plants) were unaffected by PQ. Transcriptome analysis revealed induction of reactive oxygen species (ROS) marker genes, genes involved in ROS-induced programmed cell death, and autophagy-related genes after PQ treatment. These changes did not occur in PQ-stressed plants primed with SF. In contrast, upregulation of several carbohydrate metabolism genes, growth, and hormone signaling as well as antioxidant-related genes were specific to SF-primed plants. Metabolomic analyses revealed accumulation of the stress-protective metabolite maltose and the tricarboxylic acid cycle intermediates fumarate and malate in SF-primed plants. Lipidome analysis indicated that those lipids associated with oxidative stress-induced cell death and chloroplast degradation, such as triacylglycerols (TAGs), declined upon SF priming. Our study demonstrated that SF confers tolerance to PQ-induced oxidative stress in A. thaliana, an effect achieved by modulating a range of processes at the transcriptomic, metabolic, and lipid levels.
Abiotic stresses cause oxidative damage in plants. Here, we demonstrate that foliar application of an extract from the seaweed Ascophyllum nodosum, SuperFifty (SF), largely prevents paraquat (PQ)-induced oxidative stress in Arabidopsis thaliana. While PQ-stressed plants develop necrotic lesions, plants pre-treated with SF (i.e., primed plants) were unaffected by PQ. Transcriptome analysis revealed induction of reactive oxygen species (ROS) marker genes, genes involved in ROS-induced programmed cell death, and autophagy-related genes after PQ treatment. These changes did not occur in PQ-stressed plants primed with SF. In contrast, upregulation of several carbohydrate metabolism genes, growth, and hormone signaling as well as antioxidant-related genes were specific to SF-primed plants. Metabolomic analyses revealed accumulation of the stress-protective metabolite maltose and the tricarboxylic acid cycle intermediates fumarate and malate in SF-primed plants. Lipidome analysis indicated that those lipids associated with oxidative stress-induced cell death and chloroplast degradation, such as triacylglycerols (TAGs), declined upon SF priming. Our study demonstrated that SF confers tolerance to PQ-induced oxidative stress in A. thaliana, an effect achieved by modulating a range of processes at the transcriptomic, metabolic, and lipid levels.
IntroductionTo date, most studies of natural variation and metabolite quantitative trait loci (mQTL) in tomato have focused on fruit metabolism, leaving aside the identification of genomic regions involved in the regulation of leaf metabolism.ObjectiveThis study was conducted to identify leaf mQTL in tomato and to assess the association of leaf metabolites and physiological traits with the metabolite levels from other tissues.MethodsThe analysis of components of leaf metabolism was performed by phenotypying 76 tomato ILs with chromosome segments of the wild species Solanum pennellii in the genetic background of a cultivated tomato (S. lycopersicum) variety M82. The plants were cultivated in two different environments in independent years and samples were harvested from mature leaves of non-flowering plants at the middle of the light period. The non-targeted metabolite profiling was obtained by gas chromatography time-of-flight mass spectrometry (GC-TOF-MS). With the data set obtained in this study and already published metabolomics data from seed and fruit, we performed QTL mapping, heritability and correlation analyses.ResultsChanges in metabolite contents were evident in the ILs that are potentially important with respect to stress responses and plant physiology. By analyzing the obtained data, we identified 42 positive and 76 negative mQTL involved in carbon and nitrogen metabolism.ConclusionsOverall, these findings allowed the identification of S. lycopersicum genome regions involved in the regulation of leaf primary carbon and nitrogen metabolism, as well as the association of leaf metabolites with metabolites from seeds and fruits.
CytoSeg 2.0
(2020)
Motivation:
Actin filaments (AFs) are dynamic structures that substantially change their organization over time. The dynamic behavior and the relatively low signal-to-noise ratio during live-cell imaging have rendered the quantification of the actin organization a difficult task.
Results:
We developed an automated image-based framework that extracts AFs from fluorescence microscopy images and represents them as networks, which are automatically analyzed to identify and compare biologically relevant features. Although the source code is freely available, we have now implemented the framework into a graphical user interface that can be installed as a Fiji plugin, thus enabling easy access by the research community.
Integration of high-throughput data with functional annotation by graph-theoretic methods has been postulated as promising way to unravel the function of unannotated genes. Here, we first review the existing graph-theoretic approaches for automated gene function annotation and classify them into two categories with respect to their relation to two instances of transductive learning on networks - with dynamic costs and with constant costs - depending on whether or not ontological relationship between functional terms is employed. The determined categories allow to characterize the computational complexity of the existing approaches and establish the relation to classical graph-theoretic problems, such as bisection and multiway cut. In addition, our results point out that the ontological form of the structured functional knowledge does not lower the complexity of the transductive learning with dynamic costs - one of the key problems in modern systems biology. The NP-hardness of automated gene annotation renders the development of heuristic or approximation algorithms a priority for additional research.
Robustness of biochemical systems has become one of the central questions in systems biology although it is notoriously difficult to formally capture its multifaceted nature. Maintenance of normal system function depends not only on the stoichiometry of the underlying interrelated components, but also on the multitude of kinetic parameters. Invariant flux ratios, obtained within flux coupling analysis, as well as invariant complex ratios, derived within chemical reaction network theory, can characterize robust properties of a system at steady state. However, the existing formalisms for the description of these invariants do not provide full characterization as they either only focus on the flux-centric or the concentration-centric view. Here we develop a novel mathematical framework which combines both views and thereby overcomes the limitations of the classical methodologies. Our unified framework will be helpful in analyzing biologically important system properties.
We investigated the systems response of metabolism and growth after an increase in irradiance in the nonsaturating range in the algal model Chlamydomonas reinhardtii. In a three-step process, photosynthesis and the levels of metabolites increased immediately, growth increased after 10 to 15 min, and transcript and protein abundance responded by 40 and 120 to 240 min, respectively. In the first phase, starch and metabolites provided a transient buffer for carbon until growth increased. This uncouples photosynthesis from growth in a fluctuating light environment. In the first and second phases, rising metabolite levels and increased polysome loading drove an increase in fluxes. Most Calvin-Benson cycle (CBC) enzymes were substrate-limited in vivo, and strikingly, many were present at higher concentrations than their substrates, explaining how rising metabolite levels stimulate CBC flux. Rubisco, fructose-1,6-biosphosphatase, and seduheptulose-1,7-bisphosphatase were close to substrate saturation in vivo, and flux was increased by posttranslational activation. In the third phase, changes in abundance of particular proteins, including increases in plastidial ATP synthase and some CBC enzymes, relieved potential bottlenecks and readjusted protein allocation between different processes. Despite reasonable overall agreement between changes in transcript and protein abundance (R-2 = 0.24), many proteins, including those in photosynthesis, changed independently of transcript abundance.
L-2,L-1-norm regularized multivariate regression model with applications to genomic prediction
(2021)
Motivation:
Genomic selection (GS) is currently deemed the most effective approach to speed up breeding of agricultural varieties. It has been recognized that consideration of multiple traits in GS can improve accuracy of prediction for traits of low heritability. However, since GS forgoes statistical testing with the idea of improving predictions, it does not facilitate mechanistic understanding of the contribution of particular single nucleotide polymorphisms (SNP).
Results:
Here, we propose a L-2,L-1-norm regularized multivariate regression model and devise a fast and efficient iterative optimization algorithm, called L-2,L-1-joint, applicable in multi-trait GS. The usage of the L-2,L-1-norm facilitates variable selection in a penalized multivariate regression that considers the relation between individuals, when the number of SNPs is much larger than the number of individuals. The capacity for variable selection allows us to define master regulators that can be used in a multi-trait GS setting to dissect the genetic architecture of the analyzed traits. Our comparative analyses demonstrate that the proposed model is a favorable candidate compared to existing state-of-the-art approaches. Prediction and variable selection with datasets from Brassica napus, wheat and Arabidopsis thaliana diversity panels are conducted to further showcase the performance of the proposed model.
Genomic prediction has revolutionized crop breeding despite remaining issues of transferability of models to unseen environmental conditions and environments. Usage of endophenotypes rather than genomic markers leads to the possibility of building phenomic prediction models that can account, in part, for this challenge. Here, we compare and contrast genomic prediction and phenomic prediction models for 3 growth-related traits, namely, leaf count, tree height, and trunk diameter, from 2 coffee 3-way hybrid populations exposed to a series of treatment-inducing environmental conditions. The models are based on 7 different statistical methods built with genomic markers and ChlF data used as predictors. This comparative analysis demonstrates that the best-performing phenomic prediction models show higher predictability than the best genomic prediction models for the considered traits and environments in the vast majority of comparisons within 3-way hybrid populations. In addition, we show that phenomic prediction models are transferrable between conditions but to a lower extent between populations and we conclude that chlorophyll a fluorescence data can serve as alternative predictors in statistical models of coffee hybrid performance. Future directions will explore their combination with other endophenotypes to further improve the prediction of growth-related traits for crops.
The reactive oxygen species (ROS) gene network, consisting of both ROS-generating and detoxifying enzymes, adjusts ROS levels in response to various stimuli. We performed a cross-kingdom comparison of ROS gene networks to investigate how they have evolved across all Eukaryotes, including protists, fungi, plants and animals. We included the genomes of 16 extremotolerant Eukaryotes to gain insight into ROS gene evolution in organisms that experience extreme stress conditions. Our analysis focused on ROS genes found in all Eukaryotes (such as catalases, superoxide dismutases, glutathione reductases, peroxidases and glutathione peroxidase/peroxiredoxins) as well as those specific to certain groups, such as ascorbate peroxidases, dehydroascorbate/monodehydroascorbate reductases in plants and other photosynthetic organisms. ROS-producing NADPH oxidases (NOX) were found in most multicellular organisms, although several NOX-like genes were identified in unicellular or filamentous species. However, despite the extreme conditions experienced by extremophile species, we found no evidence for expansion of ROS-related gene families in these species compared to other Eukaryotes. Tardigrades and rotifers do show ROS gene expansions that could be related to their extreme lifestyles, although a high rate of lineage-specific horizontal gene transfer events, coupled with recent tetraploidy in rotifers, could explain this observation. This suggests that the basal Eukaryotic ROS scavenging systems are sufficient to maintain ROS homeostasis even under the most extreme conditions.