UFMG, Brazil
Hypertrophic cardiomyopathy (HCM) is the most prevalent inherited cardiac disease, affecting approximately 1 in 500 individuals and representing a leading cause of sudden cardiac death in young people. Identifying genes whose expression reliably distinguishes HCM myocardium from healthy tissue remains a central challenge in cardiovascular genomics. Traditional approaches rely on univariate statistical tests applied independently to each gene, which may overlook coordinated expression changes distributed across multiple loci. Here we present a supervised gene selection method based on the solution of a regularized block linear system that simultaneously estimates a discriminant weight for every gene. The formulation is algebraically equivalent to Tikhonov-regularized least squares (ridge regression, λ = 1), but is solved in sparse block form, which preserves numerical stability in the strongly underdetermined regime typical of transcriptomics and avoids explicitly forming the dense Gram matrix. We applied the method to microarray expression profiles from 106 HCM patients and 39 healthy donor hearts (GSE36961, Illumina HumanHT-12 v4). Each of the 37,846 probes received a scalar weight (α); the 20 genes with the largest positive and the 20 with the most negative weights were retained as the discriminant panel. Singular value decomposition was then used to assess whether this reduced panel captures the phenotypic structure of the data. Principal component analysis restricted to the 40 selected genes revealed clear separation between HCM and control samples (PC1: 41.9% of variance; PC2: 15.0%), whereas the same projection computed from all probes showed no discernible grouping (PC1: 38.8%; PC2: 9.0%). The selected panel includes well-established HCM-associated sarcomeric genes (MYH6, MYH7, TNNC1, TNNI3, ACTA1, TCAP, MYOM1, PDLIM3), recovering the classical alpha-to-beta myosin heavy chain switch as the two strongest individual signals, together with the matricellular stress protein THBS4 and multiple mitochondrial complex I subunits (NDUFS4, NDUFB8). These results demonstrate that a concise algebraic formulation can recover biologically meaningful gene signatures from high-dimensional expression data without prior biological knowledge, offering a complementary perspective to conventional differential expression analyses.
Vinicius Toledo holds a degree in Mechanical Engineering from Federal University of Minas Gerais. His research applies numerical linear algebra and machine learning to high-dimensional biological data, with a current focus on supervised feature selection in cardiac transcriptomics. His work investigates how regularized algebraic formulations can recover biologically interpretable gene signatures from microarray and RNA-seq datasets.
© 2026 Mathews International LLC. All rights reserved.