Publications-Proceedings

Article View/Open

Publication Export

Google ScholarTM

NCCU Library

Citation Infomation

Related Publications in TAIR

題名 scGHSOM: Hierarchical clustering and visualization of single-cell and CRISPR data using growing hierarchical SOM
作者 郁方
Yu, Fang;Wen, Shang-Jung;Chang, Jia-Ming
貢獻者 資管系
關鍵詞 unsupervised clustering; self-organizing map; growing hierarchical self-organizing map; single-cell; cluster feature map; cluster distribution map; scRNAseq
日期 2024-08
上傳時間 4-Oct-2024 09:21:34 (UTC+8)
摘要 High-dimensional single-cell data poses significant challenges in identifying underlying biological patterns due to the complexity and heterogeneity of cellular states. We propose a comprehensive gene-cell dependency visualization via unsupervised clustering, Growing Hierarchical Self-Organizing Map (GHSOM), specifically designed for analyzing high-dimensional single-cell data like single-cell sequencing and CRISPR screens. GHSOM is applied to cluster samples in a hierarchical structure such that the self-growth structure of clusters satisfies the required variations between and within. We propose a novel Significant Attributes Identification Algorithm to identify features that distinguish clusters. This algorithm pinpoints attributes with minimal variation within a cluster but substantial variation between clusters. These key attributes can then be used for targeted data retrieval and downstream analysis. Furthermore, we present two innovative visualization tools: Cluster Feature Map and Cluster Distribution Map. The Cluster Feature Map highlights the distribution of specific features across the hierarchical structure of GHSOM clusters. This allows for rapid visual assessment of cluster uniqueness based on chosen features. The Cluster Distribution Map depicts leaf clusters as circles on the GHSOM grid, with circle size reflecting cluster data size and color customizable to visualize features like cell type or other attributes. We apply our analysis to three single-cell datasets and one CRISPR dataset (cell-gene database) and evaluate clustering methods with internal and external CH and ARI scores. GHSOM performs well, being the best performer in internal evaluation (CH=4.2). In external evaluation, GHSOM has the third-best performance of all methods.
關聯 23rd International Workshop on Data Mining in Bioinformatics (In conjuction with ACM KDD 2024), ACM
資料類型 conference
dc.contributor 資管系
dc.creator (作者) 郁方
dc.creator (作者) Yu, Fang;Wen, Shang-Jung;Chang, Jia-Ming
dc.date (日期) 2024-08
dc.date.accessioned 4-Oct-2024 09:21:34 (UTC+8)-
dc.date.available 4-Oct-2024 09:21:34 (UTC+8)-
dc.date.issued (上傳時間) 4-Oct-2024 09:21:34 (UTC+8)-
dc.identifier.uri (URI) https://nccur.lib.nccu.edu.tw/handle/140.119/153865-
dc.description.abstract (摘要) High-dimensional single-cell data poses significant challenges in identifying underlying biological patterns due to the complexity and heterogeneity of cellular states. We propose a comprehensive gene-cell dependency visualization via unsupervised clustering, Growing Hierarchical Self-Organizing Map (GHSOM), specifically designed for analyzing high-dimensional single-cell data like single-cell sequencing and CRISPR screens. GHSOM is applied to cluster samples in a hierarchical structure such that the self-growth structure of clusters satisfies the required variations between and within. We propose a novel Significant Attributes Identification Algorithm to identify features that distinguish clusters. This algorithm pinpoints attributes with minimal variation within a cluster but substantial variation between clusters. These key attributes can then be used for targeted data retrieval and downstream analysis. Furthermore, we present two innovative visualization tools: Cluster Feature Map and Cluster Distribution Map. The Cluster Feature Map highlights the distribution of specific features across the hierarchical structure of GHSOM clusters. This allows for rapid visual assessment of cluster uniqueness based on chosen features. The Cluster Distribution Map depicts leaf clusters as circles on the GHSOM grid, with circle size reflecting cluster data size and color customizable to visualize features like cell type or other attributes. We apply our analysis to three single-cell datasets and one CRISPR dataset (cell-gene database) and evaluate clustering methods with internal and external CH and ARI scores. GHSOM performs well, being the best performer in internal evaluation (CH=4.2). In external evaluation, GHSOM has the third-best performance of all methods.
dc.format.extent 7250331 bytes-
dc.format.mimetype application/pdf-
dc.relation (關聯) 23rd International Workshop on Data Mining in Bioinformatics (In conjuction with ACM KDD 2024), ACM
dc.subject (關鍵詞) unsupervised clustering; self-organizing map; growing hierarchical self-organizing map; single-cell; cluster feature map; cluster distribution map; scRNAseq
dc.title (題名) scGHSOM: Hierarchical clustering and visualization of single-cell and CRISPR data using growing hierarchical SOM
dc.type (資料類型) conference