Skip Navigation
Skip to contents

Journal of Microbiology : Journal of Microbiology

OPEN ACCESS
SEARCH
Search

Author index

Page Path
HOME > Browse Articles > Author index
Search
Seung-Hyun Jung 2 Articles
Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning
Yoojung Hwang, Woo Young Cho, Woojung Lee, Insun Joo, Jeong-Ih Shin, Mi-Ran Seo, Seung-Hun Shin, Kwan Soo Ko, Kun Taek Park, Yeun-Jun Chung, Seung-Hyun Jung
J. Microbiol. 2026;64(8):e2604011.   Published online August 31, 2026
DOI: https://doi.org/10.71150/jm.2604011
  • 424 View
  • 23 Download
AbstractAbstract PDFSupplementary Material

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

PneusPage: A WEB-BASED TOOL for the analysis of Whole-Genome Sequencing Data of Streptococcus pneumonia
Eunju Hong, Youngjin Shin, Hyunseong Kim, Woo Young Cho, Woo-Hyun Song, Seung-Hyun Jung, Minho Lee
J. Microbiol. 2025;63(1):e.2409020.   Published online January 24, 2025
DOI: https://doi.org/10.71150/jm.2409020
  • 3,675 View
  • 132 Download
  • 2 Web of Science
  • 3 Crossref
AbstractAbstract PDFSupplementary Material

With the advent of whole-genome sequencing, opportunities to investigate the population structure, transmission patterns, antimicrobial resistance profiles, and virulence determinants of Streptococcus pneumoniae at high resolution have been increasingly expanding. Consequently, a user-friendly bioinformatics tool is needed to automate the analysis of Streptococcus pneumoniae whole-genome sequencing data, summarize clinically relevant genomic features, and further guide treatment options. Here, we developed PneusPage, a web-based tool that integrates functions for species prediction, molecular typing, drug resistance determination, and data visualization of Streptococcus pneumoniae. To evaluate the performance of PneusPage, we analyzed 80 pneumococcal genomes with different serotypes from the Global Pneumococcal Sequencing Project and compared the results with those from another platform, PathogenWatch. We observed a high concordance between the two platforms in terms of serotypes (100% concordance rate), multilocus sequence typing (100% concordance rate), penicillin-binding protein typing (88.8% concordance rate), and the Global Pneumococcal Sequencing Clusters (98.8% concordance rate). In addition, PneusPage offers integrated analysis functions for the detection of virulence and mobile genetic elements that are not provided by previous platforms. By automating the analysis pipeline, PneusPage makes whole-genome sequencing data more accessible to non-specialist users, including microbiologists, epidemiologists, and clinicians, thereby enhancing the utility of whole-genome sequencing in both research and clinical settings. PneusPage is available at https://pneuspage.minholee.net/.

Citations

Citations to this article as recorded by  
  • Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning
    Yoojung Hwang, Woo Young Cho, Woojung Lee, Insun Joo, Jeong-Ih Shin, Mi-Ran Seo, Seung-Hun Shin, Kwan Soo Ko, Kun Taek Park, Yeun-Jun Chung, Seung-Hyun Jung
    Journal of Microbiology.2026; 64(8): e2604011.     CrossRef
  • Genomic analysis and pneumococcal population dynamics across PCV implementation in South Korea, 1997–2023
    Jeong-Ih Shin, Sung-Yeon Cho, Jiyon Chu, Chulmin Park, Minho Lee, Joon Young Song, Seung-Hyun Jung, Dong-Gun Lee
    Microbial Genomics .2025;[Epub]     CrossRef
  • GPS Pipeline: portable, scalable genomic pipeline for Streptococcus pneumoniae surveillance from Global Pneumococcal Sequencing Project
    Harry C. H. Hung, Narender Kumar, Victoria Dyster, Corin Yeats, Benjamin Metcalf, Yuan Li, Paulina A. Hawkins, Lesley McGee, Stephen D. Bentley, Stephanie W. Lo
    Nature Communications.2025;[Epub]     CrossRef

Journal of Microbiology : Journal of Microbiology
TOP