HOTSPOT: hierarchical host prediction for assembled plasmid contigs with transformer

Yongxin Ji, Jiayu Shang, Xubo Tang, Yanni Sun*

*Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

8 Citations (Scopus)
82 Downloads (CityUHK Scholars)

Abstract

Motivation: As prevalent extrachromosomal replicons in many bacteria, plasmids play an essential role in their hosts' evolution and adaptation. The host range of a plasmid refers to the taxonomic range of bacteria in which it can replicate and thrive. Understanding host ranges of plasmids sheds light on studying the roles of plasmids in bacterial evolution and adaptation. Metagenomic sequencing has become a major means to obtain new plasmids and derive their hosts. However, host prediction for assembled plasmid contigs still needs to tackle several challenges: different sequence compositions and copy numbers between plasmids and the hosts, high diversity in plasmids, and limited plasmid annotations. Existing tools have not yet achieved an ideal tradeoff between sensitivity and precision on metagenomic assembled contigs.

Results: In this work, we construct a hierarchical classification tool named HOTSPOT, whose backbone is a phylogenetic tree of the bacterial hosts from phylum to species. By incorporating the state-of-the-art language model, Transformer, in each node's taxon classifier, the top-down tree search achieves an accurate host taxonomy prediction for the input plasmid contigs. We rigorously tested HOTSPOT on multiple datasets, including RefSeq complete plasmids, artificial contigs, simulated metagenomic data, mock metagenomic data, the Hi-C dataset, and the CAMI2 marine dataset. All experiments show that HOTSPOT outperforms other popular methods.

Availability and implementation : The source code of HOTSPOT is available via: https://github.com/Orin-beep/HOTSPOT

© The Author(s) 2023. Published by Oxford University Press.

Original languageEnglish
Article numberbtad283
JournalBioinformatics
Volume39
Issue number5
Online published22 Apr 2023
DOIs
Publication statusPublished - May 2023

Research Keywords

  • Plasmid
  • Host prediction
  • Deep learning
  • Transformer
  • Tree-based classification.

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

Fingerprint

Dive into the research topics of 'HOTSPOT: hierarchical host prediction for assembled plasmid contigs with transformer'. Together they form a unique fingerprint.

Cite this