Skip to main navigation Skip to search Skip to main content

Protocol for the automatic extraction of epidemiological information via a pre-trained language model

  • Zhizheng Wang
  • , Xiao Fan Liu
  • , Zhanwei Du
  • , Lin Wang
  • , Ye Wu
  • , Petter Holme
  • , Michael Lachmann
  • , Hongfei Lin
  • , Zhuoyue Wang
  • , Yu Cao
  • , Zoie S.Y. Wong*
  • , Xiao-Ke Xu*
  • , Yuanyuan Sun*
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

47 Downloads (CityUHK Scholars)

Abstract

The lack of systems to automatically extract epidemiological fields from open-access COVID-19 cases restricts the timeliness of formulating prevention measures. Here we present a protocol for using CCIE, a COVID-19 Cases Information Extraction system based on the pre-trained language model.1 We describe steps for preparing supervised training data and executing python scripts for named entity recognition and text category classification. We then detail the use of machine evaluation and manual validation to illustrate the effectiveness of CCIE. For complete details on the use and execution of this protocol, please refer to Wang et al.2 © 2023 The Authors.
Original languageEnglish
Article number102392
JournalSTAR Protocols
Volume4
Issue number3
Online published1 Jul 2023
DOIs
Publication statusPublished - 15 Sept 2023

Research Keywords

  • Clinical Protocol
  • Computer Sciences
  • Health Sciences

Publisher's Copyright Statement

  • This full text is made available under CC-BY-NC-ND 4.0. https://creativecommons.org/licenses/by-nc-nd/4.0/

Fingerprint

Dive into the research topics of 'Protocol for the automatic extraction of epidemiological information via a pre-trained language model'. Together they form a unique fingerprint.

Cite this