Skip to main navigation Skip to search Skip to main content

ANG: Accelerating NFA processing on GPUs via Exploring Multi-Level Fine-Grained Parallelism

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Finite Automata (FA) processing is a core computation in various real-world applications. Over the past decades, extensive efforts have been dedicated to accelerating FA processing on modern parallel platforms, particularly GPUs, due to their high memory bandwidth and massive hardware parallelism. As Non-deterministic Finite Automata (NFA)-based applications have strong and growing demands for real-time data analytics nowadays, reducing latency in automata processing has become a critical priority. However, existing approaches face significant challenges when limited parallelism is exposed in NFA computations. In this work, we explore opportunities of introducing fine-grained parallelism from various sources and addressing the limitations of fast NFA processing. Specifically, by analyzing different NFA parallelization schemes, we identify the major performance issue caused by insufficient state-level parallelism in conventional designs. To overcome the bottleneck, this work introduces speculative parallelization tailored for GPU-based NFA processing, thus effectively exploiting fine-grained parallelism across multilevels, with a particular focus on input-chunk-level parallelism. To realize speculative parallelization in practice, we develop ANG, a latency-oriented NFA processing framework that overcomes key implementation challenges on GPUs. We evaluate the efficiency of ANG on a set of representative NFAs with diverse properties. Experimental results demonstrate that ANG achieves significant performance improvement compared to state-of-theart techniques, with reaching 11.74× speedup on average (and up to 49.88× in extreme cases).
©2025 IEEE 
Original languageEnglish
Title of host publication2025 34th International Conference on Parallel Architectures and Compilation Techniques (PACT)
PublisherIEEE
Pages135-147
Number of pages13
ISBN (Electronic)979-8-3315-8295-1
ISBN (Print)979-8-3315-8296-8
DOIs
Publication statusPublished - 2025
Event2025 34th International Conference on Parallel Architectures and Compilation Techniques (PACT) - Irvine, California, United States
Duration: 3 Nov 20256 Nov 2025
https://pact2025.github.io/

Conference

Conference2025 34th International Conference on Parallel Architectures and Compilation Techniques (PACT)
Abbreviated titlePACT 2025
PlaceUnited States
CityIrvine, California
Period3/11/256/11/25
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Funding

We are grateful to the anonymous reviewers for their constructive comments and suggestions. The work was partially supported by City University of Hong Kong internal and donation fundings (No. 9610598 and No. 9220148), and the National Science Foundation (NSF) Grant (No. 2105006).

Research Keywords

  • NFA
  • GPU
  • Speculative Parallelization

Fingerprint

Dive into the research topics of 'ANG: Accelerating NFA processing on GPUs via Exploring Multi-Level Fine-Grained Parallelism'. Together they form a unique fingerprint.

Cite this