Abstract
Road pothole segmentation is important for the driving safety of autonomous vehicles, especially in unstructured or rural environments. Recently, many effective multi-modal fusion networks have been proposed for road pothole segmentation. However, most of them adopt two encoders with the same type of structure, such as only convolutional neural network (CNN) or only Transformer, to extract features from different modalities. This overlooks the fact that the information richness of features extracted from different modalities varies across the types of features, such as CNN features or self-attention features. To provide a solution to this issue, we design a novel RGB-Disparity segmentation network, named MMFSeg, by adopting the two types of structures as encoders. Specifically, we adopt CNN and Transformer as encoders to extract two types of features from each modality, and also propose a late-fusion multi-feature alignment fusion module to fuse the two types of features with different numbers of channels. Experimental results demonstrate that our network outperforms well-known networks, and can trade-off between accuracy and efficiency. © 2025 IEEE.
| Original language | English |
|---|---|
| Pages (from-to) | 22742-22754 |
| Number of pages | 13 |
| Journal | IEEE Transactions on Automation Science and Engineering |
| Volume | 22 |
| Online published | 23 Sept 2025 |
| DOIs | |
| Publication status | Published - 2025 |
Funding
This work was supported by the City University of Hong Kong under Grant 9610675.
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Research Keywords
- autonomous vehicles
- convolution-transformer structure
- multi-modal fusion
- Pothole segmentation
Fingerprint
Dive into the research topics of 'MMFSeg: Multi-Structure Multi-Feature Fusion for Segmentation of Road Potholes'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver