Abstract
This work presents CAT++, an enhanced Context-Aware Transformer (CAT), to address the need for more effective 3D point cloud annotation. CAT++ is an end-to-end 3D automatic annotator that lifts 2D weak annotations into 3D, thereby minimizing human annotation burdens. Specifically, it employs a hierarchical-interleaved encoder (HIE) to enhance its encoder-decoder architecture. The HIE alternates between intra- and inter-object Transformer blocks, integrating local and global object features by performing self-attentions along the sequence and batch dimensions. This process continuously combines local and global object features and creates a 3D representation that links interactions between objects at different scales, leading to a more comprehensive understanding of the scene. Moreover, CAT++ features an attention-conditioned implicit neural representation (INR) network, which has been specially engineered to enable more precise and efficient continuous modeling of 3D object surfaces. This improves CAT++’s multi-task learning and its ability to handle hard samples. Results on KITTI and nuScenes datasets show CAT++’s state-of-the-art performance, outperforming other methods by up to 2% AP3D on the KITTI test split. Furthermore, CAT++ significantly reduces human annotation effort, by factors of 6.4 and 50 on the KITTI and nuScenes datasets for the Car category, respectively. © The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2026.
| Original language | English |
|---|---|
| Article number | 71 |
| Number of pages | 21 |
| Journal | International Journal of Computer Vision |
| Volume | 134 |
| Issue number | 2 |
| Online published | 20 Jan 2026 |
| DOIs | |
| Publication status | Published - Feb 2026 |
| Externally published | Yes |
Research Keywords
- 3D Annotation
- 3D Point Cloud
- Implicit Neural Representation
- LiDAR
- Transformer
Fingerprint
Dive into the research topics of 'CAT++: Enhancing 3D Annotations with Hierarchical-Interleaved Encoding and Attention-Conditioned Implicit Representation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver