Abstract
In monocular 3D human pose estimation, target motions are generally stable and continuous, which indicates that joint velocity can provide valuable information for better estimation. Therefore, it is critical to learn the joint motion trajectory and spatio-temporal information from velocity. Previous works have shown that Transformers are effective in capturing the relationship between tokens. However, in practice, only 2D position is available and 3D velocity has not been explicitly used as a model input. To address this challenge, we propose TMT (Two-step Mixed-Training strategy), a transformer-based approach that effectively incorporates 3D velocity into the input vector during training, allowing for better learning of relevant features in the shallow layers. Extensive experiments demonstrate that TMT significantly improves the performance of state-of-the-art models, such as MixSTE, MHFormer, and PoseFomer, on two datasets: Human3.6M and MPI-INF-3DHP. TMT outperforms the state-of-the-art approach by up to 13.8% on the Human3.6M dataset. © 2024 IEEE.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2024 IEEE Winter Conference on Applications of Computer Vision, WACV 2024 |
| Publisher | IEEE |
| Pages | 3320-3329 |
| ISBN (Electronic) | 979-8-3503-1892-0 |
| ISBN (Print) | 979-8-3503-1893-7 |
| DOIs | |
| Publication status | Published - 2024 |
| Event | 24th IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2024) - Waikoloa Beach Marriott Resort, Waikoloa, United States Duration: 4 Jan 2024 → 8 Jan 2024 https://wacv2024.thecvf.com/ https://openaccess.thecvf.com/menu https://ieeexplore.ieee.org/xpl/conhome/10483279/proceeding |
Publication series
| Name | Proceedings - IEEE Winter Conference on Applications of Computer Vision, WACV |
|---|---|
| ISSN (Print) | 2472-6737 |
| ISSN (Electronic) | 2642-9381 |
Conference
| Conference | 24th IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2024) |
|---|---|
| Place | United States |
| City | Waikoloa |
| Period | 4/01/24 → 8/01/24 |
| Internet address |
Bibliographical note
Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).”Funding
This work is funded by Hong Kong Innovation and Technology Commission (InnoHK Project CIMDA).
Research Keywords
- 3D computer vision
- Algorithms
Fingerprint
Dive into the research topics of '3D Human Pose Estimation with Two-step Mixed-Training Strategy'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver