Projects per year
Abstract
Monocular estimation of 3D human poses is challenging due to ambiguity in depths and partial occlusion. Most recent works define this as a 2D-to-3D lifting task, taking 2D key point sequences and using spatial and temporal relationships. However, prior works focus on capturing spatio-temporal correlations but ignore the motion of joints that is needed for continuous estimation. To extend the potential of 2D-to-3D pose estimation, we propose TSwinPose, which learns multi-scale spatio-temporal representations from 2D key point locations and patterns of motion. The input 2D key point sequences are enhanced by JointFlow, which encodes the motion of each human joint. Based on Swin-Transformer, we designed a temporal domain Swin-Unet structure to model multi-scale spatio-temporal relationships of human joints across different temporal windows. The final 3D pose generated by multi-stage representations is consistent temporally and has a higher accuracy. Experiments conducted on three benchmark datasets, Human3.6M, MPI-INF-3DHP, and HumanEva-I, demonstrate that TSwinPose achieves performance that is on par with state-of-the-art methods. Moreover, the introduction of JointFlow as a plug-in extension enhances performance significantly, particularly benefiting long-term 2D-to-3D lifting human pose estimation methods. © 2024 Elsevier Ltd
| Original language | English |
|---|---|
| Article number | 123545 |
| Journal | Expert Systems with Applications |
| Volume | 249 |
| Issue number | Part A |
| Online published | 27 Feb 2024 |
| DOIs | |
| Publication status | Published - 1 Sept 2024 |
Funding
This work is supported by Hong Kong Innovation and Technology Commission (InnoHK Project CIMDA) and Hong Kong Research Grants Council (Project CityU 11204821).
Research Keywords
- 3D human pose estimation
- Monocular video
- Transformer
RGC Funding Information
- RGC-funded
Fingerprint
Dive into the research topics of 'TSwinPose: Enhanced monocular 3D human pose estimation with JointFlow'. Together they form a unique fingerprint.Projects
- 1 Finished
-
GRF: Matching Large Feature Sets based on Hypergraph Models and Structurally Adaptive CUR Decompositions of Compatibility Tensors
YAN, H. (Principal Investigator / Project Coordinator)
1/01/22 → 3/06/26
Project: Research
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver