Skip to main navigation Skip to search Skip to main content

Utilizing motion segmentation for optimizing the temporal adjacency matrix in 3D human pose estimation

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

In monocular 3D human pose estimation, modeling the temporal relation of human joints is crucial for prediction accuracy. Currently, most methods utilize transformer to model the temporal relation among joints. However, existing transformer-based methods have limitations. The temporal adjacency matrix utilized within the self-attention of the temporal transformer inaccurately models the temporal relationships between frames, particularly in cases where distinct motions exhibit significant correlation despite having different physical interpretations and large temporal spans. To address this issue, we construct an artificial temporal adjacency matrix based on input data and introduce a temporal adjacency matrix hybrid module to blend this matrix with the model’s inherent temporal adjacency matrix, resulting in a novel composite temporal adjacency matrix. Through extensive experiments on Human3.6M and MPI-INF-3DHP datasets using state-of-the-art methods as benchmarks, our proposed method demonstrates a maximum improvement of up to 5.6% compared to the original approach. © 2024 Published by Elsevier B.V.
Original languageEnglish
Article number128153
Pages (from-to)1-12
Number of pages12
JournalNeurocomputing
Volume600
Online published5 Jul 2024
DOIs
Publication statusPublished - 1 Oct 2024

Funding

We thank Zhengwei Wang for his help with software development. This work is supported by Hong Kong Innovation and Technology Commission (InnoHK Project CIMDA) and Hong Kong Research Grants Council (Project CityU 11204821 ).

Research Keywords

  • 3D human pose estimation
  • Temporal adjacency matrix
  • Motion segmentation

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'Utilizing motion segmentation for optimizing the temporal adjacency matrix in 3D human pose estimation'. Together they form a unique fingerprint.

Cite this