Skip to main navigation Skip to search Skip to main content

Invariant descriptions for motion trajectory representation and recognition

  • Zhanpeng SHAO

    Student thesis: Doctoral Thesis

    Abstract

    Motion trajectories tracked from motions of humans, robots, and moving objects can provide an important clue for motion analysis, matching, and recognition. Many studies have adopted motion trajectories as key features in diverse motion analysis scenarios, such as human action recognition, motion retrieval, human-robot interaction, and imitation learning by demonstration. However, in most related applications, raw data or simple descriptions for motion trajectories are often used directly. An effective and robust description for motion trajectories has not been obtained in principle, particularly when noise, motion variations, and view changes exist in a motion tracking system. Instead of using simple descriptions as most of the current work, this thesis proposes a framework to study and define some invariant descriptions for motion trajectories, show their rich properties in motion trajectory representation, and present their advantages in motion trajectory recognition. First, a general definition of integral invariants of Euclidean and similarity groups is proposed for motion trajectories in both 2-dimensional (2D) and 3-dimensional (3D) Euclidean spaces, where integral invariants are defined as line integrals of a class of kernels along a motion trajectory. According to the definition, two integral invariants (distance and area integral invariants) are designed on the basis of two typical kernels. A robust estimation of the area integral invariants for discrete motion trajectories is formulated based on the maximal blurred segment of noisy discrete curves to avoid the computation of high-order derivatives. Such integral invariants possess some desirable properties, such as computational locality, uniqueness of representation, and noise insensitivity. Moreover, the definition of integral invariants allows a multiscale space analysis of motion trajectories by varying the scale of a kernel. The features of motion trajectories can be perceived at multiple scales in a coarse-to-fine manner by extending integral invariants at one fixed scale to their multiscale representation. When using integral invariants to match and retrieve similar motion trajectories, a distance of dynamic time warping (DTW) of integral invariants is defined to measure the trajectory similarity. The DTW distance can deal with the different lengths, different sampling rates, and occlusions between a pair of motion trajectories. However, such a deterministic method, the DTW distance of integral invariants, shows a low efficiency in trajectory recognition. This research resorts to a learningbased method to model the statistics of each motion class when performing a fast recognition task. To fit a learning-based model, self-similarity descriptors are proposed by exploring local temporal self-similarities of motion trajectories over time. Such temporal self-similarities are observed by transforming a motion trajectory into an image of a self-similarity matrix (SSM) of integral invariants. On analysis of SSM imges, a set of self-similarity descriptors are extracted from the diagonal of each SSM image. Furthermore, sparse codes of local self-similarity descriptors are max pooled across different temporal sub-blocks and over different temporal scales to model the statistics of self-similarity descriptors for each SSM. The max pooled features are concatenated to form a temporal pyramid representation with a linear matching kernel. By using a multiclass support vector machine (SVM) and the linear matching kernel, the training complexity of SVMs is O(n) and the testing complexity of SVMs is a constant. Such a temporal pyramid representation with the linear matching kernel contributes to improve the recognition efficiency a lot compared to a variety of other approaches. As action sequences can be abstracted as multiple motion trajectories, a hierarchical descriptor is constructed by decomposing a group of multiple motion trajectories into a root trajectory and child trajectories. While the root trajectory is represented by integral invariants, child trajectories are represented by relative orientations and distances of themselves with respect to the root trajectory in a unit sphere. Thus, the hierarchical descriptor is built by concatenating representations of the root and child trajectories to form a feature vector that can capture spatio-temporal features within a group of multiple motion trajectories in such a compact form. Finally, multiple experiments are designed and conducted on several trajectory datasets to evaluate the proposed invariant descriptions for trajectory representation and recognition. Large benchmarks of trajectory matching are run to evaluate the claimed rich properties of integral invariants, where different kinds of integral invariants are used to measure the similarities between pairs of motion trajectories. In the following sign recognition on a public dataset, the effectiveness of integral invariants and self-similarity descriptors for motion trajectory recognition are evaluated in terms of both the recognition accuracy and efficiency. In addition, in cases where groups of multiple motion trajectories occur, hierarchical descriptors are used to retrieve similar human action sequences on three action datasets given a query. Experimental results demonstrate the effectiveness and robustness of these invariant descriptions in motion trajectory representation and recognition.
    Date of Award2 Oct 2015
    Original languageEnglish
    Awarding Institution
    • City University of Hong Kong
    SupervisorYou Fu LI (Supervisor)

    Keywords

    • Optical pattern recognition
    • Motion perception (Vision)

    Cite this

    '