Skip to main navigation Skip to search Skip to main content

Mix-DANN and Dynamic-Modal-Distillation for Video Domain Adaptation

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Video domain adaptation is non-trivial due to video is inherently involved with multi-dimensional and multi-modal information. Existing works mainly adopt adversarial learning and self-supervised tasks to align features. Nevertheless, the explicit interaction between source and target in the temporal dimension, as well as the adaptation between modalities, are unexploited. In this paper, we propose Mix-Domain-Adversarial Neural Network and Dynamic-Modal-Distillation (MD-DMD), a novel multi-modal adversarial learning framework for unsupervised video domain adaptation. Our approach incorporates the temporal information between source and target domains, as well as the diversity of adaptability between modalities. On the one hand, for every single modality, we mix the frames from source and target domains to form mix-samples, then let the adversarial-discriminator predict the mix ratio of a mix-sample to further enhance the ability of the model to capture domain-invariant feature representations. On the other hand, we dynamically estimate the adaptability for different modalities during training, then pick the most adaptable modality as a teacher to guide other modalities by knowledge distillation. As a result, modalities are capable of learning transferable knowledge from each other, which leads to more effective adaptation. Experiments on two video domain adaptation benchmarks demonstrate the superiority of our proposed MD-DMD over state-of-the-art methods. © 2022 Association for Computing Machinery.
Original languageEnglish
Title of host publicationMM '22
Subtitle of host publicationProceedings of the 30th ACM International Conference on Multimedia
PublisherAssociation for Computing Machinery
Pages3224-3233
ISBN (Print)978-1-4503-9203-7
DOIs
Publication statusPublished - 10 Oct 2022
Event30th ACM International Conference on Multimedia (MM 2022) - Lisbon, Portugal
Duration: 10 Oct 202214 Oct 2022
https://2022.acmmm.org/

Publication series

NameMM - Proceedings of the 30th ACM International Conference on Multimedia

Conference

Conference30th ACM International Conference on Multimedia (MM 2022)
Abbreviated titleACM Multimedia 2022
PlacePortugal
CityLisbon
Period10/10/2214/10/22
Internet address

Bibliographical note

Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).

Research Keywords

  • adversarial learning
  • dynamic-modal-distillation
  • video domain adaptation

Fingerprint

Dive into the research topics of 'Mix-DANN and Dynamic-Modal-Distillation for Video Domain Adaptation'. Together they form a unique fingerprint.

Cite this