Skip to main navigation Skip to search Skip to main content

Deep Learning-Based Intra Mode Derivation for Versatile Video Coding

  • Linwei ZHU
  • , Yun ZHANG*
  • , Na LI
  • , Gangyi JIANG
  • , Sam KWONG
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

In intra coding, Rate Distortion Optimization (RDO) is performed to achieve the optimal intra mode from a pre-defined candidate list. The optimal intra mode is also required to be encoded and transmitted to the decoder side besides the residual signal, where lots of coding bits are consumed. To further improve the performance of intra coding in Versatile Video Coding (VVC), an intelligent intra mode derivation method is proposed in this paper, termed as Deep Learning based Intra Mode Derivation (DLIMD). In specific, the process of intra mode derivation is formulated as a multi-class classification task, which aims to skip the module of intra mode signaling for coding bits reduction. The architecture of DLIMD is developed to adapt to different quantization parameter settings and variable coding blocks including non-square ones, where only one single trained model is required. Different from the existing deep learning based classification problems, the hand-crafted features are also fed into intra mode derivation network besides the learned features from feature learning network. To compete with traditional methods, one additional binary flag is utilized in the video codec to indicate the selected scheme with RDO. Extensive experimental results reveal that the proposed method can achieve 2.28%, 1.74%, and 2.18% bit rate reduction on average for Y, U, and V components on the platform of VVC test model, which outperforms the state-of-the-art works. © 2023 Association for Computing Machinery.
Original languageEnglish
Article number96
JournalACM Transactions on Multimedia Computing Communications and Applications
Volume19
Issue number2s
Online publishedFeb 2023
DOIs
Publication statusPublished - Apr 2023

Funding

This work was supported in part by the National Natural Science Foundation of China under Grants 62172400, 61901459, 61902389 and 62271276, in part by the Shenzhen Science and Technology Program under Grant JCYJ20200109110410133, in part by the Guangdong Basic and Applied Basic Research Foundation under Grant 2022A1515011351, in part by the Membership of Youth Innovation Promotion Association, Chinese Academy of Sciences (CAS) under Grant 2018392, in part by the China Postdoctoral Science Foundation under Grant 2021T140696, in part by the CAS Present’s International Fellowship Initiative (PIFI) under Grant 2022VTA0005, in part by the Hong Kong Innovation and Technology Commission (InnoHK Project CIMDA), in part by the Hong Kong GRF-RGC General Research Fund under Grants 11209819 (CityU 9042816) and 11203820 (CityU 9042598).

Research Keywords

  • Versatile video coding
  • intra mode derivation
  • most probable mode
  • deep learning
  • multi-class classification
  • PREDICTION

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'Deep Learning-Based Intra Mode Derivation for Versatile Video Coding'. Together they form a unique fingerprint.

Cite this