Abstract
Training general-purpose vision models on purely sequential visual data, eschewing linguistic inputs, has heralded a new frontier in visual understanding. These models are intended to not only comprehend but also seamlessly transit to out-of-domain tasks. However, current endeavors are hamstrung by an over-reliance on colossal models, exemplified by models with upwards of 3B parameters, and the necessity for an extensive corpus of visual data, often comprising a staggering 400B tokens (Bai et al., 2023). In this paper, we delve into the development of an efficient, autoregression-based vision model, innovatively architected to operate on a limited dataset. We meticulously demonstrate how this model achieves proficiency in a spectrum of visual tasks spanning both high-level and low-level semantic understanding during the testing phase. Our empirical evaluations underscore the model's agility in adapting to various tasks, heralding a significant reduction in the parameter footprint, and a marked decrease in training data requirements, thereby paving the way for more sustainable and accessible advancements in the field of generalist vision models. The code is available at https://github.com/ggjy/DeLVM.
Copyright 2024 by the author(s).
Copyright 2024 by the author(s).
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 41st International Conference on Machine Learning |
| Editors | Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, Felix Berkenkamp |
| Place of Publication | United States |
| Publisher | ML Research Press |
| Pages | 17572-17596 |
| Number of pages | 25 |
| Publication status | Published - 21 Jul 2024 |
| Externally published | Yes |
| Event | 41st International Conference on Machine Learning (ICML 2024) - Messe Wien Exhibition Congress Center, Vienna, Austria Duration: 21 Jul 2024 → 27 Jul 2024 https://proceedings.mlr.press/v235/ https://icml.cc/ |
Publication series
| Name | Proceedings of Machine Learning Research |
|---|---|
| Publisher | ML Research Press |
| Volume | 235 |
| ISSN (Print) | 2640-3498 |
Conference
| Conference | 41st International Conference on Machine Learning (ICML 2024) |
|---|---|
| Place | Austria |
| City | Vienna |
| Period | 21/07/24 → 27/07/24 |
| Internet address |
Bibliographical note
Publisher Copyright:Copyright 2024 by the author(s)
Funding
This paper is supported by National Key Research and Development Program of China under No. 2021YFC3300128, and Joint Funds of the National Natural Science Foundation of China No. U2336211. Chang Xu was supported in part by the Australian Research Council under Projects DP240101848 and FT230100549.
Fingerprint
Dive into the research topics of 'Data-efficient Large Vision Models through Sequential Autoregression'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver