Skip to main navigation Skip to search Skip to main content

Surviving in Diverse Biases: Unbiased Dataset Acquisition in Online Data Market for Fair Model Training

  • Jiashi Gao
  • , Ziwei Wang
  • , Xiangyu Zhao
  • , Xin Yao
  • , Xuetao Wei*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

The online data markets have emerged as a valuable source of diverse datasets for training machine learning (ML) models. However, datasets from different data providers may exhibit varying levels of bias with respect to certain sensitive attributes in the population (such as race, sex, age, and marital status). Recent dataset acquisition research has focused on maximizing accuracy improvements for downstream model training, ignoring the negative impact of biases in the acquired datasets, which can lead to an unfair model. Can a consumer obtain an unbiased dataset from datasets with diverse biases? In this work, we propose a fairness-aware data acquisition framework (FAIRDA) to acquire high-quality datasets that maximize both accuracy and fairness for consumer local classifier training while remaining within a limited budget. Given the biases of data commodities remain opaque to consumers, the data acquisition in FAIRDA employs explore-exploit strategies. Based on whether exploration and exploitation are conducted sequentially or alternately, we introduce two algorithms: the knowledge-based offline data acquisition (KDA) and the reward-based online data acquisition algorithms (RDA). Each algorithm is tailored to specific customer needs, giving the former an advantage in computational efficiency and the latter an advantage in robustness. We conduct experiments to demonstrate the effectiveness of the proposed data acquisition framework in steering users toward fairer model training compared to existing baselines under varying market settings. © 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
Original languageEnglish
Title of host publicationProceedings of the Seventh AAAI/ACM Conference on AI, Ethics, and Society (AIES-24)
EditorsSanmay Das, Brian Patrick Green, Kush Varshney, Marianna Ganapini , Andrea Renda
Place of PublicationWashington, DC
PublisherAAAI Press
Pages451-462
ISBN (Print)978-1-57735-892-3, 1-57735-892-9
DOIs
Publication statusPublished - 2024
Event7th AAAI/ACM Conference on AI, Ethics, and Society (AIES-24) - San Jose, United States
Duration: 21 Oct 202423 Oct 2024

Publication series

NameProceedings of the AAAI/ACM Conference on AI, Ethics, and Society
Number1
Volume7

Conference

Conference7th AAAI/ACM Conference on AI, Ethics, and Society (AIES-24)
PlaceUnited States
CitySan Jose
Period21/10/2423/10/24

Funding

This work was supported by Key Programs of Guangdong Province under Grant2021QN02X166. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the funding parties.

Fingerprint

Dive into the research topics of 'Surviving in Diverse Biases: Unbiased Dataset Acquisition in Online Data Market for Fair Model Training'. Together they form a unique fingerprint.

Cite this