Skip to main navigation Skip to search Skip to main content

Beyond Majority Voting: A Coarse-to-Fine Label Filtration for Heavily Noisy Labels

  • Bo Han
  • , Ivor W. Tsang*
  • , Ling Chen
  • , Joey Tianyi Zhou
  • , Celina P. Yu
  • *Corresponding author for this work

Research output: Contribution to journalJournal articlepeer-review

22 Citations (Scopus)

Abstract

Crowdsourcing has become the most appealing way to provide a plethora of labels at a low cost. Nevertheless, labels from amateur workers are often noisy, which inevitably degenerates the robustness of subsequent learning models. To improve the label quality for subsequent use, majority voting (MV) is widely leveraged to aggregate crowdsourced labels due to its simplicity and scalability. However, when crowdsourced labels are 'heavily' noisy (e.g., 40% of noisy labels), MV may not work well because of the fact 'garbage (heavily noisy labels) in, garbage (full aggregated labels) out.' This issue inspires us to think: if the ultimate target is to learn a robust model using noisy labels, why not provide partial aggregated labels and ensure that these labels are reliable enough for learning models? To solve this challenge by improving MV, we propose a coarse-to-fine label filtration model called double filter machine (DFM), which consists of a (majority) voting filter and a sparse filter serially. Specifically, the DFM refines crowdsourced labels from coarse filtering to fine filtering. In the stage of coarse filtering, the DFM aggregates crowdsourced labels by voting filter, which yields (quality-acceptable) full aggregated labels. In the stage of fine filtering, DFM further digs out a set of high-quality labels from full aggregated labels by sparse filter, since this filter can identify high-quality labels by the methodology of support selection. Based on the insight of compressed sensing, DFM recovers a ground-truth signal from heavily noisy data under a restricted isometry property. To sum up, the primary benefits of DFM are to keep the scalability by voting filter, while improve the robustness by sparse filter. We also derive theoretical guarantees for the convergence and recovery of DFM and reveal its complexity. We conduct comprehensive experiments on both the UCI simulated and the AMT crowdsourced datasets. Empirical results show that partial aggregated labels provided by DFM effectively improve the robustness of learning models.

Original languageEnglish
Pages (from-to)3774-3787
Number of pages14
JournalIEEE Transactions on Neural Networks and Learning Systems
Volume30
Issue number12
Early online date15 Mar 2019
DOIs
Publication statusPublished - Dec 2019

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure

User-Defined Keywords

  • Coarse-to-fine label filtration
  • double filter machine (DFM)
  • heavily noisy labels
  • majority voting (MV)
  • robustness
  • scalability
  • support selection

Fingerprint

Dive into the research topics of 'Beyond Majority Voting: A Coarse-to-Fine Label Filtration for Heavily Noisy Labels'. Together they form a unique fingerprint.

Cite this