TINF-Consensus for Multi-Annotator Learning: Separating Neutral Responses from Polarized Disagreement

Authors

  • Mukhtar Ahmad Doctoral Program of Mathematics, Faculty of Mathematics and Natural Sciences, Institut Teknologi Bandung, Jalan Ganesha No. 10, Bandung 40132, Indonesia;

Keywords:

multi-annotator learning; crowdsourcing; neutrosophic logic; label aggregation; annotator disagreement; abstention; human label variation; robust consensus.

Abstract

 Aggregating multiple annotations into one label usually compresses two different phenomena into the same low
confidence outcome: explicit neutrality, where annotators decline to support either class, and polarized disagreement, where
decisive annotators support opposite classes. The distinction matters operationally. Neutral items may need clarification or
additional evidence, whereas polarized items may require adjudication, subgroup analysis, or retention of plural labels. We
introduce TINF-Consensus, a reliability-weighted binary label-fusion framework built from the truth, indeterminacy, neutrality,
falsehood (tinf) representation. For each item, weighted positive, neutral, and negative response masses define truth T, neutrality
N, and falsehood F. Pure indeterminacy is then I = 4TF/(T +F), which measures polarization only among decisive responses.
A derived consensus coordinate C = (T −F)2/(T +F) yields the exact decomposition C +I +N = 1. Thus, all-neutral and
evenly split decisive crowds—which have the same signed mean—receive different profiles. Gold questions provide Beta-smoothed
annotator orientation and reliability weights. We establish permutation invariance, an adversarial mass bound, consistency of
gold-based orientation, and an exponential conditional error bound.
Deterministic semi-synthetic experiments use 21 annotators, 40 repetitions, and four public binary datasets. The contrast
I −N distinguishes constructed polarization from explicit neutrality with macro-average AUROC 0.99999, and the rule I > N
routes the reason for review with 0.9928 accuracy. With a nominal 40% adversarial setting (eight of 21 annotators), gold-weighted
TINF-Consensus retains 0.9942 macro-average accuracy, compared with 0.5930 for majority vote and 0.9949 for a gold-anchored
Dawid–Skene model. Importantly, weighted margin deficit, not indeterminacy, is best for predicting aggregation errors (AUROC
0.9552). The results support a two-output design: margin answers whether the fused label is fragile, whereas tinf answers why
the annotation process is unresolved. The study is a controlled methodological validation; real multi-rater and domain-specific
evaluation remains necessary.

 

DOI 10.5281/zenodo.22188626

Downloads

Download data is not yet available.

Downloads

Published

2026-06-25

How to Cite

Mukhtar Ahmad. (2026). TINF-Consensus for Multi-Annotator Learning: Separating Neutral Responses from Polarized Disagreement. Neutrosophic Sets and Systems, 100, 380-400. https://fs.unm.edu/nss8/index.php/111/article/view/7730