{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T05:55:08Z","timestamp":1782280508989,"version":"3.54.5"},"reference-count":75,"publisher":"Association for Computing Machinery (ACM)","issue":"CSCW2","license":[{"start":{"date-parts":[[2021,10,13]],"date-time":"2021-10-13T00:00:00Z","timestamp":1634083200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Hum.-Comput. Interact."],"published-print":{"date-parts":[[2021,10,13]]},"abstract":"<jats:p>Human ratings have become a crucial resource for training and evaluating machine learning systems. However, traditional elicitation methods for absolute and comparative rating suffer from issues with consistency and often do not distinguish between uncertainty due to disagreement between annotators and ambiguity inherent to the item being rated. In this work, we present Goldilocks, a novel crowd rating elicitation technique for collecting calibrated scalar annotations that also distinguishes inherent ambiguity from inter-annotator disagreement. We introduce two main ideas: grounding absolute rating scales with examples and using a two-step bounding process to establish a range for an item's placement. We test our designs in three domains: judging toxicity of online comments, estimating satiety of food depicted in images, and estimating age based on portraits. We show that (1) Goldilocks can improve consistency in domains where interpretation of the scale is not universal, and that (2) representing items with ranges lets us simultaneously capture different sources of uncertainty leading to better estimates of pairwise relationship distributions.<\/jats:p>","DOI":"10.1145\/3476076","type":"journal-article","created":{"date-parts":[[2021,10,19]],"date-time":"2021-10-19T02:30:11Z","timestamp":1634610611000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty"],"prefix":"10.1145","volume":"5","author":[{"given":"Quan Ze","family":"Chen","sequence":"first","affiliation":[{"name":"University of Washington, Seattle, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel S.","family":"Weld","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amy X.","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,18]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Abhaya Agarwal and A. Lavie. 2008. Meteor M-BLEU and M-TER: Evaluation Metrics for High-Correlation with Human Rankings of Machine Translation Output. In WMT@ACL.","DOI":"10.3115\/1626394.1626406"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.alw-1.2"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3308560.3317083"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1609\/aimag.v36i1.2564"},{"key":"e_1_2_1_5_1","volume-title":"R. Krishnan, Jason Stanley, O. Tickoo, L. Nachman, R. Chunara, Adrian Weller, and Alice Xiang.","author":"Bhatt Umang","year":"2020","unstructured":"Umang Bhatt, Y. Zhang, J. Antor\u00e1n, Q. Liao, P. Sattigeri, Riccardo Fogliato, Gabrielle Gauthier Melan\u00e7on, R. Krishnan, Jason Stanley, O. Tickoo, L. Nachman, R. Chunara, Adrian Weller, and Alice Xiang. 2020. Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty. ArXiv, Vol. abs\/2011.07586 (2020)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3415164"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Lukas Bossard M. Guillaumin and L. Gool. 2014. Food-101 - Mining Discriminative Components with Random Forests. In ECCV.","DOI":"10.1007\/978-3-319-10599-4_29"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2009.03.027"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242598"},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","unstructured":"G. Brown I. Neath and N. Chater. 2007. A temporal ratio model of memory. Psychological review Vol. 114 3 (2007) 539--76.","DOI":"10.1037\/0033-295X.114.3.539"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.appet.2008.04.017"},{"key":"e_1_2_1_12_1","unstructured":"Chris Callison-Burch M. Osborne and Philipp Koehn. 2006. Re-evaluation the Role of Bleu in Machine Translation Research. In EACL."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3025453.3026044"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858411"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300761"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359164"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0190393"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.2307\/2346806"},{"key":"e_1_2_1_19_1","unstructured":"Michael J. Denkowski and A. Lavie. 2010. Choosing the Right Evaluation for Machine Translation: an Examination of Annotator and Automatic Metric Performance on Human Judgment Tasks."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858268"},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Ryan Drapeau Lydia B Chilton Jonathan Bragg and Daniel S Weld. 2016. MicroTalk: Using Argumentation to Improve Crowdsourcing Accuracy.. In Hcomp. 32--41.","DOI":"10.1609\/hcomp.v4i1.13270"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"A. Dumitrache. 2015. Crowdsourcing Disagreement for Collecting Semantic Annotation. In ESWC.","DOI":"10.1007\/978-3-319-18818-8_43"},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"A. Dumitrache Lora Aroyo and Chris Welty. 2018a. Capturing Ambiguity in Crowdsourcing Frame Disambiguation. In HCOMP.","DOI":"10.1609\/hcomp.v6i1.13330"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3152889"},{"key":"e_1_2_1_25_1","volume-title":"SummEval: Re-evaluating Summarization Evaluation. ArXiv","author":"Fabbri A. R.","year":"2020","unstructured":"A. R. Fabbri, Wojciech Kryscinski, Bryan McCann, R. Socher, and Dragomir Radev. 2020. SummEval: Re-evaluating Summarization Evaluation. ArXiv, Vol. abs\/2007.12626 (2020)."},{"key":"e_1_2_1_26_1","unstructured":"Yanwei Fu Timothy M. Hospedales Tao Xiang Yuan Yao and Shaogang Gong. 2014. Interestingness Prediction by Robust Learning to Rank. In ECCV."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1162\/pres.16.4.439"},{"key":"e_1_2_1_28_1","volume-title":"Smith","author":"Gehman Samuel","year":"2020","unstructured":"Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In EMNLP."},{"key":"e_1_2_1_29_1","volume-title":"Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets. ArXiv","author":"Geva Mor","year":"2019","unstructured":"Mor Geva, Y. Goldberg, and Jonathan Berant. 2019. Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets. ArXiv, Vol. abs\/1908.07898 (2019)."},{"key":"e_1_2_1_30_1","volume-title":"Explaining and Harnessing Adversarial Examples. CoRR","author":"Goodfellow Ian J.","year":"2015","unstructured":"Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. CoRR, Vol. abs\/1412.6572 (2015)."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445423"},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse. Association for Computational Linguistics","author":"Graham Yvette","year":"2013","unstructured":"Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013. Continuous Measurement Scales in Human Evaluation of Machine Translation. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse. Association for Computational Linguistics, Sofia, Bulgaria, 33--41. https:\/\/www.aclweb.org\/anthology\/W13--2305"},{"key":"e_1_2_1_33_1","volume-title":"Weinberger","author":"Guo Chuan","year":"2017","unstructured":"Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. ArXiv, Vol. abs\/1706.04599 (2017)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3025453.3025781"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376375"},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Scott Huffman. 2008. Search evaluation at Google. https:\/\/googleblog.blogspot.com\/2008\/09\/search-evaluation-at-google.html.","DOI":"10.5040\/9798400658495"},{"key":"e_1_2_1_37_1","volume-title":"Machine Learning: An Introduction to Concepts and Methods. arXiv: Learning","author":"Hullermeier E.","year":"2019","unstructured":"E. Hullermeier and W. Waegeman. 2019. Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods. arXiv: Learning (2019)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Tao Jin Pan Xu Quanquan Gu and F. Farnoud. 2020. Rank Aggregation via Heterogeneous Thurstone Preference Models. In AAAI.","DOI":"10.1609\/aaai.v34i04.5860"},{"key":"e_1_2_1_39_1","volume-title":"Embracing Ambiguity: A Comparison of Annotation Methodologies for Crowdsourcing Word Sense Labels. In HLT-NAACL.","author":"Jurgens David","year":"2013","unstructured":"David Jurgens. 2013. Embracing Ambiguity: A Comparison of Annotation Methodologies for Crowdsourcing Word Sense Labels. In HLT-NAACL."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2818048.2820016"},{"key":"e_1_2_1_41_1","volume-title":"Genie: A leaderboard for human-in-the-loop evaluation of text generation. arXiv preprint arXiv:2101.06561","author":"Khashabi Daniel","year":"2021","unstructured":"Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld. 2021. Genie: A leaderboard for human-in-the-loop evaluation of text generation. arXiv preprint arXiv:2101.06561 (2021)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16--1095"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2556288.2557238"},{"key":"e_1_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Samuel L\"aubli Rico Sennrich and M. Volk. 2018. Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation. ArXiv Vol. abs\/1808.07048 (2018).","DOI":"10.18653\/v1\/D18-1512"},{"key":"e_1_2_1_45_1","doi-asserted-by":"crossref","unstructured":"Weixin Liang J. Zou and Zhou Yu. 2020. Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog Evaluation. In ACL.","DOI":"10.18653\/v1\/2020.acl-main.126"},{"key":"e_1_2_1_46_1","volume-title":"A technique for the measurement of attitudes. Archives of psychology","author":"Likert Rensis","year":"1932","unstructured":"Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of psychology (1932)."},{"key":"e_1_2_1_47_1","volume-title":"Weld","author":"Lin C. H.","year":"2014","unstructured":"C. H. Lin, Mausam, and Daniel S. Weld. 2014. To Re(label), or Not To Re(label). In HCOMP."},{"key":"e_1_2_1_48_1","volume-title":"Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence","author":"Lin Christopher H.","year":"1845","unstructured":"Christopher H. Lin, Mausam, and Daniel S. Weld. 2016. Re-Active Learning: Active Learning with Relabeling. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (Phoenix, Arizona) (AAAI'16). AAAI Press, 1845--1852."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1837885.1837907"},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of NAACL and HLT","author":"Liu Angli","year":"2016","unstructured":"Angli Liu, Stephen Soderland, Jonathan Bragg, Christopher H. Lin, Xiao Ling, and Daniel S. Weld. 2016. Effective Crowd Annotation for Relation Extraction. In Proceedings of NAACL and HLT 2016."},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","volume":"9","author":"Mitra Tanushree","year":"2015","unstructured":"Tanushree Mitra and Eric Gilbert. 2015. Credbank: A large-scale social media corpus with associated credibility annotations. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 9."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1609\/icwsm.v13i01.3261"},{"key":"e_1_2_1_53_1","volume-title":"The measurement of meaning. Number 47","author":"Osgood Charles Egerton","unstructured":"Charles Egerton Osgood, George J Suci, and Percy H Tannenbaum. 1957. The measurement of meaning. Number 47. University of Illinois press."},{"key":"e_1_2_1_54_1","volume-title":"Huai hsin Chi","author":"Qin Y.","year":"2020","unstructured":"Y. Qin, Xuezhi Wang, Alex Beutel, and Ed Huai hsin Chi. 2020. Improving Uncertainty Estimates through the Relationship with Adversarial Robustness. ArXiv, Vol. abs\/2006.16375 (2020)."},{"key":"e_1_2_1_55_1","volume-title":"Know What You Don't Know: Unanswerable Questions for SQuAD. ArXiv","author":"Rajpurkar Pranav","year":"2018","unstructured":"Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don't Know: Unanswerable Questions for SQuAD. ArXiv, Vol. abs\/1806.03822 (2018)."},{"key":"e_1_2_1_56_1","volume-title":"Efficient Online Scalar Annotation with Bounded Support. ArXiv","author":"Sakaguchi Keisuke","year":"2018","unstructured":"Keisuke Sakaguchi and Benjamin Van Durme. 2018. Efficient Online Scalar Annotation with Bounded Support. ArXiv, Vol. abs\/1806.01170 (2018)."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/SNAMS.2018.8554954"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3274423"},{"key":"e_1_2_1_59_1","unstructured":"Jo ao Sedoc Daphne Ippolito Arun Kirubarajan Jai Thirani L. Ungar and Chris Callison-Burch. 2019. ChatEval: A Tool for Chatbot Evaluation. In NAACL-HLT."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00026"},{"key":"e_1_2_1_61_1","volume-title":"Proceedings of the Fifth Conference on Machine Translation. Association for Computational Linguistics, Online, 76--91","author":"Specia Lucia","year":"2020","unstructured":"Lucia Specia, Zhenhao Li, Juan Pino, Vishrav Chaudhary, Francisco Guzm\u00e1n, Graham Neubig, Nadir Durrani, Yonatan Belinkov, Philipp Koehn, Hassan Sajjad, Paul Michel, and Xian Li. 2020. Findings of the WMT 2020 Shared Task on Machine Translation Robustness. In Proceedings of the Fifth Conference on Machine Translation. Association for Computational Linguistics, Online, 76--91. https:\/\/www.aclweb.org\/anthology\/2020.wmt-1.4"},{"key":"e_1_2_1_62_1","doi-asserted-by":"crossref","unstructured":"N. Stewart G. Brown and N. Chater. 2005. Absolute identification by relative judgment. Psychological review Vol. 112 4 (2005) 881--911.","DOI":"10.1037\/0033-295X.112.4.881"},{"key":"e_1_2_1_63_1","doi-asserted-by":"crossref","unstructured":"N. Stewart N. Chater and G. Brown. 2006. Decision by sampling.","DOI":"10.1016\/j.cogpsych.2005.10.003"},{"key":"e_1_2_1_64_1","volume-title":"A law of comparative judgment. Psychological review","author":"Thurstone Louis L","year":"1927","unstructured":"Louis L Thurstone. 1927. A law of comparative judgment. Psychological review, Vol. 34, 4 (1927), 273."},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2860987"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1080\/14992027.2016.1220680"},{"key":"e_1_2_1_67_1","volume-title":"Similarity Comparisons for Interactive Fine-Grained Categorization. 2014 IEEE Conference on Computer Vision and Pattern Recognition","author":"Wah C.","year":"2014","unstructured":"C. Wah, Grant Van Horn, Steve Branson, Subhransu Maji, P. Perona, and Serge J. Belongie. 2014. Similarity Comparisons for Interactive Fine-Grained Categorization. 2014 IEEE Conference on Computer Vision and Pattern Recognition (2014), 859--866."},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijresmar.2016.05.003"},{"key":"e_1_2_1_69_1","volume-title":"Metrology for AI: From Benchmarks to Instruments. arXiv preprint arXiv:1911.01875","author":"Welty Chris","year":"2019","unstructured":"Chris Welty, Praveen Paritosh, and Lora Aroyo. 2019. Metrology for AI: From Benchmarks to Instruments. arXiv preprint arXiv:1911.01875 (2019)."},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052591"},{"key":"e_1_2_1_71_1","volume-title":"Ranking with Uncertain Labels. 2007 IEEE International Conference on Multimedia and Expo (2007)","author":"Yan S.","unstructured":"S. Yan, H. Wang, T. Huang, Q. Yang, and X. Tang. 2007. Ranking with Uncertain Labels. 2007 IEEE International Conference on Multimedia and Expo (2007), 96--99."},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2008.2006585"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/2818048.2819953"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3158226"},{"key":"e_1_2_1_75_1","volume-title":"Quantifying Facial Age by Posterior of Age Comparisons. ArXiv","author":"Zhang Yunxuan","year":"2017","unstructured":"Yunxuan Zhang, Li Liu, Cheng Li, and Chen Change Loy. 2017. Quantifying Facial Age by Posterior of Age Comparisons. ArXiv, Vol. abs\/1708.09687 (2017)."}],"container-title":["Proceedings of the ACM on Human-Computer Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3476076","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3476076","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T04:55:49Z","timestamp":1752468949000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3476076"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,13]]},"references-count":75,"journal-issue":{"issue":"CSCW2","published-print":{"date-parts":[[2021,10,13]]}},"alternative-id":["10.1145\/3476076"],"URL":"https:\/\/doi.org\/10.1145\/3476076","relation":{},"ISSN":["2573-0142"],"issn-type":[{"value":"2573-0142","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,13]]},"assertion":[{"value":"2021-10-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}