{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T23:37:58Z","timestamp":1782517078966,"version":"3.54.5"},"reference-count":148,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T00:00:00Z","timestamp":1760572800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["GE1745016"],"award-info":[{"award-number":["GE1745016"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"UL Research Institutes through the Center for Advancing Safety of Machine Intelligence (CASMI) at Northwestern University"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Hum.-Comput. Interact."],"published-print":{"date-parts":[[2025,10,18]]},"abstract":"<jats:p>\n            Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the ''authenticity'' of student writing or the ''healthcare need'' of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply\n            <jats:italic toggle=\"yes\">problem (re)formulation strategies,<\/jats:italic>\n            such as\n            <jats:italic toggle=\"yes\">swapping<\/jats:italic>\n            out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or\n            <jats:italic toggle=\"yes\">composing<\/jats:italic>\n            multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction.\n          <\/jats:p>","DOI":"10.1145\/3757628","type":"journal-article","created":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T16:59:10Z","timestamp":1760633950000},"page":"1-37","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-3566-9429","authenticated-orcid":false,"given":"Luke","family":"Guerdan","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5566-7409","authenticated-orcid":false,"given":"Devansh","family":"Saxena","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison, Madison, Wisconsin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0620-0903","authenticated-orcid":false,"given":"Stevie","family":"Chancellor","sequence":"additional","affiliation":[{"name":"University of Minnesota, Minneapolis, MN, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8125-8227","authenticated-orcid":false,"given":"Zhiwei Steven","family":"Wu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6730-922X","authenticated-orcid":false,"given":"Kenneth","family":"Holstein","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,16]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2016. https:\/\/www2.ed.gov\/rschstat\/eval\/high-school\/early-warning-systems-brief.pdf"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594083"},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Robert Adcock and David Collier. 2001. Measurement validity: A shared standard for qualitative and quantitative research. American political science review 95 3 529-546.","DOI":"10.1017\/S0003055401003100"},{"key":"e_1_2_1_4_1","volume-title":"Fairsight: Visual analytics for fairness in decision making","author":"Ahn Yongsu","year":"2019","unstructured":"Yongsu Ahn and Yu-Ru Lin. 2019. Fairsight: Visual analytics for fairness in decision making. IEEE transactions on visualization and computer graphics 26, 1, 1086-1095."},{"key":"e_1_2_1_5_1","volume-title":"Inioluwa Deborah Raji, and Travis Zack","author":"Alaa Ahmed","year":"2025","unstructured":"Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji, and Travis Zack. 2025. Medical Large Language Model Benchmarks Should Prioritize Construct Validity. arXiv preprint arXiv:2503.10694 (2025)."},{"key":"e_1_2_1_6_1","volume-title":"Futzing and moseying: Interviews with professional data analysts on exploration practices","author":"Alspaugh Sara","unstructured":"Sara Alspaugh, Nava Zokaei, Andrea Liu, Cindy Jin, and Marti A Hearst. 2018. Futzing and moseying: Interviews with professional data analysts on exploration practices. IEEE transactions on visualization and computer graphics 25, 1, 22-31."},{"key":"e_1_2_1_7_1","volume-title":"Rethinking value-added models in education: Critical perspectives on tests and assessment-based accountability","author":"Amrein-Beardsley Audrey","unstructured":"Audrey Amrein-Beardsley. 2014. Rethinking value-added models in education: Critical perspectives on tests and assessment-based accountability. Routledge."},{"key":"e_1_2_1_8_1","volume-title":"Elizabeth M Daly, Rahul Nair, Tejaswini Pedapati, Swapnaja Achintalwar, and Werner Geyer.","author":"Ashktorab Zahra","year":"2024","unstructured":"Zahra Ashktorab, Michael Desmond, Qian Pan, James M Johnson, Martin Santillan Cooper, Elizabeth M Daly, Rahul Nair, Tejaswini Pedapati, Swapnaja Achintalwar, and Werner Geyer. 2024. Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences. arXiv preprint arXiv:2410.00873 (2024)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Ted Baker and Reed E Nelson. 2005. Creating something from nothing: Resource construction through entrepreneurial bricolage. Administrative science quarterly 50 3 329-366.","DOI":"10.2189\/asqu.2005.50.3.329"},{"key":"e_1_2_1_10_1","volume-title":"Latent variable models and factor analysis: A unified approach","author":"Bartholomew David J","unstructured":"David J Bartholomew, Martin Knott, and Irini Moustaki. 2011. Latent variable models and factor analysis: A unified approach. John Wiley & Sons."},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Anton Barua Stephen W Thomas and Ahmed E Hassan. 2014. What are developers talking about? an analysis of topics and trends in stack overflow. Empirical software engineering 19 619-654.","DOI":"10.1007\/s10664-012-9231-y"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1147\/JRD.2019.2942287"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642106"},{"key":"e_1_2_1_14_1","volume-title":"Learning in graphical models","author":"Bishop Christopher M","unstructured":"Christopher M Bishop. 1998. Latent variable models. In Learning in graphical models. Springer, 371-403."},{"key":"e_1_2_1_15_1","first-page":"709","article-title":"Dataset discovery in data lakes. In 2020 ieee 36th international conference on data engineering (icde)","author":"Bogatu Alex","year":"2020","unstructured":"Alex Bogatu, Alvaro AA Fernandes, Norman W Paton, and Nikolaos Konstantinou. 2020. Dataset discovery in data lakes. In 2020 ieee 36th international conference on data engineering (icde). IEEE, 709-720.","journal-title":"IEEE"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/380681"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242598"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","unstructured":"Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3 2 77-101.","DOI":"10.1191\/1478088706qp063oa"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011293210539"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/VAST47406.2019.8986948"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544548.3581268"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3542921"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589317"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3025453.3026044"},{"key":"e_1_2_1_25_1","first-page":"1","article-title":"Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement","author":"Chen Quan Ze","year":"2023","unstructured":"Quan Ze Chen and Amy X Zhang. 2023. Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2, 1-26.","journal-title":"Proceedings of the ACM on Human-Computer Interaction 7, CSCW2"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3491102.3501831"},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Victoria Clarke and Virginia Braun. 2017. Thematic analysis. The journal of positive psychology 12 3 297-298.","DOI":"10.1080\/17439760.2016.1262613"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/SaTML54575.2023.00050"},{"key":"e_1_2_1_29_1","volume-title":"The measure of reality: Quantification in Western Europe, 1250-1600","author":"Crosby Alfred W","unstructured":"Alfred W Crosby. 1997. The measure of reality: Quantification in Western Europe, 1250-1600. Cambridge University Press."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the Twelfth Language Resources and Evaluation Conference. 7053-7059","author":"Daudert Tobias","year":"2020","unstructured":"Tobias Daudert. 2020. A web-based collaborative annotation and consolidation tool. In Proceedings of the Twelfth Language Resources and Evaluation Conference. 7053-7059."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533113"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397481.3450698"},{"key":"e_1_2_1_33_1","volume-title":"Erik H H M Korsten, Arthur R A Bouwman, and Jarke Van Wijk.","author":"Dingen Dennis","year":"2018","unstructured":"Dennis Dingen, Marcel van 't Veer, Patrick Houthuizen, Eveline H J M Mestrom, Erik H H M Korsten, Arthur R A Bouwman, and Jarke Van Wijk. 2018. RegressionExplorer: Interactive exploration of logistic regression models with subgroup analysis. IEEE transactions on visualization and computer graphics 25, 1, 246-255."},{"key":"e_1_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Ellen A Drost. 2011. Validity and reliability in social science research. Education Research and perspectives 38 1 105-123.","DOI":"10.70953\/ERPv38.11005"},{"key":"e_1_2_1_35_1","volume-title":"The art of feature engineering: essentials for machine learning","author":"Duboue Pablo","unstructured":"Pablo Duboue. 2020. The art of feature engineering: essentials for machine learning. Cambridge University Press."},{"key":"e_1_2_1_36_1","volume-title":"Towards a foundation of bricolage in organization and management theory","author":"Duymedjian Raffi","unstructured":"Raffi Duymedjian and Charles-Clemens R\u00fcling. 2010. Towards a foundation of bricolage in organization and management theory. Organization studies 31, 2, 133-151."},{"key":"e_1_2_1_37_1","volume-title":"Design for services","author":"Evenson Shelley","unstructured":"Shelley Evenson. 2016. Driving Service Design By Directed Storytelling. In Design for services. Routledge, 66-72."},{"key":"e_1_2_1_38_1","volume-title":"An introduction to latent variable models","author":"Everett B","unstructured":"B Everett. 2013. An introduction to latent variable models. Springer Science & Business Media."},{"key":"e_1_2_1_39_1","volume-title":"2018 IEEE 34th International Conference on Data Engineering (ICDE). IEEE, 1001-1012","author":"Fernandez Raul Castro","year":"2018","unstructured":"Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker. 2018. Aurum: A data discovery system. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). IEEE, 1001-1012."},{"key":"e_1_2_1_40_1","volume-title":"An Analysis of Jane Jacobs's The Death and Life of Great American Cities","author":"Fuller Martin","unstructured":"Martin Fuller and Ryan Moore. 2017. An Analysis of Jane Jacobs's The Death and Life of Great American Cities. Macat Library."},{"key":"e_1_2_1_41_1","volume-title":"Ray Eitel-Porter, et al.","author":"Gala Dalia","year":"2024","unstructured":"Dalia Gala, Milo Phillips-Brown, Naman Goel, Carinal Prunkl, Laura Alvarez Jubete, Ray Eitel-Porter, et al. 2024. FairTargetSim: An Interactive Simulator for Understanding and Explaining the Fairness Effects of Target Variable Definition. arXiv preprint arXiv:2403.06031."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458723"},{"key":"e_1_2_1_43_1","volume-title":"Ways of Knowing in HCI","author":"Gergle Darren","unstructured":"Darren Gergle and Desney S Tan. 2014. Experimental research in HCI. In Ways of Knowing in HCI. Springer, 191-227."},{"key":"e_1_2_1_44_1","volume-title":"Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley.","author":"Goldberger Ary L","year":"2000","unstructured":"Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley. 2000. PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals. circulation 101, 23, e215-e220."},{"key":"e_1_2_1_45_1","first-page":"1","article-title":"Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation","author":"Goyal Nitesh","year":"2022","unstructured":"Nitesh Goyal, Ian D Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2, 1-28.","journal-title":"Proceedings of the ACM on Human-Computer Interaction 6, CSCW2"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594101"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594036"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2047196.2047205"},{"key":"e_1_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Ian Hacking. 1999. The social construction of what? Harvard university press.","DOI":"10.2307\/j.ctv1bzfp1z"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642738"},{"key":"e_1_2_1_51_1","unstructured":"Bill Harding. 2021. Software effort estimates vs popular developer productivity metrics. Technical Report. GitClear."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642004"},{"key":"e_1_2_1_53_1","volume-title":"Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems. arXiv preprint arXiv:2506","author":"Harvey Emma","year":"2025","unstructured":"Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, and Hanna Wallach. 2025. Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems. arXiv preprint arXiv:2506.04482 (2025)."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-24853-8_45"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3714319"},{"key":"e_1_2_1_56_1","first-page":"355","volume-title":"Comment: Snowball versus respondent-driven sampling. Sociological methodology 41, 1","author":"Heckathorn Douglas D","year":"2011","unstructured":"Douglas D Heckathorn. 2011. Comment: Snowball versus respondent-driven sampling. Sociological methodology 41, 1, 355-366."},{"key":"e_1_2_1_57_1","volume-title":"Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories. arXiv preprint arXiv:2501.15114","author":"Hoess Nicole","year":"2025","unstructured":"Nicole Hoess, Carlos Paradis, Rick Kazman, and Wolfgang Mauerer. 2025. Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories. arXiv preprint arXiv:2501.15114 (2025)."},{"key":"e_1_2_1_58_1","unstructured":"Jake M Hofman Angelos Chatzimparmpas Amit Sharma Duncan J Watts and Jessica Hullman. 2023. Pre-registration for predictive modeling. arXiv preprint arXiv:2311.18807."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300830"},{"key":"e_1_2_1_60_1","first-page":"1","article-title":"Hacking with NPOs: collaborative analytics and broker roles in civic data hackathons","author":"Hou Youyang","year":"2017","unstructured":"Youyang Hou and Dakuo Wang. 2017. Hacking with NPOs: collaborative analytics and broker roles in civic data hackathons. Proceedings of the ACM on Human-Computer Interaction 1, CSCW, 1-16.","journal-title":"Proceedings of the ACM on Human-Computer Interaction 1, CSCW"},{"key":"e_1_2_1_61_1","volume-title":"NCES 2004-405","author":"Ingels Steven J","year":"2004","unstructured":"Steven J Ingels, Daniel J Pratt, James E Rogers, Peter H Siegel, and Ellen S Stutts. 2004. Education Longitudinal Study of 2002: Base Year Data File User's Manual. NCES 2004-405. National Center for Education Statistics."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445901"},{"key":"e_1_2_1_63_1","volume-title":"Leo Anthony Celi, and Roger Mark","author":"Johnson Alistair","year":"2020","unstructured":"Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. 2020. Mimic-iv. PhysioNet. Available online at: https:\/\/physionet.org\/content\/mimiciv\/1.0\/(accessed August 23, 2021), 49-55."},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3476980"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3491102.3501888"},{"key":"e_1_2_1_66_1","volume-title":"Making the Right Thing: Bridging HCI and Responsible AI in Early-Stage AI Concept Selection. arXiv preprint arXiv:2506.17494","author":"Jung Ji-Youn","year":"2025","unstructured":"Ji-Youn Jung, Devansh Saxena, Minjung Park, Jini Kim, Jodi Forlizzi, Kenneth Holstein, and John Zimmerman. 2025. Making the Right Thing: Bridging HCI and Responsible AI in Early-Stage AI Concept Selection. arXiv preprint arXiv:2506.17494 (2025)."},{"key":"e_1_2_1_67_1","volume-title":"Companion Proceedings of The 2019 World Wide Web Conference. 1121-1130","author":"Chaithanya Manam V K.","unstructured":"V K. Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J. Quinn. 2019. Taskmate: A mechanism to improve the quality of instructions in crowdsourcing. In Companion Proceedings of The 2019 World Wide Web Conference. 1121-1130."},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/1978942.1979444"},{"key":"e_1_2_1_69_1","volume-title":"Enterprise data analysis and visualization: An interview study","author":"Kandel Sean","unstructured":"Sean Kandel, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer. 2012. Enterprise data analysis and visualization: An interview study. IEEE transactions on visualization and computer graphics 18, 12, 2917-2926."},{"key":"e_1_2_1_70_1","volume-title":"Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4, 9","author":"Kapoor Sayash","year":"2023","unstructured":"Sayash Kapoor and Arvind Narayanan. 2023. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4, 9 (2023)."},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642849"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3491102.3517439"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLHCC.2017.8103446"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3173748"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/2884781.2884783"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/1922649.1922658"},{"key":"e_1_2_1_77_1","first-page":"1","article-title":"Orienting, framing, bridging, magic, and counseling: How data scientists navigate the outer loop of client collaborations in industry and academia","author":"Kross Sean","year":"2021","unstructured":"Sean Kross and Philip Guo. 2021. Orienting, framing, bridging, magic, and counseling: How data scientists navigate the outer loop of client collaborations in industry and academia. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2, 1-28.","journal-title":"Proceedings of the ACM on Human-Computer Interaction 5, CSCW2"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544548.3580882"},{"key":"e_1_2_1_79_1","volume-title":"Laboratory life: The construction of scientific facts","author":"Latour Bruno","unstructured":"Bruno Latour and Steve Woolgar. 2013. Laboratory life: The construction of scientific facts. Princeton university press."},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP52600.2021.00021"},{"key":"e_1_2_1_81_1","volume-title":"The savage mind","author":"Levi-Strauss Claude","unstructured":"Claude Levi-Strauss. 1966. The savage mind. University of Chicago Press."},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.5772\/intechopen.111293"},{"key":"e_1_2_1_83_1","volume-title":"Jackie Chi Kit Cheung, Q Vera Liao, Alexandra Olteanu, and Ziang Xiao.","author":"Liu Yu Lu","year":"2024","unstructured":"Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Q Vera Liao, Alexandra Olteanu, and Ziang Xiao. 2024. ECBD: Evidence-centered benchmark design for NLP. arXiv preprint arXiv:2406.08723 (2024)."},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0142-694X(98)00044-1"},{"key":"e_1_2_1_85_1","unstructured":"Scott Lundberg. 2017. A unified approach to interpreting model predictions. arXiv preprint arXiv:1705.07874."},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1609\/hcomp.v6i1.13338"},{"key":"e_1_2_1_87_1","first-page":"1","article-title":"How data scientistswork together with domain experts in scientific collaborations: To find the right answer or to ask the right question","author":"Mao Yaoli","year":"2019","unstructured":"Yaoli Mao, Dakuo Wang, Michael Muller, Kush R Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilovic. 2019. How data scientistswork together with domain experts in scientific collaborations: To find the right answer or to ask the right question? Proceedings of the ACM on Human-Computer Interaction 3, GROUP, 1-23.","journal-title":"Proceedings of the ACM on Human-Computer Interaction 3, GROUP"},{"key":"e_1_2_1_88_1","first-page":"1","article-title":"Bricolage-a systematic review, conceptualization, and research agenda","author":"Mateus Sara","year":"2024","unstructured":"Sara Mateus and Soumodip Sarkar. 2024. Bricolage-a systematic review, conceptualization, and research agenda. Entrepreneurship & Regional Development, 1-22.","journal-title":"Entrepreneurship & Regional Development"},{"key":"e_1_2_1_89_1","volume-title":"Racial\/Ethnic Categories in AI and Algorithmic Fairness: Why They Matter and What They Represent. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2484-2494","author":"Mickel Jennifer","year":"2024","unstructured":"Jennifer Mickel. 2024. Racial\/Ethnic Categories in AI and Algorithmic Fairness: Why They Matter and What They Represent. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2484-2494."},{"key":"e_1_2_1_90_1","volume-title":"Handbook of research design and social measurement","author":"Miller Delbert C","unstructured":"Delbert C Miller and Neil J Salkind. 2002. Handbook of research design and social measurement. Sage."},{"key":"e_1_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445933"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287560.3287596"},{"key":"e_1_2_1_93_1","volume-title":"The body multiple: Ontology in medical practice","author":"Mol A","unstructured":"A Mol. 2002. The body multiple: Ontology in medical practice. Duke University Press."},{"key":"e_1_2_1_94_1","first-page":"37","article-title":"On the inequity of predicting A while hoping for B. In AEA Papers and Proceedings, Vol. 111. American Economic Association 2014 Broadway","volume":"37203","author":"Mullainathan Sendhil","year":"2021","unstructured":"Sendhil Mullainathan and Ziad Obermeyer. 2021. On the inequity of predicting A while hoping for B. In AEA Papers and Proceedings, Vol. 111. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 37-42.","journal-title":"Suite 305, Nashville, TN"},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300356"},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445402"},{"key":"e_1_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13194-022-00484-8"},{"key":"e_1_2_1_98_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2016.200"},{"key":"e_1_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510209"},{"key":"e_1_2_1_100_1","unstructured":"Hiroki Nakayama Takahiro Kubo Junya Kamura Yasufumi Taniguchi and Xu Liang. 2018. doccano: Text Annotation Tool for Human. https:\/\/github.com\/doccano\/doccano Software available from https:\/\/github.com\/doccano\/doccano."},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352116"},{"key":"e_1_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.aax2342"},{"key":"e_1_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1109\/SEAA.2011.69"},{"key":"e_1_2_1_104_1","volume-title":"James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer.","author":"Pan Qian","year":"2024","unstructured":"Qian Pan, Zahra Ashktorab, Michael Desmond, Martin Santillan Cooper, James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer. 2024. Human-Centered Design Recommendations for LLM-as-a-judge. arXiv preprint arXiv:2407.03479 (2024)."},{"key":"e_1_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.3102\/1076998619872761"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287560.3287567"},{"key":"e_1_2_1_107_1","first-page":"1","article-title":"Trust in data science: Collaboration, translation, and accountability in corporate data science projects","author":"Passi Samir","year":"2018","unstructured":"Samir Passi and Steven J Jackson. 2018. Trust in data science: Collaboration, translation, and accountability in corporate data science projects. Proceedings of the ACM on human-computer interaction 2, CSCW, 1-28.","journal-title":"Proceedings of the ACM on human-computer interaction 2, CSCW"},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626521"},{"key":"e_1_2_1_109_1","unstructured":"Juan Carlos Perdomo. 2023. The Relative Value of Prediction in Algorithmic Decision Making. arXiv preprint arXiv:2312.08511."},{"key":"e_1_2_1_110_1","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702298"},{"key":"e_1_2_1_111_1","volume-title":"Trust in numbers: The pursuit of objectivity in science and public life","author":"Porter Theodore M","unstructured":"Theodore M Porter. 1996. Trust in numbers: The pursuit of objectivity in science and public life. Princeton University Press."},{"key":"e_1_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533231"},{"key":"e_1_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1138-y"},{"key":"e_1_2_1_114_1","volume-title":"AI and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366","author":"Raji Inioluwa Deborah","year":"2021","unstructured":"Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021. AI and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366 (2021)."},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533158"},{"key":"e_1_2_1_116_1","volume-title":"Principles of data wrangling: Practical techniques for data preparation. '' O'Reilly Media","author":"Rattenbury Tye","unstructured":"Tye Rattenbury, Joseph M Hellerstein, Jeffrey Heer, Sean Kandel, and Connor Carreras. 2017. Principles of data wrangling: Practical techniques for data preparation. '' O'Reilly Media, Inc.''."},{"key":"e_1_2_1_117_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-3020"},{"key":"e_1_2_1_118_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11491"},{"key":"e_1_2_1_119_1","volume-title":"Aequitas: A bias and fairness audit toolkit. arXiv preprint arXiv:1811.05577.","author":"Saleiro Pedro","year":"2018","unstructured":"Pedro Saleiro, Benedict Kuester, Loren Hinkson, Jesse London, Abby Stevens, Ari Anisfeld, Kit T Rodolfa, and Rayid Ghani. 2018. Aequitas: A bias and fairness audit toolkit. arXiv preprint arXiv:1811.05577."},{"key":"e_1_2_1_120_1","volume-title":"Saturation in qualitative research: exploring its conceptualization and operationalization. Quality & quantity 52","author":"Saunders Benjamin","year":"1893","unstructured":"Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. 2018. Saturation in qualitative research: exploring its conceptualization and operationalization. Quality & quantity 52, 1893-1907."},{"key":"e_1_2_1_121_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3714098"},{"key":"e_1_2_1_122_1","volume-title":"Case-control studies: design, conduct, analysis","author":"Schlesselman James J","unstructured":"James J Schlesselman. 1982. Case-control studies: design, conduct, analysis. Vol. 2. Oxford university press."},{"key":"e_1_2_1_123_1","doi-asserted-by":"publisher","DOI":"10.3389\/frma.2022.861944"},{"key":"e_1_2_1_124_1","doi-asserted-by":"crossref","unstructured":"James C Scott. 2020. Seeing like a state: How certain schemes to improve the human condition have failed. yale university Press.","DOI":"10.12987\/9780300252989"},{"key":"e_1_2_1_125_1","volume-title":"Human-centered software engineering - integrating usability in the software development lifecycle","author":"Seffah Ahmed","unstructured":"Ahmed Seffah, Jan Gulliksen, and Michel C Desmarais. 2005. Human-centered software engineering - integrating usability in the software development lifecycle. Vol. 8. Springer Science & Business Media."},{"key":"e_1_2_1_126_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654777.3676450"},{"key":"e_1_2_1_127_1","unstructured":"Amit Sharma and Emre Kiciman. 2020. DoWhy: An end-to-end library for causal inference. arXiv preprint arXiv:2011.04216."},{"key":"e_1_2_1_128_1","volume-title":"Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems 36.","author":"Shen Yongliang","year":"2024","unstructured":"Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems 36."},{"key":"e_1_2_1_129_1","volume-title":"Quantitative social research methods","author":"Singh Kultar","unstructured":"Kultar Singh. 2007. Quantitative social research methods. Sage."},{"key":"e_1_2_1_130_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3713664"},{"key":"e_1_2_1_131_1","volume-title":"Proceedings from the policing research institute meetings. Department of Justice, 37-88","author":"Skogan Wesley G","year":"1999","unstructured":"Wesley G Skogan. 1999. Measuring what matters. In Proceedings from the policing research institute meetings. Department of Justice, 37-88."},{"key":"e_1_2_1_132_1","unstructured":"Jonathan Stray Ivan Vendrov Jeremy Nixon Steven Adler and Dylan Hadfield-Menell. 2021. What are youo ptimizing for? aligning recommender systems with human values. arXiv preprint arXiv:2107.10939."},{"key":"e_1_2_1_133_1","doi-asserted-by":"crossref","unstructured":"Lucy Suchman. 1993. Do categories have politics? The language\/action perspective reconsidered. Computer supported cooperative work (CSCW) 2 177-190.","DOI":"10.1007\/BF00749015"},{"key":"e_1_2_1_134_1","volume-title":"Oghenemaro Anuyah, Ronald A Metoyer, and Toby Jia-Jun Li.","author":"Szymanski Annalisa","year":"2024","unstructured":"Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A Metoyer, and Toby Jia-Jun Li. 2024. Comparing Criteria Development Across Domain Experts, Lay Users, and Models in Large Language Model Evaluation. arXiv preprint arXiv:2410.02054 (2024)."},{"key":"e_1_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.1109\/EuroSP.2017.29"},{"key":"e_1_2_1_136_1","volume-title":"Life on the Screen","author":"Turkle Sherry","unstructured":"Sherry Turkle. 2011. Life on the Screen. Simon and Schuster."},{"key":"e_1_2_1_137_1","unstructured":"Bill Turque. 2012. Creative... motivating'and fired. The Washington Post 6."},{"key":"e_1_2_1_138_1","doi-asserted-by":"publisher","DOI":"10.1145\/2677199.2680594"},{"key":"e_1_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jesp.2016.03.004"},{"key":"e_1_2_1_140_1","volume-title":"Emily Corvi, et al.","author":"Wallach Hanna","year":"2024","unstructured":"Hanna Wallach, Meera Desai, Nicholas Pangakis, A Feder Cooper ,Angelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, et al. 2024. Evaluating Generative AI Systems is a Social Science Measurement Challenge. arXiv preprint arXiv:2411.10939 (2024)."},{"key":"e_1_2_1_141_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3593998"},{"issue":"257","key":"e_1_2_1_142_1","first-page":"1","article-title":"Fairlearn: Assessing and improving fairness of ai systems","volume":"24","author":"Weerts Hilde","year":"2023","unstructured":"Hilde Weerts, Miroslav Dud\u00edk, Richard Edgar, Adrin Jalali, Roman Lutz, and Michael Madaio. 2023. Fairlearn: Assessing and improving fairness of ai systems. Journal of Machine Learning Research 24, 257, 1-8.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_143_1","volume-title":"The what-if tool: Interactive probing of machine learning models","author":"Wexler James","unstructured":"James Wexler, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda Viegas, and Jimbo Wilson. 2019. The what-if tool: Interactive probing of machine learning models. IEEE transactions on visualization and computer graphics 26, 1, 56-65."},{"key":"e_1_2_1_144_1","unstructured":"Kanit Wongsuphasawat Yang Liu and Jeffrey Heer. 2019. Goals process and challenges of exploratory data analysis: An interview study. arXiv preprint arXiv:1911.00568."},{"key":"e_1_2_1_145_1","volume-title":"Voyager: Exploratory analysis via faceted browsing of visualization recommendations","author":"Wongsuphasawat Kanit","year":"2015","unstructured":"Kanit Wongsuphasawat, Dominik Moritz, Anushka Anand, Jock Mackinlay, Bill Howe, and Jeffrey Heer. 2015. Voyager: Exploratory analysis via faceted browsing of visualization recommendations. IEEE transactions on visualization and computer graphics 22, 1, 649-658."},{"key":"e_1_2_1_146_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41597-022-01782-9"},{"key":"e_1_2_1_147_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445728"},{"key":"e_1_2_1_148_1","first-page":"1","article-title":"How do data science workers collaborate? roles, workflows, and tools","author":"Zhang Amy X","year":"2020","unstructured":"Amy X Zhang, Michael Muller, and Dakuo Wang. 2020. How do data science workers collaborate? roles, workflows, and tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1, 1-23.","journal-title":"Proceedings of the ACM on Human-Computer Interaction 4, CSCW1"}],"container-title":["Proceedings of the ACM on Human-Computer Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3757628","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3757628","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T01:55:50Z","timestamp":1760666150000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3757628"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,16]]},"references-count":148,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2025,10,18]]}},"alternative-id":["10.1145\/3757628"],"URL":"https:\/\/doi.org\/10.1145\/3757628","relation":{},"ISSN":["2573-0142"],"issn-type":[{"value":"2573-0142","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,16]]},"assertion":[{"value":"2025-10-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}