{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,4]],"date-time":"2026-03-04T03:04:26Z","timestamp":1772593466505,"version":"3.50.1"},"publisher-location":"Cham","reference-count":34,"publisher":"Springer International Publishing","isbn-type":[{"value":"9783030337773","type":"print"},{"value":"9783030337780","type":"electronic"}],"license":[{"start":{"date-parts":[[2019,1,1]],"date-time":"2019-01-01T00:00:00Z","timestamp":1546300800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Real-world datasets are often characterised by outliers, points far from the majority of the points, which might negatively influence modelling of the data. In data analysis it is hence important to use methods that are robust to outliers. In this paper we develop a robust regression method for finding the largest subset in the data that can be approximated using a sparse linear model to a given precision. We show that the problem is NP-hard and hard to approximate. We present an efficient algorithm, termed<jats:sc>slise<\/jats:sc>, to find solutions to the problem. Our method extends current state-of-the-art robust regression methods, especially in terms of scalability on large datasets. Furthermore, we show that our method can be used to yield interpretable explanations for individual decisions by opaque, black box, classifiers. Our approach solves shortcomings in other recent explanation methods by not requiring sampling of new data points and by being usable without modifications across various data domains. We demonstrate our method using both synthetic and real-world regression and classification problems.<\/jats:p>","DOI":"10.1007\/978-3-030-33778-0_27","type":"book-chapter","created":{"date-parts":[[2019,10,18]],"date-time":"2019-10-18T14:28:50Z","timestamp":1571408930000},"page":"351-366","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Sparse Robust Regression for Explaining Classifiers"],"prefix":"10.1007","author":[{"given":"Anton","family":"Bj\u00f6rklund","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andreas","family":"Henelius","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Emilia","family":"Oikarinen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kimmo","family":"Kallonen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Puolam\u00e4ki","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2019,10,16]]},"reference":[{"key":"27_CR1","doi-asserted-by":"crossref","unstructured":"Adler, P., et al.: Auditing black-box models for indirect influence. In: ICDM, pp. 1\u201310 (2016)","DOI":"10.1109\/ICDM.2016.0011"},{"issue":"1","key":"27_CR2","doi-asserted-by":"publisher","first-page":"226","DOI":"10.1214\/12-AOAS575","volume":"7","author":"A Alfons","year":"2013","unstructured":"Alfons, A., Croux, C., Gelper, S.: Sparse least trimmed squares regression for analyzing high-dimensional large data sets. Ann. Appl. Stat. 7(1), 226\u2013248 (2013)","journal-title":"Ann. Appl. Stat."},{"issue":"1","key":"27_CR3","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1016\/0304-3975(94)00254-G","volume":"147","author":"E Amaldi","year":"1995","unstructured":"Amaldi, E., Kann, V.: The complexity and approximability of finding maximum feasible subsystems of linear relations. Theor. Comput. Sci. 147(1), 181\u2013210 (1995)","journal-title":"Theor. Comput. Sci."},{"key":"27_CR4","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-58412-1","volume-title":"Complexity and Approximation: Combinatorial Optimization Problems and their Approximability Properties","author":"G Ausiello","year":"1999","unstructured":"Ausiello, G., Crescenzi, P., Gambosi, G., Kann, V., Marchetti-Spaccamela, A., Protasi, M.: Complexity and Approximation: Combinatorial Optimization Problems and their Approximability Properties, 2nd edn. Springer, Heidelberg (1999). https:\/\/doi.org\/10.1007\/978-3-642-58412-1","edition":"2"},{"key":"27_CR5","first-page":"1803","volume":"11","author":"D Baehrens","year":"2010","unstructured":"Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., M\u00fcller, K.: How to explain individual classification decisions. JMLR 11, 1803\u20131831 (2010)","journal-title":"JMLR"},{"key":"27_CR6","doi-asserted-by":"crossref","unstructured":"Caruana, R., Lou, Y., Gehrke, J., Koch, P., Sturm, M., Elhadad, N.: Intelligible models for healthcare: predicting pneumonia risk and hospital 30-day readmission. In: SIGKDD, pp. 1721\u20131730 (2015)","DOI":"10.1145\/2783258.2788613"},{"key":"27_CR7","unstructured":"CMS Collaboration: Performance of quark\/gluon discrimination in 8 TeV pp data. CMS-PAS-JME-13-002 (2013)"},{"key":"27_CR8","unstructured":"CMS Collaboration: Dataset QCD$$\\_$$Pt15to3000$$\\_$$TuneZ2star$$\\_$$Flat$$\\_$$8TeV$$\\_$$pythia6 in AODSIM format for 2012 collision data. CERN Open Data Portal (2017)"},{"key":"27_CR9","doi-asserted-by":"crossref","unstructured":"Cohen, G., Afshar, S., Tapson, J., van Schaik, A.: EMNIST: an extension of MNIST to handwritten letters. arXiv:1702.05373 (2017)","DOI":"10.1109\/IJCNN.2017.7966217"},{"key":"27_CR10","doi-asserted-by":"crossref","unstructured":"Datta, A., Sen, S., Zick, Y.: Algorithmic transparency via quantitative input influence: theory and experiments with learning systems. In: IEEE S&P, pp. 598\u2013617 (2016)","DOI":"10.1109\/SP.2016.42"},{"key":"27_CR11","unstructured":"Donoho, D.L., Huber, P.J.: The notion of breakdown point. In: A festschrift for Erich L. Lehmann, pp. 157\u2013184 (1983)"},{"key":"27_CR12","unstructured":"Finnish Grid and Cloud Infrastructure, urn:nbn:fi:research-infras-2016072533"},{"key":"27_CR13","doi-asserted-by":"crossref","unstructured":"Fong, R.C., Vedaldi, A.: Interpretable explanations of black boxes by meaningful perturbation. arXiv:1704.03296 (2017)","DOI":"10.1109\/ICCV.2017.371"},{"key":"27_CR14","unstructured":"Guidotti, R., Monreale, A., Ruggieri, S., Pedreschi, D., Turini, F., Giannotti, F.: Local rule-based explanations of black box decision systems. arXiv:1805.10820 (2018)"},{"issue":"5","key":"27_CR15","doi-asserted-by":"publisher","first-page":"93:1","DOI":"10.1145\/3236009","volume":"51","author":"R Guidotti","year":"2018","unstructured":"Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., Pedreschi, D.: A survey of methods for explaining black box models. CSUR 51(5), 93:1\u201393:42 (2018). https:\/\/doi.org\/10.1145\/3236009","journal-title":"CSUR"},{"issue":"5\u20136","key":"27_CR16","first-page":"1503","volume":"28","author":"A Henelius","year":"2014","unstructured":"Henelius, A., Puolam\u00e4ki, K., Bostr\u00f6m, H., Asker, L., Papapetrou, P.: A peek into the black box: exploring classifiers by randomization. DAMI 28(5\u20136), 1503\u20131529 (2014)","journal-title":"DAMI"},{"key":"27_CR17","unstructured":"Henelius, A., Puolam\u00e4ki, K., Ukkonen, A.: Interpreting classifiers through attribute interactions in datasets. In: WHI, pp. 8\u201313 (2017)"},{"key":"27_CR18","doi-asserted-by":"publisher","first-page":"110","DOI":"10.1007\/JHEP01(2017)110","volume":"01","author":"PT Komiske","year":"2017","unstructured":"Komiske, P.T., Metodiev, E.M., Schwartz, M.D.: Deep learning in color: towards automated quark\/gluon jet discrimination. JHEP 01, 110 (2017)","journal-title":"JHEP"},{"key":"27_CR19","doi-asserted-by":"crossref","unstructured":"Lakkaraju, H., Bach, S.H., Leskovec, J.: Interpretable decision sets: a joint framework for description and prediction. In: SIGKDD, pp. 1675\u20131684 (2016)","DOI":"10.1145\/2939672.2939874"},{"key":"27_CR20","unstructured":"Loh, P.L.: Scale calibration for high-dimensional robust regression. arXiv preprint arXiv:1811.02096 (2018)"},{"key":"27_CR21","unstructured":"Lundberg, S.M., Lee, S.I.: A unified approach to interpreting model predictions. In: NIPS, pp. 4765\u20134774 (2017)"},{"key":"27_CR22","unstructured":"Maas, A.L., Daly, R.E., Pham, P.T., Huang, D., Ng, A.Y., Potts, C.: Learning word vectors for sentiment analysis. In: ACL HLT, pp. 142\u2013150 (2011)"},{"key":"27_CR23","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1007\/978-3-319-14612-6_4","volume-title":"Energy Minimization Methods in Computer Vision and Pattern Recognition","author":"H Mobahi","year":"2015","unstructured":"Mobahi, H., Fisher, J.W.: On the link between gaussian homotopy continuation and convex envelopes. In: Tai, X.-C., Bae, E., Chan, T.F., Lysaker, M. (eds.) EMMCVPR 2015. LNCS, vol. 8932, pp. 43\u201356. Springer, Cham (2015). https:\/\/doi.org\/10.1007\/978-3-319-14612-6_4"},{"key":"27_CR24","unstructured":"Molnar, C.: Interpretable Machine Learning (2019). https:\/\/christophm.github.io\/interpretable-ml-book"},{"key":"27_CR25","doi-asserted-by":"crossref","unstructured":"Ribeiro, M.T., Singh, S., Guestrin, C.: Why should I trust you? Explaining the predictions of any classifier. In: SIGKDD, pp. 1135\u20131144 (2016)","DOI":"10.1145\/2939672.2939778"},{"issue":"388","key":"27_CR26","doi-asserted-by":"publisher","first-page":"871","DOI":"10.1080\/01621459.1984.10477105","volume":"79","author":"PJ Rousseeuw","year":"1984","unstructured":"Rousseeuw, P.J.: Least median of squares regression. J. Am. Stat. Assoc. 79(388), 871\u2013880 (1984)","journal-title":"J. Am. Stat. Assoc."},{"issue":"1","key":"27_CR27","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1002\/widm.2","volume":"1","author":"PJ Rousseeuw","year":"2011","unstructured":"Rousseeuw, P.J., Hubert, M.: Robust statistics for outlier detection. WIRES Data Min. Knowl. Discov. 1(1), 73\u201379 (2011)","journal-title":"WIRES Data Min. Knowl. Discov."},{"key":"27_CR28","first-page":"335","volume-title":"Data Analysis. Studies in Classification, Data Analysis, and Knowledge Organization","author":"PJ Rousseeuw","year":"2000","unstructured":"Rousseeuw, P.J., Van Driessen, K.: An algorithm for positive-breakdown regression based on concentration steps. In: Gaul, W., Opitz, O., Schader, M. (eds.) Data Analysis. Studies in Classification, Data Analysis, and Knowledge Organization, pp. 335\u2013346. Springer, Heidelberg (2000)"},{"key":"27_CR29","unstructured":"Schmidt, M., Berg, E., Friedlander, M., Murphy, K.: Optimizing costly functions with simple constraints: a limited-memory projected quasi-newton algorithm. In: AISTATS, pp. 456\u2013463 (2009)"},{"key":"27_CR30","doi-asserted-by":"publisher","first-page":"116","DOI":"10.1016\/j.csda.2017.02.002","volume":"111","author":"E Smucler","year":"2017","unstructured":"Smucler, E., Yohai, V.J.: Robust and sparse estimators for linear regression models. Comput. Stat. Data Anal. 111, 116\u2013130 (2017)","journal-title":"Comput. Stat. Data Anal."},{"issue":"1","key":"27_CR31","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","volume":"58","author":"R Tibshirani","year":"1996","unstructured":"Tibshirani, R.: Regression shrinkage and selection via the Lasso. J. R. Stat. Soc. Series. B Stat. Methodol. 58(1), 267\u2013288 (1996)","journal-title":"J. R. Stat. Soc. Series. B Stat. Methodol."},{"key":"27_CR32","unstructured":"Ustun, B., Traca, S., Rudin, C.: Supersparse linear integer models for interpretable classification. arXiv:1306.6677v6 (2014)"},{"issue":"3","key":"27_CR33","doi-asserted-by":"publisher","first-page":"347","DOI":"10.1198\/073500106000000251","volume":"25","author":"H Wang","year":"2007","unstructured":"Wang, H., Li, G., Jiang, G.: Robust regression shrinkage and consistent variable selection through the LAD-Lasso. J. Bus. Econ. Stat. 25(3), 347\u2013355 (2007)","journal-title":"J. Bus. Econ. Stat."},{"issue":"2","key":"27_CR34","doi-asserted-by":"publisher","first-page":"642","DOI":"10.1214\/aos\/1176350366","volume":"15","author":"VJ Yohai","year":"1987","unstructured":"Yohai, V.J.: High breakdown-point and high efficiency robust estimates for regression. Ann. Stat. 15(2), 642\u2013656 (1987). https:\/\/doi.org\/10.1214\/aos\/1176350366","journal-title":"Ann. Stat."}],"container-title":["Lecture Notes in Computer Science","Discovery Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-030-33778-0_27","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,25]],"date-time":"2024-07-25T06:31:33Z","timestamp":1721889093000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-030-33778-0_27"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019]]},"ISBN":["9783030337773","9783030337780"],"references-count":34,"URL":"https:\/\/doi.org\/10.1007\/978-3-030-33778-0_27","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019]]},"assertion":[{"value":"16 October 2019","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"DS","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Discovery Science","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Split","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Croatia","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2019","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"28 October 2019","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"30 October 2019","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"22","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"dis2019","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"http:\/\/ds2019.irb.hr\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Single-blind","order":1,"name":"type","label":"Type","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"EasyChair","order":2,"name":"conference_management_system","label":"Conference Management System","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"63","order":3,"name":"number_of_submissions_sent_for_review","label":"Number of Submissions Sent for Review","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"21","order":4,"name":"number_of_full_papers_accepted","label":"Number of Full Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"19","order":5,"name":"number_of_short_papers_accepted","label":"Number of Short Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"33% - The value is computed by the equation \"Number of Full Papers Accepted \/ Number of Submissions Sent for Review * 100\" and then rounded to a whole number.","order":6,"name":"acceptance_rate_of_full_papers","label":"Acceptance Rate of Full Papers","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"3","order":7,"name":"average_number_of_reviews_per_paper","label":"Average Number of Reviews per Paper","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"4","order":8,"name":"average_number_of_papers_per_reviewer","label":"Average Number of Papers per Reviewer","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Yes","order":9,"name":"external_reviewers_involved","label":"External Reviewers Involved","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}}]}}