{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T19:55:28Z","timestamp":1783367728893,"version":"3.54.6"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2020,12,31]],"date-time":"2020-12-31T00:00:00Z","timestamp":1609372800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000780","name":"European Commission","doi-asserted-by":"crossref","award":["823914, 951911, and 871042"],"award-info":[{"award-number":["823914, 951911, and 871042"]}],"id":[{"id":"10.13039\/501100000780","id-type":"DOI","asserted-by":"crossref"}]},{"name":"AI4Media project"},{"name":"SoBigData++ project"},{"name":"H2020 Programme","award":["INFRAIA-2018-1, ICT-48-2020, and INFRAIA-2019-1"],"award-info":[{"award-number":["INFRAIA-2018-1, ICT-48-2020, and INFRAIA-2019-1"]}]},{"name":"ARIADNEplus project"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2021,4,30]]},"abstract":"<jats:p>We critically re-examine the Saerens-Latinne-Decaestecker (SLD) algorithm, a well-known method for estimating class prior probabilities (\u201cpriors\u201d) and adjusting posterior probabilities (\u201cposteriors\u201d) in scenarios characterized by distribution shift, i.e., difference in the distribution of the priors between the training and the unlabelled documents. Given a machine learned classifier and a set of unlabelled documents for which the classifier has returned posterior probabilities and estimates of the prior probabilities, SLD updates them both in an iterative, mutually recursive way, with the goal of making both more accurate; this is of key importance in downstream tasks such as single-label multiclass classification and cost-sensitive text classification. Since its publication, SLD has become the standard algorithm for improving the quality of the posteriors in the presence of distribution shift, and SLD is still considered a top contender when we need to estimate the priors (a task that has become known as \u201cquantification\u201d). However, its real effectiveness in improving the quality of the posteriors has been questioned. We here present the results of systematic experiments conducted on a large, publicly available dataset, across multiple amounts of distribution shift and multiple learners. Our experiments show that SLD improves the quality of the posterior probabilities and of the estimates of the prior probabilities, but only when the number of classes in the classification scheme is very small and the classifier is calibrated. As the number of classes grows, or as we use non-calibrated classifiers, SLD converges more slowly (and often does not converge at all), performance degrades rapidly, and the impact of SLD on the quality of the prior estimates and of the posteriors becomes negative rather than positive.<\/jats:p>","DOI":"10.1145\/3433164","type":"journal-article","created":{"date-parts":[[2020,12,31]],"date-time":"2020-12-31T23:08:59Z","timestamp":1609456139000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["A Critical Reassessment of the Saerens-Latinne-Decaestecker Algorithm for Posterior Probability Adjustment"],"prefix":"10.1145","volume":"39","author":[{"given":"Andrea","family":"Esuli","sequence":"first","affiliation":[{"name":"Consiglio Nazionale delle Ricerche, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alessio","family":"Molinari","sequence":"additional","affiliation":[{"name":"Consiglio Nazionale delle Ricerche, Italy and Universit\u00e0 di Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fabrizio","family":"Sebastiani","sequence":"additional","affiliation":[{"name":"Consiglio Nazionale delle Ricerche, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,12,31]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3385656","article-title":"Better classifier calibration for small data sets","volume":"14","author":"Alasalmi Tuomo","year":"2020","unstructured":"Tuomo Alasalmi , Jaakko Suutala , Heli Koskim\u00e4ki , and Juha R\u00f6ning . 2020 . Better classifier calibration for small data sets . ACM Trans. Knowl. Discov. Data 14 , 3 (2020), 1 -- 19 . DOI:https:\/\/doi.org\/10.1145\/3385656 10.1145\/3385656 Tuomo Alasalmi, Jaakko Suutala, Heli Koskim\u00e4ki, and Juha R\u00f6ning. 2020. Better classifier calibration for small data sets. ACM Trans. Knowl. Discov. Data 14, 3 (2020), 1--19. DOI:https:\/\/doi.org\/10.1145\/3385656","journal-title":"ACM Trans. Knowl. Discov. Data"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-013-0308-z"},{"key":"e_1_2_1_3_1","volume-title":"Approaches for credit scorecard calibration: An empirical analysis. Knowl.-based Syst. 134","author":"Bequ\u00e9 Artem","year":"2017","unstructured":"Artem Bequ\u00e9 , Kristof Coussement , Ross W. Gayler , and Stefan Lessmann . 2017. Approaches for credit scorecard calibration: An empirical analysis. Knowl.-based Syst. 134 ( 2017 ), 213--227. DOI:https:\/\/doi.org\/10.1016\/j.knosys.2017.07.034 10.1016\/j.knosys.2017.07.034 Artem Bequ\u00e9, Kristof Coussement, Ross W. Gayler, and Stefan Lessmann. 2017. Approaches for credit scorecard calibration: An empirical analysis. Knowl.-based Syst. 134 (2017), 213--227. DOI:https:\/\/doi.org\/10.1016\/j.knosys.2017.07.034"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1175\/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000006"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2011.05.027"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.2307\/2987588"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.2517-6161.1977.tb01600.x"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 13th International Conference on Machine Learning (ICML\u201996)","author":"Pedro","unstructured":"Pedro M. Domingos and Michael J. Pazzani. 1996. Beyond independence: Conditions for the optimality of the simple Bayesian classifier . In Proceedings of the 13th International Conference on Machine Learning (ICML\u201996) . 105--112. Pedro M. Domingos and Michael J. Pazzani. 1996. Beyond independence: Conditions for the optimality of the simple Bayesian classifier. In Proceedings of the 13th International Conference on Machine Learning (ICML\u201996). 105--112."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3269287"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2020.2979203"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2005.10.010"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-005-5256-4"},{"key":"e_1_2_1_14_1","article-title":"Quantification under prior probability shift: The ratio estimator and its extensions","volume":"20","author":"Vaz Afonso Fernandes","year":"2019","unstructured":"Afonso Fernandes Vaz , Rafael Izbicki , and Rafael Bassi Stern . 2019 . Quantification under prior probability shift: The ratio estimator and its extensions . J. Mach. Learn. Res. 20 (2019), 79:1--79:33. Afonso Fernandes Vaz, Rafael Izbicki, and Rafael Bassi Stern. 2019. Quantification under prior probability shift: The ratio estimator and its extensions. J. Mach. Learn. Res. 20 (2019), 79:1--79:33.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_1_15_1","volume-title":"Encyclopedia of Machine Learning","author":"Flach Peter A.","unstructured":"Peter A. Flach . 2017. Classifier calibration . In Encyclopedia of Machine Learning ( 2 nd ed.), Claude Sammut and Geoffrey I. Webb (Eds.). Springer , DE, 212--219. Peter A. Flach. 2017. Classifier calibration. In Encyclopedia of Machine Learning (2nd ed.), Claude Sammut and Geoffrey I. Webb (Eds.). Springer, DE, 212--219.","edition":"2"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-008-0097-y"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13278-016-0327-z"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1198\/016214506000001437"},{"key":"e_1_2_1_19_1","volume-title":"A review on quantification learning. Comput. Surveys 50, 5","author":"Gonz\u00e1lez Pablo","year":"2017","unstructured":"Pablo Gonz\u00e1lez , Alberto Casta\u00f1o , Nitesh V. Chawla , and Juan Jos\u00e9 del Coz . 2017. A review on quantification learning. Comput. Surveys 50, 5 ( 2017 ), 74:1--74:40. DOI:https:\/\/doi.org\/10.1145\/3117807 10.1145\/3117807 Pablo Gonz\u00e1lez, Alberto Casta\u00f1o, Nitesh V. Chawla, and Juan Jos\u00e9 del Coz. 2017. A review on quantification learning. Comput. Surveys 50, 5 (2017), 74:1--74:40. DOI:https:\/\/doi.org\/10.1145\/3117807"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/1484611.1484627"},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 17th ACM International Conference on Research and Development in Information Retrieval (SIGIR\u201994)","author":"David","year":"2099","unstructured":"David D. Lewis and William A. Gale. 1994. A sequential algorithm for training text classifiers . In Proceedings of the 17th ACM International Conference on Research and Development in Information Retrieval (SIGIR\u201994) . 3--12. DOI:https:\/\/doi.org\/10.1007\/978-1-4471- 2099 -5_1 10.1007\/978-1-4471-2099-5_1 David D. Lewis and William A. Gale. 1994. A sequential algorithm for training text classifiers. In Proceedings of the 17th ACM International Conference on Research and Development in Information Retrieval (SIGIR\u201994). 3--12. DOI:https:\/\/doi.org\/10.1007\/978-1-4471-2099-5_1"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the 8th BCS-IRSG Symposium on Future Directions in Information Access (FDIA\u201919)","author":"Molinari Alessio","year":"2019","unstructured":"Alessio Molinari . 2019 . Leveraging the transductive nature of e-discovery in cost-sensitive technology-assisted review . In Proceedings of the 8th BCS-IRSG Symposium on Future Directions in Information Access (FDIA\u201919) . 72--78. Alessio Molinari. 2019. Leveraging the transductive nature of e-discovery in cost-sensitive technology-assisted review. In Proceedings of the 8th BCS-IRSG Symposium on Future Directions in Information Access (FDIA\u201919). 72--78."},{"key":"e_1_2_1_23_1","volume-title":"Risk Minimization Models for Technology-assisted Review and Their Application to e-discovery. Master\u2019s thesis. Department of Computer Science","author":"Molinari Alessio","unstructured":"Alessio Molinari . 2019. Risk Minimization Models for Technology-assisted Review and Their Application to e-discovery. Master\u2019s thesis. Department of Computer Science , University of Pisa , Pisa, IT . Alessio Molinari. 2019. Risk Minimization Models for Technology-assisted Review and Their Application to e-discovery. Master\u2019s thesis. Department of Computer Science, University of Pisa, Pisa, IT."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2011.06.019"},{"key":"e_1_2_1_25_1","unstructured":"Alejandro Moreo and Fabrizio Sebastiani. 2020. Tweet sentiment quantification: An experimental re-evaluation. Submitted for publication. https:\/\/arxiv.org\/abs\/2011.08091.  Alejandro Moreo and Fabrizio Sebastiani. 2020. Tweet sentiment quantification: An experimental re-evaluation. Submitted for publication. https:\/\/arxiv.org\/abs\/2011.08091."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1175\/1520-0450(1973)012<0595:ANVPOT>2.0.CO;2"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 29th AAAI Conference on Artificial Intelligence (AAAI\u201915)","author":"Naeini Mahdi P.","year":"2015","unstructured":"Mahdi P. Naeini , Gregory F. Cooper , and Milos Hauskrecht . 2015 . Obtaining well-calibrated probabilities using Bayesian binning . In Proceedings of the 29th AAAI Conference on Artificial Intelligence (AAAI\u201915) . 2901--2907. Mahdi P. Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining well-calibrated probabilities using Bayesian binning. In Proceedings of the 29th AAAI Conference on Artificial Intelligence (AAAI\u201915). 2901--2907."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 21st Conference Annual Conference on Uncertainty in Artificial Intelligence (UAI\u201905)","author":"Niculescu-Mizil Alexandru","year":"2005","unstructured":"Alexandru Niculescu-Mizil and Rich Caruana . 2005 . Obtaining calibrated probabilities from boosting . In Proceedings of the 21st Conference Annual Conference on Uncertainty in Artificial Intelligence (UAI\u201905) . 413--420. Alexandru Niculescu-Mizil and Rich Caruana. 2005. Obtaining calibrated probabilities from boosting. In Proceedings of the 21st Conference Annual Conference on Uncertainty in Artificial Intelligence (UAI\u201905). 413--420."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102430"},{"key":"e_1_2_1_30_1","article-title":"Jointly minimizing the expected costs of review for responsiveness and privilege in e-discovery","volume":"37","author":"Oard Douglas W.","year":"2018","unstructured":"Douglas W. Oard , Fabrizio Sebastiani , and Jyothi K. Vinjumur . 2018 . Jointly minimizing the expected costs of review for responsiveness and privilege in e-discovery . ACM Trans. Inf. Syst. 37 , 1 (2018), 11:1--11:35 pages. DOI:https:\/\/doi.org\/10.1145\/3268928 10.1145\/3268928 Douglas W. Oard, Fabrizio Sebastiani, and Jyothi K. Vinjumur. 2018. Jointly minimizing the expected costs of review for responsiveness and privilege in e-discovery. ACM Trans. Inf. Syst. 37, 1 (2018), 11:1--11:35 pages. DOI:https:\/\/doi.org\/10.1145\/3268928","journal-title":"ACM Trans. Inf. Syst."},{"key":"e_1_2_1_31_1","volume-title":"Advances in Large Margin Classifiers, Alexander Smola, Peter Bartlett","author":"Platt John C.","unstructured":"John C. Platt . 2000. Probabilistic outputs for support vector machines and comparison to regularized likelihood methods . In Advances in Large Margin Classifiers, Alexander Smola, Peter Bartlett , Bernard Sch\u00f6lkopf , and Dale Schuurmans (Eds.). The MIT Press , Cambridge, MA, 61--74. John C. Platt. 2000. Probabilistic outputs for support vector machines and comparison to regularized likelihood methods. In Advances in Large Margin Classifiers, Alexander Smola, Peter Bartlett, Bernard Sch\u00f6lkopf, and Dale Schuurmans (Eds.). The MIT Press, Cambridge, MA, 61--74."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2018.01.001"},{"key":"#cr-split#-e_1_2_1_33_1.1","doi-asserted-by":"crossref","unstructured":"Joaquin Qui\u00f1onero-Candela Masashi Sugiyama Anton Schwaighofer and Neil D. Lawrence (Eds.). 2009. Dataset Shift in Machine Learning. The MIT Press Cambridge MA. DOI:https:\/\/doi.org\/10.7551\/mitpress\/9780262170055.001.0001 10.7551\/mitpress","DOI":"10.7551\/mitpress\/9780262170055.001.0001"},{"key":"#cr-split#-e_1_2_1_33_1.2","doi-asserted-by":"crossref","unstructured":"Joaquin Qui\u00f1onero-Candela Masashi Sugiyama Anton Schwaighofer and Neil D. Lawrence (Eds.). 2009. Dataset Shift in Machine Learning. The MIT Press Cambridge MA. DOI:https:\/\/doi.org\/10.7551\/mitpress\/9780262170055.001.0001","DOI":"10.7551\/mitpress\/9780262170055.001.0001"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1162\/089976602753284446"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-019-09363-y"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341161.3342948"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1175\/2007WAF2006116.1"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10044-016-0578-3"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-3264-1"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.5555\/1005332.1016791"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/775047.775151"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3433164","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3433164","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:10Z","timestamp":1750193290000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3433164"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12,31]]},"references-count":42,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,4,30]]}},"alternative-id":["10.1145\/3433164"],"URL":"https:\/\/doi.org\/10.1145\/3433164","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12,31]]},"assertion":[{"value":"2020-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-12-31","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}