{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T14:42:08Z","timestamp":1740148928157,"version":"3.37.3"},"reference-count":19,"publisher":"Springer Science and Business Media LLC","issue":"9-10","license":[{"start":{"date-parts":[[2020,8,29]],"date-time":"2020-08-29T00:00:00Z","timestamp":1598659200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,8,29]],"date-time":"2020-08-29T00:00:00Z","timestamp":1598659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Ann. Telecommun."],"published-print":{"date-parts":[[2020,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Many current and future applications plan to provide entity-specific predictions. These range from individualized healthcare applications to user-specific purchase recommendations. In our previous stream-based work on Amazon review data, we could show that error-weighted ensembles that combine entity-centric classifiers, which are only trained on reviews of one particular product (entity), and entity-ignorant classifiers, which are trained on all reviews irrespective of the product, can improve prediction quality. This came at the cost of storing multiple entity-centric models in primary memory, many of which would never be used again as their entities would not receive future instances in the stream. To overcome this drawback and make entity-centric learning viable in these scenarios, we investigated two different methods of reducing the primary memory requirement of our entity-centric approach. Our first method uses the lossy counting algorithm for data streams to identify entities whose instances make up a certain percentage of the total data stream within an error-margin. We then store all models which do not fulfil this requirement in secondary memory, from which they can be retrieved in case future instances belonging to them should arrive later in the stream. The second method replaces entity-centric models with a much more naive model which only stores the past labels and predicts the majority label seen so far. We applied our methods on the previously used Amazon data sets which contained up to 1.4M reviews and added two subsets of the Yelp data set which contain up to 4.2M reviews. Both methods were successful in reducing the primary memory requirements while still outperforming an entity-ignorant model.<\/jats:p>","DOI":"10.1007\/s12243-020-00800-4","type":"journal-article","created":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T07:50:59Z","timestamp":1598860259000},"page":"549-561","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Resource management for model learning at entity level"],"prefix":"10.1007","volume":"75","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8604-9523","authenticated-orcid":false,"given":"Christian","family":"Beyer","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vishnu","family":"Unnikrishnan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert","family":"Br\u00fcggemann","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vincent","family":"Toulouse","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hafez Kader","family":"Omar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eirini","family":"Ntoutsi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Myra","family":"Spiliopoulou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,8,29]]},"reference":[{"key":"800_CR1","doi-asserted-by":"crossref","unstructured":"Spitz A, Almasian S, Gertz M (2019) Topexnet: entity-centric network topic exploration in news streams. In: Proceedings of the twelfth ACM international conference on web search and data mining, pp 798\u2013801","DOI":"10.1145\/3289600.3290619"},{"key":"800_CR2","doi-asserted-by":"crossref","unstructured":"Spitz A, Gertz M (2018) Exploring entity-centric networks in entangled news streams. In: Companion proceedings of the the web conference 2018, pp 555\u2013563","DOI":"10.1145\/3184558.3188726"},{"key":"800_CR3","doi-asserted-by":"crossref","unstructured":"Bahri M, Pfahringer B, Bifet A, Maniu S (2020) Efficient batch-incremental classification using umap for evolving data streams. In: Advances in intelligent data analysis XVIII. Springer International Publishing, pp 40\u201353","DOI":"10.1007\/978-3-030-44584-3_4"},{"key":"800_CR4","doi-asserted-by":"publisher","first-page":"278","DOI":"10.1016\/j.jss.2016.07.005","volume":"127","author":"JP Barddal","year":"2017","unstructured":"Barddal JP, Gomes HM, Enembreck F, Pfahringer B (2017) A survey on feature drift adaptation: definition, benchmark, challenges and future directions. J Syst Softw 127:278\u2013294","journal-title":"J Syst Softw"},{"key":"800_CR5","doi-asserted-by":"crossref","unstructured":"Beyer C, Niemann U, Unnikrishnan V, Ntoutsi E, Spiliopoulou M (2018) Predicting polarities of entity-centered documents without reading their contents. In: Proceedings of the 33rd annual ACM symposium on applied computing. ACM, pp 525\u2013528","DOI":"10.1145\/3167132.3172870"},{"key":"800_CR6","doi-asserted-by":"crossref","unstructured":"Beyer C, Unnikrishnan V, Niemann U, Matuszyk P, Ntoutsi E, Spiliopoulou M (2019) Exploiting entity information for stream classification over a stream of reviews. In: Proceedings of the 34th ACM\/SIGAPP symposium on applied computing. ACM, pp 564\u2013573","DOI":"10.1145\/3297280.3297333"},{"issue":"9","key":"800_CR7","doi-asserted-by":"publisher","first-page":"1166","DOI":"10.1109\/TKDE.2006.137","volume":"18","author":"B-R Dai","year":"2006","unstructured":"Dai B-R, Huang J-W, Yeh M-Y, Chen M-S (2006) Adaptive clustering for multiple evolving streams. IEEE Trans Knowl Data Eng 18(9):1166\u20131180","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"800_CR8","doi-asserted-by":"crossref","unstructured":"Bifet A, Read J, \u017eliobait\u0117 I, Pfahringer B, Holmes G (2013) Pitfalls in benchmarking data stream classification and how to avoid them. In: Joint european conference on machine learning and knowledge discovery in databases. Springer, pp 465\u2013479","DOI":"10.1007\/978-3-642-40988-2_30"},{"key":"800_CR9","unstructured":"Erdogan BE, Aky\u00fcz S\u00d6, Atas PK (2019) A novel approach for panel data: an ensemble of weighted functional margin svm models. Information Sciences"},{"key":"800_CR10","doi-asserted-by":"crossref","unstructured":"He R, McAuley J (2016) Ups and downs: modeling the visual evolution of fashion trends with one-class collaborative filtering. In: Proceedings of the 25th international conference on world wide web, pp 507\u2013517","DOI":"10.1145\/2872427.2883037"},{"key":"800_CR11","doi-asserted-by":"crossref","unstructured":"Liao J, Dai B (2014) An ensemble learning approach for concept drift. In: 2014 international conference on information science applications (ICISA), pp 1\u20134","DOI":"10.1109\/ICISA.2014.6847357"},{"key":"800_CR12","doi-asserted-by":"crossref","unstructured":"Liu Z, Hauskrecht M (2016) Learning adaptive forecasting models from irregularly sampled multivariate clinical data. In: Thirtieth AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v30i1.10181"},{"key":"800_CR13","unstructured":"Lu H, Huang S (2011) Clustering panel data. In: SIAM international workshop on data mining held in conjunction with the 2011 SIAM international conference on data mining, pp 1\u201310"},{"key":"800_CR14","doi-asserted-by":"crossref","unstructured":"Manku GS, Motwani R (2002) Approximate frequency counts over data streams. In: VLDB\u201902: proceedings of the 28th international conference on very large databases. Elsevier, pp 346\u2013357","DOI":"10.1016\/B978-155860869-6\/50038-X"},{"key":"800_CR15","doi-asserted-by":"crossref","unstructured":"Melidis DP, Spiliopoulou M, Ntoutsi E (2018) Learning under feature drifts in textual streams. In: Proceedings of the 27th ACM international conference on information and knowledge management. ACM, pp 527\u2013536","DOI":"10.1145\/3269206.3271717"},{"key":"800_CR16","unstructured":"Rudovic O, Utsumi Y, Guerrero R, Peterson K, Rueckert D, Picard RW (2019) Meta-weighted gaussian process experts for personalized forecasting of ad cognitive changes. In: Machine learning for healthcare conference, pp 181\u2013196"},{"key":"800_CR17","doi-asserted-by":"crossref","unstructured":"Saadallah A, Priebe F, Morik K (2020) A drift-based dynamic ensemble members selection using clustering for time series forecasting. In: Brefeld U, Fromont E, Hotho A, Knobbe A, Maathuis M, Robardet C (eds) Machine learning and knowledge discovery in databases. Springer International Publishing, Cham, pp 678\u2013694","DOI":"10.1007\/978-3-030-46150-8_40"},{"key":"800_CR18","doi-asserted-by":"crossref","unstructured":"Unnikrishnan V, Beyer C, Matuszyk P, Niemann U, Pryss R, Schlee W, Ntoutsi E, Spiliopoulou M (2019) Entity-level stream classification: exploiting entity similarity to label the future observations referring to an entity. International Journal of Data Science and Analytics","DOI":"10.1109\/DSAA.2018.00035"},{"key":"800_CR19","doi-asserted-by":"crossref","unstructured":"Wagner S, Zimmermann M, Ntoutsi E, Spiliopoulou M (2015) Ageing-based multinomial naive bayes classifiers over opinionated data streams. In: European conference on machine learning and principles and practice of knowledge discovery in databases, ECMLPKDD\u201915, vol 9284, pp 401\u2013416","DOI":"10.1007\/978-3-319-23528-8_25"}],"container-title":["Annals of Telecommunications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12243-020-00800-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12243-020-00800-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12243-020-00800-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,11]],"date-time":"2022-11-11T05:44:18Z","timestamp":1668145458000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12243-020-00800-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,29]]},"references-count":19,"journal-issue":{"issue":"9-10","published-print":{"date-parts":[[2020,10]]}},"alternative-id":["800"],"URL":"https:\/\/doi.org\/10.1007\/s12243-020-00800-4","relation":{},"ISSN":["0003-4347","1958-9395"],"issn-type":[{"type":"print","value":"0003-4347"},{"type":"electronic","value":"1958-9395"}],"subject":[],"published":{"date-parts":[[2020,8,29]]},"assertion":[{"value":"7 November 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 August 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 August 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}