{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,2]],"date-time":"2026-04-02T16:10:28Z","timestamp":1775146228046,"version":"3.50.1"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2019,12,1]],"date-time":"2019-12-01T00:00:00Z","timestamp":1575158400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2019,12,3]],"date-time":"2019-12-03T00:00:00Z","timestamp":1575331200000},"content-version":"vor","delay-in-days":2,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Accenture Technology Labs, Beijing, China","award":["RDS10120180003"],"award-info":[{"award-number":["RDS10120180003"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"published-print":{"date-parts":[[2019,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Introduction<\/jats:title>\n                <jats:p>This paper presents a lifelong learning framework which constantly adapts with changing data patterns over time through incremental learning approach. In many big data systems, iterative re-training high dimensional data from scratch is computationally infeasible since constant data stream ingestion on top of a historical data pool increases the training time exponentially. Therefore, the need arises on how to retain past learning and fast update the model incrementally based on the new data. Also, the current machine learning approaches do the model prediction without providing a comprehensive root cause analysis. To resolve these limitations, our framework lays foundations on an ensemble process between stream data with historical batch data for an incremental lifelong learning (LML) model.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Case description<\/jats:title>\n                <jats:p>A cancer patient\u2019s pathological tests like blood, DNA, urine or tissue analysis provide a unique signature based on the DNA combinations. Our analysis allows personalized and targeted medications and achieves a therapeutic response. Model is evaluated through data from The National Cancer Institute\u2019s Genomic Data Commons unified data repository. The aim is to prescribe personalized medicine based on the thousands of genotype and phenotype parameters for each patient.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Discussion and evaluation<\/jats:title>\n                <jats:p>The model uses a dimension reduction method to reduce training time at an online sliding window setting. We identify the Gleason score as a determining factor for cancer possibility and substantiate our claim through Lilliefors and Kolmogorov\u2013Smirnov test. We present clustering and Random Decision Forest results. The model\u2019s prediction accuracy is compared with standard machine learning algorithms for numeric and categorical fields.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>We propose an ensemble framework of stream and batch data for incremental lifelong learning. The framework successively applies first streaming clustering technique and then Random Decision Forest Regressor\/Classifier to isolate anomalous patient data and provides reasoning through root cause analysis by feature correlations with an aim to improve the overall survival rate. While the stream clustering technique creates groups of patient profiles, RDF further drills down into each group for comparison and reasoning for useful actionable insights. The proposed MALA architecture retains the past learned knowledge and transfer to future learning and iteratively becomes more knowledgeable over time.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s40537-019-0261-9","type":"journal-article","created":{"date-parts":[[2019,12,3]],"date-time":"2019-12-03T11:04:01Z","timestamp":1575371041000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Lifelong Machine Learning and root cause analysis for large-scale cancer patient data"],"prefix":"10.1186","volume":"6","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2594-9699","authenticated-orcid":false,"given":"Gautam","family":"Pal","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xianbin","family":"Hong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhuo","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongyi","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gangmin","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Katie","family":"Atkinson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2019,12,3]]},"reference":[{"key":"261_CR1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4613-1381-6","volume-title":"Explanation-based neural network learning: a lifelong learning approach","author":"S Thrun","year":"1996","unstructured":"Thrun S. Explanation-based neural network learning: a lifelong learning approach. Boston: Kluwer Academic Publishers; 1996."},{"issue":"2","key":"261_CR2","doi-asserted-by":"publisher","first-page":"277","DOI":"10.1080\/095400996116929","volume":"8","author":"DL Silver","year":"1996","unstructured":"Silver DL. The parallel transfer of task knowledge using dynamic learning rates based on a measure of relatedness. Connect Sci. 1996;8(2):277\u201394. https:\/\/doi.org\/10.1080\/095400996116929.","journal-title":"Connect Sci"},{"key":"261_CR3","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1007\/3-540-47922-8_8","volume-title":"Advances in artificial intelligence","author":"DL Silver","year":"2002","unstructured":"Silver DL, Mercer RE. The task rehearsal method of life-long learning: overcoming impoverished data. In: Cohen R, Spencer B, editors. Advances in artificial intelligence. Berlin: Springer; 2002. p. 90\u2013101."},{"key":"261_CR4","doi-asserted-by":"publisher","first-page":"217","DOI":"10.1007\/978-3-540-24840-8_16","volume-title":"Advances in artificial intelligence","author":"DL Silver","year":"2004","unstructured":"Silver DL, Poirier R. Sequential consolidation of learned task knowledge. In: Tawfik AY, Goodwin SD, editors. Advances in artificial intelligence. Berlin: Springer; 2004. p. 217\u201332."},{"key":"261_CR5","doi-asserted-by":"publisher","first-page":"307","DOI":"10.1007\/978-3-319-18356-5_27","volume-title":"Advances in artificial intelligence","author":"DL Silver","year":"2015","unstructured":"Silver DL, Mason G, Eljabu L. Consolidation using sweep task rehearsal: overcoming the stability-plasticity problem. In: Barbosa D, Milios E, editors. Advances in artificial intelligence. Cham: Springer; 2015. p. 307\u201322."},{"key":"261_CR6","doi-asserted-by":"crossref","unstructured":"Hong X, Wong P, Liu D, Guan S-U, Man KL, Huang X. Lifelong machine learning: outlook and direction. In: Proceedings of the 2nd international conference on big data research. New York: ACM; 2018. p. 76\u201379.","DOI":"10.1145\/3291801.3291829"},{"key":"261_CR7","doi-asserted-by":"publisher","unstructured":"Hong X, Pal G, Guan S-U, Wong P, Liu D, Man KL, Huang X. Semi-unsupervised lifelong learning for sentiment classification: Less manual data annotation and more self-studying. In: Proceedings of the 2019 3rd high performance computing and cluster technologies conference. HPCCT 2019. New York: ACM; 2019. p. 87\u201392. https:\/\/doi.org\/10.1145\/3341069.3342992.","DOI":"10.1145\/3341069.3342992"},{"key":"261_CR8","doi-asserted-by":"publisher","unstructured":"Fei G, Wang S, Liu B. Learning cumulatively to become more knowledgeable. In: Proceedings of the 22Nd ACM SIGKDD international conference on knowledge discovery and data mining. KDD \u201916. New York: ACM; 2016. p. 1565\u20131574. https:\/\/doi.org\/10.1145\/2939672.2939835.","DOI":"10.1145\/2939672.2939835"},{"key":"261_CR9","unstructured":"Ruvolo P, Eaton E. ELLA: an efficient lifelong learning algorithm. In: Dasgupta S, McAllester D, editors. Proceedings of the 30th international conference on machine learning. Proceedings of machine learning research, vol. 28. Atlanta: PMLR; 2013. p. 507\u2013515. http:\/\/proceedings.mlr.press\/v28\/ruvolo13.html. Accessed 4 June 2019."},{"key":"261_CR10","unstructured":"Kumar A, Daume III, H. Learning task grouping and overlap in multi-task learning. 2012; arXiv preprint arXiv:1206.6417."},{"key":"261_CR11","unstructured":"Chen Z, Liu B. Topic modeling using topics from many domains, lifelong learning and big data. In: International conference on machine learning; 2014. p. 703\u2013711."},{"key":"261_CR12","doi-asserted-by":"crossref","unstructured":"Wang S, Chen Z, Liu B. Mining aspect-specific opinion using a holistic lifelong topic model. In: Proceedings of the 25th international conference on world wide web; 2016; International World Wide Web Conferences Steering Committee. p. 167\u2013176.","DOI":"10.1145\/2872427.2883086"},{"key":"261_CR13","first-page":"2986","volume-title":"Improving opinion aspect extraction using semantic similarity and aspect associations","author":"Q Liu","year":"2016","unstructured":"Liu Q, Liu B, Zhang Y, Kim DS, Gao Z. Improving opinion aspect extraction using semantic similarity and aspect associations. Menlo Park: AAAI; 2016. p. 2986\u201392."},{"key":"261_CR14","doi-asserted-by":"crossref","unstructured":"Carlson A, Betteridge J, Wang RC, Hruschka Jr ER, Mitchell TM. Coupled semi-supervised learning for information extraction. In: Proceedings of the third ACM international conference on web search and data mining. New York: ACM; 2010. p. 101\u2013110.","DOI":"10.1145\/1718487.1718501"},{"issue":"5","key":"261_CR15","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1145\/3191513","volume":"61","author":"T Mitchell","year":"2018","unstructured":"Mitchell T, Cohen W, Hruschka E, Talukdar P, Yang B, Betteridge J, Carlson A, Dalvi B, Gardner M, Kisiel B, et al. Never-ending learning. Commun ACM. 2018;61(5):103\u201315.","journal-title":"Commun ACM"},{"key":"261_CR16","unstructured":"Li L, Yang Q. Lifelong machine learning test. In: Proceedings of the workshop on \u201cBeyond the Turing Test\u201d of AAAI conference on artificial intelligence; 2015."},{"issue":"3\u20134","key":"261_CR17","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1007\/s41060-016-0027-9","volume":"1","author":"S Salloum","year":"2016","unstructured":"Salloum S, Dautov R, Chen X, Peng PX, Huang JZ. Big data analytics on apache spark. Int J Data Sci Anal. 2016;1(3\u20134):145\u201364.","journal-title":"Int J Data Sci Anal"},{"key":"261_CR18","doi-asserted-by":"crossref","unstructured":"Solaimani M, Iftekhar M, Khan L, Thuraisingham B, Ingram JB. Spark-based anomaly detection over multi-source vmware performance data in real-time. In: 2014 IEEE symposium on computational intelligence in cyber security (CICS). New York: IEEE; p. 1\u20138 2014.","DOI":"10.1109\/CICYBS.2014.7013369"},{"key":"261_CR19","doi-asserted-by":"crossref","unstructured":"Rettig L, Khayati M, Cudr\u00e9-Mauroux P, Pi\u00f3rkowski M. Online anomaly detection over big data streams. In: 2015 IEEE international conference on big data (Big Data). New York: IEEE; 2015. p. 1113\u20131122.","DOI":"10.1109\/BigData.2015.7363865"},{"key":"261_CR20","unstructured":"Guha S, Mishra N, Motwani R, O\u2019Callaghan L. Clustering data streams. In: 41st annual symposium On foundations of computer science, 2000. Proceedings. New York: IEEE; 2000. p. 359\u2013366."},{"issue":"9","key":"261_CR21","doi-asserted-by":"publisher","first-page":"2250","DOI":"10.1109\/TKDE.2013.184","volume":"26","author":"M Gupta","year":"2014","unstructured":"Gupta M, Gao J, Aggarwal CC, Han J. Outlier detection for temporal data: a survey. IEEE Trans Knowl Data Eng. 2014;26(9):2250\u201367.","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"261_CR22","doi-asserted-by":"publisher","first-page":"120","DOI":"10.1017\/CBO9781139565868","volume-title":"Statistical methods for recommender systems, Chap. 7","author":"DK Agarwal","year":"2016","unstructured":"Agarwal DK, Chen B-C. Statistical methods for recommender systems, Chap. 7. New York: Cambridge University Press; 2016. p. 120\u201341."},{"key":"261_CR23","doi-asserted-by":"publisher","first-page":"919","DOI":"10.1109\/TPDS.2016.2603511","volume":"28","author":"J Chen","year":"2017","unstructured":"Chen J, Li K, Tang Z, Bilal K, Yu S, Weng C, Li K. A parallel random forest algorithm for big data in a spark cloud computing environment. IEEE Trans Parallel Distrib Syst. 2017;28:919\u201333.","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"261_CR24","doi-asserted-by":"crossref","unstructured":"Pal G, Li G, Atkinson K. Big data ingestion and lifelong learning architecture. In: 2018 IEEE international conference on Big Data (Big Data). New York: IEEE; 2018. p. 5420\u20135423.","DOI":"10.1109\/BigData.2018.8621859"},{"issue":"4","key":"261_CR25","doi-asserted-by":"publisher","first-page":"58","DOI":"10.3390\/data3040058","volume":"3","author":"G Pal","year":"2018","unstructured":"Pal G, Li G, Atkinson K. Multi-agent big-data lambda architecture model for e-commerce analytics. Data. 2018;3(4):58.","journal-title":"Data"},{"key":"261_CR26","unstructured":"https:\/\/gdc.cancer.gov\/. Accessed 1 June 2019."},{"key":"261_CR27","unstructured":"https:\/\/spark.apache.org\/docs\/latest\/mllib-clustering.html. Accessed 27 Oct 2018."},{"issue":"3","key":"261_CR28","doi-asserted-by":"publisher","first-page":"515","DOI":"10.1109\/TKDE.2003.1198387","volume":"15","author":"S Guha","year":"2003","unstructured":"Guha S, Meyerson A, Mishra N, Motwani R, O\u2019Callaghan L. Clustering data streams: theory and practice. IEEE Trans Knowl Data Eng. 2003;15(3):515\u201328.","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"261_CR29","unstructured":"https:\/\/spark.apache.org\/docs\/2.2.0\/mllib-statistics.html#stratified-sampling. Accessed 22 Jan 2019."},{"key":"261_CR30","unstructured":"https:\/\/spark.apache.org\/docs\/2.2.0\/api\/java\/org\/apache\/spark\/ml\/classification\/RandomForestClassificationModel.html. Accessed 22 Jan 2019."},{"key":"261_CR31","doi-asserted-by":"publisher","DOI":"10.14569\/SpecialIssue.2014.040203","author":"B Hssina","year":"2014","unstructured":"Hssina B, Merbouha A, Ezzikouri H, Erritali M. A comparative study of decision tree id3 and c4.5. Int J Adv Comput Sci Appl. 2014;. https:\/\/doi.org\/10.14569\/SpecialIssue.2014.040203.","journal-title":"Int J Adv Comput Sci Appl"},{"issue":"2","key":"261_CR32","doi-asserted-by":"publisher","first-page":"438","DOI":"10.1109\/69.991727","volume":"14","author":"S Ruggieri","year":"2002","unstructured":"Ruggieri S. Efficient c4.5 [classification algorithm]. IEEE Trans Knowl Data Eng. 2002;14(2):438\u201344.","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"261_CR33","unstructured":"https:\/\/spark.apache.org\/docs\/2.2.0\/api\/java\/org\/apache\/spark\/ml\/evaluation\/MulticlassClassificationEvaluator.html. Accessed 22 Jan 2019."},{"key":"261_CR34","unstructured":"https:\/\/splunkbase.splunk.com\/app\/2890\/. Accessed 2 Feb 2019."},{"key":"261_CR35","unstructured":"https:\/\/splunkbase.splunk.com\/. Accessed 2 Feb 2019."},{"issue":"7172","key":"261_CR36","doi-asserted-by":"publisher","first-page":"1572","DOI":"10.1136\/bmj.317.7172.1572","volume":"317","author":"JM Bland","year":"1998","unstructured":"Bland JM, Altman DG. Survival probabilities (the kaplan-meier method). BMJ. 1998;317(7172):1572\u201380.","journal-title":"BMJ"},{"issue":"360a","key":"261_CR37","doi-asserted-by":"publisher","first-page":"854","DOI":"10.1080\/01621459.1977.10479970","volume":"72","author":"AV Peterson Jr","year":"1977","unstructured":"Peterson AV Jr. Expressing the kaplan-meier estimator as a function of empirical subsurvival functions. J Am Stat Assoc. 1977;72(360a):854\u20138.","journal-title":"J Am Stat Assoc"},{"issue":"1","key":"261_CR38","first-page":"21","volume":"2","author":"NM Razali","year":"2011","unstructured":"Razali NM, Wah YB, et al. Power comparisons of shapiro-wilk, Kolmogorov\u2013Smirnov, lilliefors and anderson-darling tests. J Stat Model Anal. 2011;2(1):21\u201333.","journal-title":"J Stat Model Anal"},{"key":"261_CR39","first-page":"540","volume-title":"Encyclopedia of measurement and statistics","author":"H Abdi","year":"2007","unstructured":"Abdi H, Molin P. Lilliefors\/van soest\u2019s test of normality. In: Salkind NJ, Rasmussen K, editors. Encyclopedia of measurement and statistics. Thousand Oaks: Sage; 2007. p. 540\u20134."},{"issue":"318","key":"261_CR40","doi-asserted-by":"publisher","first-page":"399","DOI":"10.1080\/01621459.1967.10482916","volume":"62","author":"HW Lilliefors","year":"1967","unstructured":"Lilliefors HW. On the Kolmogorov\u2013Smirnov test for normality with mean and variance unknown. J Am Stat Assoc. 1967;62(318):399\u2013402.","journal-title":"J Am Stat Assoc"},{"issue":"253","key":"261_CR41","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1080\/01621459.1951.10500769","volume":"46","author":"FJ Massey Jr","year":"1951","unstructured":"Massey FJ Jr. The Kolmogorov\u2013Smirnov test for goodness of fit. J Am Stat Assoc. 1951;46(253):68\u201378.","journal-title":"J Am Stat Assoc"},{"key":"261_CR42","doi-asserted-by":"crossref","unstructured":"Davis J, Goadrich M. The relationship between precision-recall and roc curves. In: Proceedings of the 23rd international conference on machine learning. New York: ACM; 2006. p. 233\u2013240.","DOI":"10.1145\/1143844.1143874"},{"issue":"3","key":"261_CR43","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/BF01062525","volume":"20","author":"RD Purves","year":"1992","unstructured":"Purves RD. Optimum numerical integration methods for estimation of area-under-the-curve (auc) and area-under-the-moment-curve (aumc). J Pharm Biopharm. 1992;20(3):211\u201326.","journal-title":"J Pharm Biopharm"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-019-0261-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s40537-019-0261-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-019-0261-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,12,1]],"date-time":"2020-12-01T19:15:20Z","timestamp":1606850120000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-019-0261-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12]]},"references-count":43,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,12]]}},"alternative-id":["261"],"URL":"https:\/\/doi.org\/10.1186\/s40537-019-0261-9","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,12]]},"assertion":[{"value":"4 June 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 October 2019","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 December 2019","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"108"}}