{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T15:35:06Z","timestamp":1772120106981,"version":"3.50.1"},"reference-count":35,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T00:00:00Z","timestamp":1690502400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T00:00:00Z","timestamp":1690502400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Data Sci Anal"],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Certain categorical sequence clustering applications require path connectivity, such as the clustering of DNA, click-paths through web-user sessions, or paths of care clustering with sequences of patient medical billing codes.\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$K$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mi>K<\/mml:mi>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    -means and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$k$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mi>k<\/mml:mi>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    -medoids clustering with non-Euclidean distance metrics such as the Jaccard or edit distances maintains such path connectivity. Although\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$k$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mi>k<\/mml:mi>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    -means and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$k$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mi>k<\/mml:mi>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    -medoids clustering with the Jaccard and edit distances have enjoyed success in these domains, the limits of accurate cluster recovery in these conditions have not yet been defined. As a first step in approaching this goal, we performed a simulated study using\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$k$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mi>k<\/mml:mi>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    -means and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$k$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mi>k<\/mml:mi>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    -medoids clustering with non-Euclidean distances and show the performance deteriorates at a certain level of noise and when the number of clusters increases. However, we identify initialization strategies that improve upon cluster recovery in the presence of noise. We employ the use of the Tibshirani and Guenther (J Comput Graph Stat 14(3):511\u2013528, 2005) Prediction Strength method, which creates a hypothesis testing scenario that determines if there is clustering structure to the data (if the clusters are reproducible), with the null hypothesis being there is none. We then applied the framework to perinatal episodes of care and the clusters reproducibly and organically split between Cesarean and vaginal deliveries, which itself is not a clinical finding but sensibly validates the approach. Further visualizations of the clusters did bring insights into subclusters that split along groups of physicians, cost and risk scores, warranting the outlined future work into ways of improving this framework for better resolution.\n                  <\/jats:p>","DOI":"10.1007\/s41060-023-00429-1","type":"journal-article","created":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T12:02:01Z","timestamp":1690545721000},"page":"87-106","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Reproducible clustering with non-Euclidean distances: a simulation and case study"],"prefix":"10.1007","volume":"20","author":[{"given":"Lauren","family":"Staples","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Janelle","family":"Ring","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Scott","family":"Fontana","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christina","family":"Stradwick","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joe","family":"DeMaio","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Herman","family":"Ray","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yifan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,7,28]]},"reference":[{"issue":"3","key":"429_CR1","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1109\/TPAMI.2005.56","volume":"27","author":"A Robles-Kelly","year":"2005","unstructured":"Robles-Kelly, A., Hancock, E.R.: Graph edit distance from spectral seriation. IEEE Trans. Pattern Anal. Mach. Intell. 27(3), 365\u2013378 (2005)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"3","key":"429_CR2","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1023\/A:1009769707641","volume":"2","author":"Z Huang","year":"1998","unstructured":"Huang, Z.: Extensions to the k-means algorithm for clustering large data sets with categorical values. Data Min. Knowl. Disc. 2(3), 283\u2013304 (1998)","journal-title":"Data Min. Knowl. Disc."},{"issue":"2","key":"429_CR3","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1016\/j.datak.2007.03.016","volume":"63","author":"A Ahmad","year":"2007","unstructured":"Ahmad, A., Dey, L.: A k-mean clustering algorithm for mixed numeric and categorical data. Data Knowl. Eng. 63(2), 503\u2013527 (2007)","journal-title":"Data Knowl. Eng."},{"issue":"2","key":"429_CR4","doi-asserted-by":"publisher","first-page":"466","DOI":"10.1007\/s10618-014-0357-y","volume":"29","author":"M Garc\u00eda-Magari\u00f1os","year":"2015","unstructured":"Garc\u00eda-Magari\u00f1os, M., Vilar, J.A.: A framework for dissimilarity-based partitioning clustering of categorical time series. Data Min. Knowl. Disc. 29(2), 466\u2013502 (2015)","journal-title":"Data Min. Knowl. Disc."},{"issue":"8","key":"429_CR5","doi-asserted-by":"publisher","first-page":"2228","DOI":"10.1016\/j.patcog.2013.01.027","volume":"46","author":"Y-M Cheung","year":"2013","unstructured":"Cheung, Y.-M., Jia, H.: Categorical-and-numerical-attribute data clustering based on a unified similarity metric without knowing cluster number. Pattern Recogn. 46(8), 2228\u20132238 (2013)","journal-title":"Pattern Recogn."},{"issue":"5","key":"429_CR6","doi-asserted-by":"publisher","first-page":"1065","DOI":"10.1109\/TNNLS.2015.2436432","volume":"27","author":"H Jia","year":"2015","unstructured":"Jia, H., Cheung, Y.-M., Liu, J.: A new distance metric for unsupervised learning of categorical data. IEEE Trans. Neural Netw. Learn. Syst. 27(5), 1065\u20131079 (2015)","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"429_CR7","unstructured":"Chen, L., Wang, S.: Central clustering of categorical data with automated feature weighting. In: IJCAI, pp. 1260\u20131266 (2013)"},{"key":"429_CR8","doi-asserted-by":"crossref","unstructured":"Tierney, S., Gao, J., Guo, Y.: Subspace clustering for sequential data. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1019\u20131026 (2014)","DOI":"10.1109\/CVPR.2014.134"},{"issue":"12","key":"429_CR9","doi-asserted-by":"publisher","first-page":"2936","DOI":"10.1109\/TNNLS.2016.2608354","volume":"28","author":"G Guo","year":"2016","unstructured":"Guo, G., Chen, L., Ye, Y., Jiang, Q.: Cluster validation method for determining the number of clusters in categorical sequences. IEEE Trans. Neural Netw. Learn. Syst. 28(12), 2936\u20132948 (2016)","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"429_CR10","unstructured":"Kim, S.M., Pena, M.I., Moll, M., Giannakopoulos, G., Bennett, G.N., Kavraki, L.E., Demokritos, I.N.: An evaluation of different clustering methods and distance measures used for grouping metabolic pathways. In: 2016 International Conference on Bioinformatics and Computational Biology. ISCA, pp. 115\u2013122 (2016)"},{"key":"429_CR11","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2020.103668","volume":"115","author":"E Aspland","year":"2021","unstructured":"Aspland, E., Harper, P.R., Gartner, D., Webb, P., Barrett-Lee, P.: Modified Needleman\u2013Wunsch algorithm for clinical pathway clustering. J. Biomed. Inform. 115, 103668 (2021)","journal-title":"J. Biomed. Inform."},{"key":"429_CR12","unstructured":"Zhang, Y.: Model-based clustering of sequential and directional data. PhD thesis, The University of Alabama (2020)"},{"key":"429_CR13","volume-title":"Introduction to Information Retrieval","author":"H Sch\u00fctze","year":"2008","unstructured":"Sch\u00fctze, H., Manning, C.D., Raghavan, P.: Introduction to Information Retrieval, vol. 39. Cambridge University Press, Cambridge (2008)"},{"key":"429_CR14","doi-asserted-by":"crossref","unstructured":"Yang, J., Wang, W.: Cluseq: efficient and effective sequence clustering. In: Proceedings 19th International Conference on Data Engineering (Cat. No. 03CH37405), pp. 101\u2013112 (2003). IEEE","DOI":"10.1109\/ICDE.2003.1260785"},{"issue":"1","key":"429_CR15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41598-021-83340-8","volume":"11","author":"G Preud\u2019homme","year":"2021","unstructured":"Preud\u2019homme, G., Duarte, K., Dalleau, K., Lacomblez, C., Bresso, E., Sma\u00efl-Tabbone, M., Couceiro, M., Devignes, M.-D., Kobayashi, M., Huttin, O., et al.: Head-to-head comparison of clustering methods for heterogeneous data: a simulation-driven benchmark. Sci. Rep. 11(1), 1\u201314 (2021)","journal-title":"Sci. Rep."},{"key":"429_CR16","doi-asserted-by":"crossref","unstructured":"Bobroske, K., Larish, C., Cattrell, A., Bjarnad\u00f3ttir, M.V., Huan, L.: The bird\u2019s-eye view: A data-driven approach to understanding patient journeys from claims data. Journal of the American Medical Informatics Association (2020)","DOI":"10.1093\/jamia\/ocaa052"},{"issue":"2","key":"429_CR17","first-page":"1","volume":"4","author":"H-H Bock","year":"2008","unstructured":"Bock, H.-H.: Origins and extensions of the k-means algorithm in cluster analysis. Electron. J. Hist. Probab. Stat. 4(2), 1\u201318 (2008)","journal-title":"Electron. J. Hist. Probab. Stat."},{"issue":"3\u20134","key":"429_CR18","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1080\/03461238.1950.10432042","volume":"1950","author":"T Dalenius","year":"1950","unstructured":"Dalenius, T.: The problem of optimum stratification. Scand. Actuar. J. 1950(3\u20134), 203\u2013213 (1950)","journal-title":"Scand. Actuar. J."},{"key":"429_CR19","unstructured":"Kaufman, P.J., Rdusseeun, L.: Clustering by means of medoids (1987)"},{"key":"429_CR20","unstructured":"Arthur, D., Vassilvitskii, S.: k-means++: The advantages of careful seeding. Technical report, Stanford (2006)"},{"key":"429_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2021\/4862451","volume":"2021","author":"G Du","year":"2021","unstructured":"Du, G., Li, X., Zhang, L., Liu, L., Zhao, C.: Novel automated k-means++ algorithm for financial data sets. Math. Prob. Eng. 2021, 1\u201312 (2021)","journal-title":"Math. Prob. Eng."},{"key":"429_CR22","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1016\/0377-0427(87)90125-7","volume":"20","author":"PJ Rousseeuw","year":"1987","unstructured":"Rousseeuw, P.J.: Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 20, 53\u201365 (1987)","journal-title":"J. Comput. Appl. Math."},{"issue":"3","key":"429_CR23","doi-asserted-by":"publisher","first-page":"511","DOI":"10.1198\/106186005X59243","volume":"14","author":"R Tibshirani","year":"2005","unstructured":"Tibshirani, R., Walther, G.: Cluster validation by prediction strength. J. Comput. Graph. Stat. 14(3), 511\u2013528 (2005)","journal-title":"J. Comput. Graph. Stat."},{"key":"429_CR24","unstructured":"Staples, L.: Simulation Framework for Categorical Clustering Study. Github Repository. https:\/\/github.com\/laurenleesc\/Reproducible-clustering-with-non-Euclidean-distances-a-simulation-and-case-study"},{"key":"429_CR25","volume-title":"Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit","author":"S Bird","year":"2009","unstructured":"Bird, S., Klein, E., Loper, E.: Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit. O\u2019Reilly Media Inc, Sebastopol, Calif (2009)"},{"key":"429_CR26","doi-asserted-by":"publisher","unstructured":"Novikov, A.: PyClustering: Data mining library. J. Open Source Softw. 4(36), 1230 (2019). https:\/\/doi.org\/10.21105\/joss.01230836","DOI":"10.21105\/joss.01230836"},{"key":"429_CR27","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al.: Scikit-learn: machine learning in python. J. Mach. Learn. Res. 12, 2825\u20132830 (2011)","journal-title":"J. Mach. Learn. Res."},{"issue":"3","key":"429_CR28","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1109\/MCSE.2007.55","volume":"9","author":"JD Hunter","year":"2007","unstructured":"Hunter, J.D.: Matplotlib: a 2d graphics environment. Comput. Sci. Eng. 9(3), 90\u201395 (2007). https:\/\/doi.org\/10.1109\/MCSE.2007.55","journal-title":"Comput. Sci. Eng."},{"key":"429_CR29","unstructured":"Technical Documents. https:\/\/www.tn.gov\/tenncare\/health-care-innovation\/episodes-of-care\/technical-documents.html. Accessed: 2021\u201305\u20133"},{"key":"429_CR30","unstructured":"Healthcare Cost and Utilization Project (HCUP): Clinical Classifications Software Refined (CCSR) for ICD-10-CM Diagnoses. Online. https:\/\/www.hcup-us.ahrq.gov\/toolssoftware\/ccs10\/ccs_dx_icd10cm_2018_1.zip. Accessed February 9, 2019. (2018)"},{"key":"429_CR31","unstructured":"Healthcare Cost and Utilization Project (HCUP): Clinical Classification Software (CCS) for ICD-9-CM. Online. https:\/\/www.hcup-us.ahrq.gov\/toolssoftware\/ccs\/Multi_Level_CCS_2015.zip. Accessed February 9, 2019 (2017)"},{"key":"429_CR32","unstructured":"Henshaw, A.: Thread plot: a matplotlib plot for a recursive data structure. Github Repository. https:\/\/github.com\/ahenshaw\/thread_plot"},{"key":"429_CR33","unstructured":"Fingar, K.R., Mabry-Hernandez, I., Ngo-Metzger, Q., Wolff, T., Steiner, C.A., Elixhauser, A.: Delivery hospitalizations involving preeclampsia and eclampsia, 2005\u20132014: statistical brief# 222 (2017)"},{"issue":"9","key":"429_CR34","doi-asserted-by":"publisher","first-page":"926","DOI":"10.1109\/34.232078","volume":"15","author":"A Marzal","year":"1993","unstructured":"Marzal, A., Vidal, E.: Computation of normalized edit distance and applications. IEEE Trans. Pattern Anal. Mach. Intell. 15(9), 926\u2013932 (1993)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"2","key":"429_CR35","doi-asserted-by":"publisher","first-page":"306","DOI":"10.1109\/TPAMI.2008.76","volume":"31","author":"P-F Marteau","year":"2008","unstructured":"Marteau, P.-F.: Time warp edit distance with stiffness adjustment for time series matching. IEEE Trans. Pattern Anal. Mach. Intell. 31(2), 306\u2013318 (2008)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["International Journal of Data Science and Analytics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41060-023-00429-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s41060-023-00429-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41060-023-00429-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,28]],"date-time":"2025-05-28T03:35:42Z","timestamp":1748403342000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s41060-023-00429-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,28]]},"references-count":35,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["429"],"URL":"https:\/\/doi.org\/10.1007\/s41060-023-00429-1","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-2739602\/v1","asserted-by":"object"}]},"ISSN":["2364-415X","2364-4168"],"issn-type":[{"value":"2364-415X","type":"print"},{"value":"2364-4168","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,28]]},"assertion":[{"value":"27 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 July 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 July 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"We have no conflicts of interest in this work. We disclose that our research was funded by BlueCross BlueShield of Tennessee through a corporate grant to Kennesaw State University, which was used to fund a research lab that supported the authors. Since the claims data is proprietary, we are not able to make the claims data publicly available. However, all code and data for the simulation study will be made publicly available through Github upon publication.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}