{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:19:42Z","timestamp":1750306782726,"version":"3.41.0"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2014,1,1]],"date-time":"2014-01-01T00:00:00Z","timestamp":1388534400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100002920","name":"Research Grants Council, University Grants Committee, Hong Kong","doi-asserted-by":"publisher","award":["GRF 617610 and T12-403\/11-2"],"award-info":[{"award-number":["GRF 617610 and T12-403\/11-2"]}],"id":[{"id":"10.13039\/501100002920","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["60970043"],"award-info":[{"award-number":["60970043"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Database Syst."],"published-print":{"date-parts":[[2014,1]]},"abstract":"<jats:p>\n            <jats:italic>Order-preserving submatrices<\/jats:italic>\n            (OPSMs) capture consensus trends over columns shared by rows in a data matrix. Mining OPSM patterns discovers important and interesting local correlations in many real applications, such as those involving biological data or sensor data. The prevalence of uncertain data in various applications, however, poses new challenges for OPSM mining, since data uncertainty must be incorporated into OPSM modeling and the algorithmic aspects.\n          <\/jats:p>\n          <jats:p>\n            In this article, we define new probabilistic matrix representations to model uncertain data with continuous distributions. A novel\n            <jats:italic>probabilistic order-preserving submatrix<\/jats:italic>\n            (POPSM) model is formalized in order to capture similar local correlations in probabilistic matrices. The POPSM model adopts a new probabilistic support measure that evaluates the extent to which a row belongs to a POPSM pattern. Due to the intrinsic high computational complexity of the POPSM mining problem, we utilize the anti-monotonic property of the probabilistic support measure and propose an efficient Apriori-based mining framework called\n            <jats:sc>ProbApri<\/jats:sc>\n            to mine POPSM patterns. The framework consists of two mining methods,\n            <jats:sc>UniApri<\/jats:sc>\n            and\n            <jats:sc>NormApri<\/jats:sc>\n            , which are developed for mining POPSM patterns, respectively, from two representative types of probabilistic matrices, the\n            <jats:italic>UniDist matrix<\/jats:italic>\n            (assuming uniform data distributions) and the\n            <jats:italic>NormDist matrix<\/jats:italic>\n            (assuming normal data distributions). We show that the\n            <jats:sc>NormApri<\/jats:sc>\n            method is practical enough for mining POPSM patterns from probabilistic matrices that model more general data distributions.\n          <\/jats:p>\n          <jats:p>We demonstrate the superiority of our approach by two applications. First, we use two biological datasets to illustrate that the POPSM model better captures the characteristics of the expression levels of biologically correlated genes and greatly promotes the discovery of patterns with high biological significance. Our result is significantly better than the counterpart OPSMRM (OPSM with repeated measurement) model which adopts a set-valued matrix representation to capture data uncertainty. Second, we run the experiments on an RFID trace dataset and show that our POPSM model is effective and efficient in capturing the common visiting subroutes among users.<\/jats:p>","DOI":"10.1145\/2533712","type":"journal-article","created":{"date-parts":[[2014,2,4]],"date-time":"2014-02-04T14:16:21Z","timestamp":1391523381000},"page":"1-43","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Mining order-preserving submatrices from probabilistic matrices"],"prefix":"10.1145","volume":"39","author":[{"given":"Qiong","family":"Fang","sequence":"first","affiliation":[{"name":"Hong Kong University of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wilfred","family":"Ng","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianlin","family":"Feng","sequence":"additional","affiliation":[{"name":"Sun Yat-Sen University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuliang","family":"Li","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2014,1,6]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557030"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/304182.304188"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/276304.276314"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/645480.655281"},{"volume-title":"An Introduction to Order Statistics. Atlantis Studies in Probability and Statistics","author":"Ahsanullah Mohammad","key":"e_1_2_1_5_1","unstructured":"Mohammad Ahsanullah , Valery Nevzorov , and Mohammad Shakil . 2013. An Introduction to Order Statistics. Atlantis Studies in Probability and Statistics , Vol. 3 , Atlantis Press . Mohammad Ahsanullah, Valery Nevzorov, and Mohammad Shakil. 2013. An Introduction to Order Statistics. Atlantis Studies in Probability and Statistics, Vol. 3, Atlantis Press."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1014052.1014111"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/565196.565203"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 2nd SIAM ICDM Workshop on Clustering High Dimensional Data. SIAM","author":"Busygin Stanislav","year":"2002","unstructured":"Stanislav Busygin , Gerrit Jacobsen , Ewald Kramer , and Contentsoft Ag . 2002 . Double conjugated clustering applied to leukemia microarray data . In Proceedings of the 2nd SIAM ICDM Workshop on Clustering High Dimensional Data. SIAM , Philadelphia, PA. Stanislav Busygin, Gerrit Jacobsen, Ewald Kramer, and Contentsoft Ag. 2002. Double conjugated clustering applied to leukemia microarray data. In Proceedings of the 2nd SIAM ICDM Workshop on Clustering High Dimensional Data. SIAM, Philadelphia, PA."},{"volume-title":"Proceedings of the 8th International Conference on Intelligent Systems for Molecular Biology. 93--103","author":"Cheng Yizong","key":"e_1_2_1_9_1","unstructured":"Yizong Cheng and George M. Church . 2000. Biclustering of expression data . In Proceedings of the 8th International Conference on Intelligent Systems for Molecular Biology. 93--103 . Yizong Cheng and George M. Church. 2000. Biclustering of expression data. In Proceedings of the 8th International Conference on Intelligent Systems for Molecular Biology. 93--103."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1504\/IJBRA.2007.011834"},{"key":"e_1_2_1_11_1","volume-title":"Krishna Murthy Karuturi","author":"Hui Chia Burton Kuan","year":"2010","unstructured":"Burton Kuan Hui Chia and R. Krishna Murthy Karuturi . 2010 . Differential co-expression framework to quantify goodness of biclusters and compare biclustering algorithms. Algor. Molecular Bio . 5, 23 (2010). Burton Kuan Hui Chia and R. Krishna Murthy Karuturi. 2010. Differential co-expression framework to quantify goodness of biclusters and compare biclustering algorithms. Algor. Molecular Bio. 5, 23 (2010)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972740.11"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2008.12"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557140"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/956750.956764"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150420"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1835804.1835861"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.180"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150529"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.244"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.210134797"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.184"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339588"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2010.03.002"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 9th International Workshop on Data Mining in Bioinformatics (BIOKDD'10)","author":"Gupta Rohit","year":"2010","unstructured":"Rohit Gupta , Navneet Rao , and Vipin Kumar . 2010 . Discovery of error-tolerant biclusters from noisy gene expression data . In Proceedings of the 9th International Workshop on Data Mining in Bioinformatics (BIOKDD'10) . ACM, New York, NY. Rohit Gupta, Navneet Rao, and Vipin Kumar. 2010. Discovery of error-tolerant biclusters from noisy gene expression data. In Proceedings of the 9th International Workshop on Data Mining in Bioinformatics (BIOKDD'10). ACM, New York, NY."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1972.10481214"},{"key":"e_1_2_1_27_1","volume-title":"Scientific Computing: An Introductory Survey","author":"Heath Michael T.","year":"2002","unstructured":"Michael T. Heath . 2002 . Scientific Computing: An Introductory Survey . McGraw-Hill Higher Education . Michael T. Heath. 2002. Scientific Computing: An Introductory Survey. McGraw-Hill Higher Education."},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Timothy R. Hughes Matthew J. Marton Allan R. Jones etal 2000. Functional discovery via a compendium of expression profiles. Cell 102 (2000) 1 109--126.  Timothy R. Hughes Matthew J. Marton Allan R. Jones et al. 2000. Functional discovery via a compendium of expression profiles. Cell 102 (2000) 1 109--126.","DOI":"10.1016\/S0092-8674(00)00015-5"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2011.10"},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Trey Ideker Vesteinn Thorsson Jeffrey A. Ranish R. Christmas J. Bunler J. Eng R. Bumgarner D. Goodlett R. Aebersold and L. Hood. 2001. Integrated genomic and proteomic analyses of a systematically perturbed metabolic network. Science 292 5518 (2001) 929--934.  Trey Ideker Vesteinn Thorsson Jeffrey A. Ranish R. Christmas J. Bunler J. Eng R. Bumgarner D. Goodlett R. Aebersold and L. Hood. 2001. Integrated genomic and proteomic analyses of a systematically perturbed metabolic network. Science 292 5518 (2001) 929--934.","DOI":"10.1126\/science.292.5518.929"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1038\/ng941"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339586"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972740.23"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1497577.1497578"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972733.36"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.97.18.9834"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/gkp491"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920923"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/951949.952138"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1081870.1081949"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2008.34"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2004.2"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1401956"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the Pacific Symposium on Biocomputing. 77--88","author":"Murali T. M.","year":"2003","unstructured":"T. M. Murali and S Kasif . 2003 . Extracting conserved gene expression motifs from gene expression data . In Proceedings of the Pacific Symposium on Biocomputing. 77--88 . T. M. Murali and S Kasif. 2003. Extracting conserved gene expression motifs from gene expression data. In Proceedings of the Pacific Symposium on Biocomputing. 77--88."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/2022850.2022868"},{"key":"e_1_2_1_46_1","volume-title":"Importance of replication in analyzing time-series gene expression data: Corticosteroid dynamics and circadian patterns in rat liver. BMC Bioinform. 11, 279","author":"Nguyen Tung T.","year":"2010","unstructured":"Tung T. Nguyen , Richard R. Almon , Debra C. DuBois , William J Jusko , and Ioannis P Androulakis . 2010. Importance of replication in analyzing time-series gene expression data: Corticosteroid dynamics and circadian patterns in rat liver. BMC Bioinform. 11, 279 ( 2010 ). Tung T. Nguyen, Richard R. Almon, Debra C. DuBois, William J Jusko, and Ioannis P Androulakis. 2010. Importance of replication in analyzing time-series gene expression data: Corticosteroid dynamics and circadian patterns in rat liver. BMC Bioinform. 11, 279 (2010)."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376637"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557095"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007730.1007731"},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the IEEE 17th International Conference on Data Engineering (ICDE'01)","author":"Pei Jian","year":"2001","unstructured":"Jian Pei , Jiawei Han , Behzad Mortazavi-Asl , Helen Pinto , Q. Chen , U. Dayal , and M.-C. Hsu . 2001 . Prefixspan: Mining sequential patterns efficiently by prefix-projected pattern growth . In Proceedings of the IEEE 17th International Conference on Data Engineering (ICDE'01) . IEEE Computer Society, Los Alamitos, CA, 215--224. Jian Pei, Jiawei Han, Behzad Mortazavi-Asl, Helen Pinto, Q. Chen, U. Dayal, and M.-C. Hsu. 2001. Prefixspan: Mining sequential patterns efficiently by prefix-projected pattern growth. In Proceedings of the IEEE 17th International Conference on Data Engineering (ICDE'01). IEEE Computer Society, Los Alamitos, CA, 215--224."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2004.77"},{"key":"e_1_2_1_52_1","first-page":"293","article-title":"Improved biclustering on expression data through overlapping control","volume":"3","author":"Pontes Beatriz","year":"2010","unstructured":"Beatriz Pontes , Federico Divina , Ra\u00fal Gir\u00e1ldez , and J. S. Aguilar-Ruiz . 2010 . Improved biclustering on expression data through overlapping control . Int. J. Intell. Comput. Cybernet. 3 (2010), 293 -- 309 . Beatriz Pontes, Federico Divina, Ra\u00fal Gir\u00e1ldez, and J. S. Aguilar-Ruiz. 2010. Improved biclustering on expression data through overlapping control. Int. J. Intell. Comput. Cybernet. 3 (2010), 293--309.","journal-title":"Int. J. Intell. Comput. Cybernet."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btl060"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.148"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376688"},{"volume-title":"Introduction to DNA Microarrays","author":"Seidel Chris","key":"e_1_2_1_56_1","unstructured":"Chris Seidel . 2008. Introduction to DNA Microarrays . Wiley-VCH Verlag GmbH & Co. KGaA , 1--26. Chris Seidel. 2008. Introduction to DNA Microarrays. Wiley-VCH Verlag GmbH & Co. KGaA, 1--26."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2009.102"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/645337.650382"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/18.suppl_1.S136"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1287\/ijoc.1090.0358"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.5555\/2117684.2118202"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/1378600.1378631"},{"key":"e_1_2_1_63_1","volume-title":"Clustering gene-expression data with repeated measurements. Gen. Biol. 4, 5","author":"Yeung Ka Yee","year":"2003","unstructured":"Ka Yee Yeung , Mario Medvedovic , and Roger Bumgarner . 2003. Clustering gene-expression data with repeated measurements. Gen. Biol. 4, 5 ( 2003 ). Ka Yee Yeung, Mario Medvedovic, and Roger Bumgarner. 2003. Clustering gene-expression data with repeated measurements. Gen. Biol. 4, 5 (2003)."},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.167"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2008.4497424"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2247596.2247606"}],"container-title":["ACM Transactions on Database Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2533712","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2533712","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:34:05Z","timestamp":1750232045000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2533712"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,1]]},"references-count":66,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,1]]}},"alternative-id":["10.1145\/2533712"],"URL":"https:\/\/doi.org\/10.1145\/2533712","relation":{},"ISSN":["0362-5915","1557-4644"],"issn-type":[{"type":"print","value":"0362-5915"},{"type":"electronic","value":"1557-4644"}],"subject":[],"published":{"date-parts":[[2014,1]]},"assertion":[{"value":"2012-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-01-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}