{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T00:29:42Z","timestamp":1777854582017,"version":"3.51.4"},"reference-count":54,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2023,9,19]],"date-time":"2023-09-19T00:00:00Z","timestamp":1695081600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100010248","name":"zhejiang province public welfare technology application research project","doi-asserted-by":"publisher","award":["LGF18F020019"],"award-info":[{"award-number":["LGF18F020019"]}],"id":[{"id":"10.13039\/501100010248","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100022963","name":"key research and development program of zhejiang province","doi-asserted-by":"publisher","award":["2022C01220"],"award-info":[{"award-number":["2022C01220"]}],"id":[{"id":"10.13039\/100022963","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Information Science"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:p>Author name disambiguation (AND) is the task of resolving the ambiguity problem in bibliographic databases, where distinct real-world authors may share the same name or same author may have distinct names. The aim of AND is to split the name-ambiguous entities (articles) into the corresponding authors. Existing AND algorithms mainly focus on designing different similarity metrics between two ambiguous articles. However, most previous methods empirically select and process the features of entities, then use features to predict the similarity by data-driven models. In this article, we are motivated by natural questions: Which features are most useful for splitting name-ambiguous entities? Can they be automatically determined by an optimisation approach rather than heuristic feature engineering? Therefore, we proposed a novel end-to-end differentiable feature selection algorithm, automatically searching the optimal features for AND task (AAND). AAND optimises the discrete feature selection by differentiable Gumbel-Softmax, leading to the joint learning of feature selection policy and similarity prediction model. The experiments are conducted on a benchmark data set, S2AND, which harmonises eight different AND data sets. The results show that the performance of our proposal is superior to the advanced AND methods and feature selection algorithms. Meanwhile, deep insights into AND features are also given.<\/jats:p>","DOI":"10.1177\/01655515231193859","type":"journal-article","created":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T02:34:50Z","timestamp":1695177290000},"page":"309-323","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Automatic author name disambiguation by differentiable feature selection"],"prefix":"10.1177","volume":"52","author":[{"given":"ZhiJian","family":"Fang","sequence":"first","affiliation":[{"name":"Zhejiang Sci-Tech University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Zhuo","sequence":"additional","affiliation":[{"name":"Zhejiang Sci-Tech University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinying","family":"Xu","sequence":"additional","affiliation":[{"name":"Zhejiang Province Science and Technology Department, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5937-1488","authenticated-orcid":false,"given":"Zhechong","family":"Tang","sequence":"additional","affiliation":[{"name":"Zhejiang Sci-Tech University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zijie","family":"Jia","sequence":"additional","affiliation":[{"name":"Zhejiang Sci-Tech University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"HuaXiong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Zhejiang Sci-Tech University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2023,9,19]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551519888605"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1002\/aris.2009.1440430113"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0269888917000182"},{"issue":"8","key":"e_1_3_3_5_2","first-page":"15","article-title":"Author name disambiguation techniques for academic literature: a review","volume":"4","author":"Zhe S","year":"2020","unstructured":"Zhe S, Yi W, Yifan Y, et al. Author name disambiguation techniques for academic literature: a review. Data Anal Knowl Disc 2020; 4(8): 15\u201327.","journal-title":"Data Anal Knowl Disc"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11192-014-1289-4"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.3390\/e22040416"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551518761011"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13278-015-0249-1"},{"key":"e_1_3_3_10_2","doi-asserted-by":"crossref","unstructured":"Navarro G. A Guided Tour to Approximate String Matching. ACM Comput Surv 2001 33: 31\u201388. http:\/\/dx.doi.org\/10.1145\/375360.375365","DOI":"10.1145\/375360.375365"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/1555400.1555408"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICWS.2018.00041"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/11575832_13"},{"key":"e_1_3_3_14_2","first-page":"29","volume-title":"Proceedings of the 1st instructional conference on machine learning","volume":"242","author":"Ramos J","unstructured":"Ramos J. Using TF-IDF to determine word relevance in document queries. In: Proceedings of the 1st instructional conference on machine learning, December, vol. 242. pp. 29\u201348."},{"key":"e_1_3_3_15_2","volume-title":"Proceedings of the 26th international conference on neural information processing systems","author":"Mikolov T","unstructured":"Mikolov T, Sutskever I, Chen K, et al. Distributed representations of words and phrases and their compositionality. In: Proceedings of the 26th international conference on neural information processing systems, Lake Tahoe, NV, 5\u201310 December 2013."},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Cohan A Feldman S Beltagy I et al. SPECTER: document-level representation learning using citation-informed transformers 2020 https:\/\/arxiv.org\/abs\/2004.07180","DOI":"10.18653\/v1\/2020.acl-main.207"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1590\/S1415-47571999000300024"},{"key":"e_1_3_3_18_2","first-page":"380","volume-title":"Proceedings of the international multiconference of engineers and computer scientists","volume":"1","author":"Niwattanakul S","unstructured":"Niwattanakul S, Singthongchai J, Naenudorn E, et al. Using of Jaccard coefficient for keywords similarity. In: Proceedings of the international multiconference of engineers and computer scientists, Hong Kong, 13\u201315 March 2013, vol. 1, pp. 380\u2013384. IAENG"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557107"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.23063"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106622"},{"key":"e_1_3_3_22_2","unstructured":"Jang E Gu S Poole B. Categorical reparameterization with Gumbel-Softmax 2016 https:\/\/arxiv.org\/abs\/1611.01144"},{"key":"e_1_3_3_23_2","unstructured":"Liu H Simonyan K Yang Y. DARTS: differentiable architecture search 2018 https:\/\/arxiv.org\/abs\/1806.09055"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Li Y Hu G Wang Y et al. DADA: differentiable automatic data augmentation 2020 https:\/\/arxiv.org\/abs\/2003.03780","DOI":"10.1007\/978-3-030-58542-6_35"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-44874-8_3"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2015.2494218"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.301"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-25832-9_16"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2014.03.013"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-67008-9_24"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11192-017-2341-y"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.24212"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDMW.2019.00150"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13369-018-3099-0"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMPSAC.2018.10226"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-55705-2_13"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.joi.2015.08.004"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/1891879.1891883"},{"key":"e_1_3_3_39_2","first-page":"444","volume-title":"Proceedings of the 36th international conference on machine learning","volume":"97","author":"Baln MF","unstructured":"Baln MF, Abid A, Zou J. Concrete autoencoders: differentiable feature selection and reconstruction. In: Proceedings of the 36th international conference on machine learning (ed Chaudhuri K, Salakhutdinov R), Long Beach, CA, 9\u201315 June 2019, vol. 97, pp. 444\u2013453. Proceedings of Machine Learning Research (PMLR). Cambridge, MA: JMLR"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1175\/1520-0434(1996)011<0003:TFAASE>2.0.CO;2"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.2307\/1932409"},{"key":"e_1_3_3_42_2","first-page":"707","article-title":"Binary codes capable of correcting deletions, insertions and reversals","volume":"10","author":"Levenshtein VI","year":"1966","unstructured":"Levenshtein VI. Binary codes capable of correcting deletions, insertions and reversals. Sov Phys: Doklady 1966; 10: 707\u2013710.","journal-title":"Sov Phys: Doklady"},{"key":"e_1_3_3_43_2","first-page":"957","volume-title":"Proceedings of the 32nd international conference on machine learning","author":"Kusner MJ","unstructured":"Kusner MJ, Sun Y, Kolkin NI, et al. From word embeddings to document distances. In: Proceedings of the 32nd international conference on machine learning, Lille, 6\u201311 July 2015, pp. 957\u2013966. Proceedings of Machine Learning Research (PMLR). Cambridge, MA: JMLR"},{"key":"e_1_3_3_44_2","unstructured":"Mikolov T Chen K Corrado G et al. Efficient estimation of word representations in vector space 2013 https:\/\/arxiv.org\/abs\/1301.3781"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00357-014-9161-z"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219859"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.13"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-45880-9_21"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2010.10.001"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1093\/jamia\/ocz028"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-015-9261-3"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11192-017-2363-5"},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"M\u00fcllner D. Fastcluster: fast hierarchical agglomerative clustering routines for R and Python. J Stat Softw 2013; 53(i09) https:\/\/ideas.repec.org\/a\/jss\/jstsof\/v053i09.html","DOI":"10.18637\/jss.v053.i09"},{"key":"e_1_3_3_54_2","unstructured":"Kim J. A fast and integrative algorithm for clustering performance evaluation in author name disambiguation 2021 https:\/\/arxiv.org\/abs\/2102.03251"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/949\/1\/012009"}],"container-title":["Journal of Information Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/01655515231193859","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/01655515231193859","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/01655515231193859","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T23:10:14Z","timestamp":1777504214000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/01655515231193859"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,19]]},"references-count":54,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["10.1177\/01655515231193859"],"URL":"https:\/\/doi.org\/10.1177\/01655515231193859","relation":{},"ISSN":["0165-5515","1741-6485"],"issn-type":[{"value":"0165-5515","type":"print"},{"value":"1741-6485","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,19]]}}}