{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T13:00:34Z","timestamp":1787317234138,"version":"build-2736575974"},"publisher-location":"Cham","reference-count":19,"publisher":"Springer International Publishing","isbn-type":[{"value":"9783030461492","type":"print"},{"value":"9783030461508","type":"electronic"}],"license":[{"start":{"date-parts":[[2020,1,1]],"date-time":"2020-01-01T00:00:00Z","timestamp":1577836800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,4,30]],"date-time":"2020-04-30T00:00:00Z","timestamp":1588204800000},"content-version":"vor","delay-in-days":120,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    T-distributed stochastic neighbour embedding (t-SNE) is a widely used data visualisation technique. It differs from its predecessor SNE by the low-dimensional similarity kernel: the Gaussian kernel was replaced by the heavy-tailed Cauchy kernel, solving the \u2018crowding problem\u2019 of SNE. Here, we develop an efficient implementation of t-SNE for a t-distribution kernel with an arbitrary degree of freedom\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\nu $$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    , with\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\nu \\rightarrow \\infty $$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    corresponding to SNE and\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\nu =1$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    corresponding to the standard t-SNE. Using theoretical analysis and toy examples, we show that\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\nu &lt;1$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    can further reduce the crowding problem and reveal finer cluster structure that is invisible in standard t-SNE. We further demonstrate the striking effect of heavier-tailed kernels on large real-life data sets such as MNIST, single-cell RNA-sequencing data, and the HathiTrust library. We use domain knowledge to confirm that the revealed clusters are meaningful. Overall, we argue that modifying the tail heaviness of the t-SNE kernel can yield additional insight into the cluster structure of the data.\n                  <\/jats:p>","DOI":"10.1007\/978-3-030-46150-8_8","type":"book-chapter","created":{"date-parts":[[2020,4,30]],"date-time":"2020-04-30T21:02:40Z","timestamp":1588280560000},"page":"124-139","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Heavy-Tailed Kernels Reveal a Finer Cluster Structure in t-SNE Visualisations"],"prefix":"10.1007","author":[{"given":"Dmitry","family":"Kobak","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"George","family":"Linderman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stefan","family":"Steinerberger","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuval","family":"Kluger","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Philipp","family":"Berens","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,4,30]]},"reference":[{"issue":"6","key":"8_CR1","doi-asserted-by":"publisher","first-page":"545","DOI":"10.1038\/nbt.2594","volume":"31","author":"EAD Amir","year":"2013","unstructured":"Amir, E.A.D., et al.: viSNE enables visualization of high dimensional single-cell data and reveals phenotypic heterogeneity of leukemia. Nat. Biotechnol. 31(6), 545 (2013)","journal-title":"Nat. Biotechnol."},{"key":"8_CR2","doi-asserted-by":"crossref","unstructured":"Belkina, A.C., Ciccolella, C.O., Anno, R., Spidlen, J., Halpert, R., Snyder-Cappione, J.: Automated optimized parameters for T-distributed stochastic neighbor embedding improve visualization and analysis of large datasets. Nat. Commun. 10, 5415 (2019)","DOI":"10.1038\/s41467-019-13055-y"},{"key":"8_CR3","unstructured":"Bernhardsson, E.: Annoy. https:\/\/github.com\/spotify\/annoy (2013)"},{"key":"8_CR4","doi-asserted-by":"crossref","unstructured":"Diaz-Papkovich, A., Anderson-Trocme, L., Ben-Eghan, C., Gravel, S.: UMAP reveals cryptic population structure and phenotype heterogeneity in large genomic cohorts. PLoS Genet. 15(11), e1008432 (2019)","DOI":"10.1371\/journal.pgen.1008432"},{"key":"8_CR5","unstructured":"Hinton, G., Roweis, S.: Stochastic neighbor embedding. In: Advances in Neural Information Processing Systems, pp. 857\u2013864 (2003)"},{"key":"8_CR6","unstructured":"Im, D.J., Verma, N., Branson, K.: Stochastic neighbor embedding under f-divergences. arXiv (2018)"},{"key":"8_CR7","doi-asserted-by":"crossref","unstructured":"Kobak, D., Berens, P.: The art of using t-SNE for single-cell transcriptomics. Nat. Commun. 10, 5416 (2019)","DOI":"10.1038\/s41467-019-13056-x"},{"issue":"7\u20139","key":"8_CR8","doi-asserted-by":"publisher","first-page":"1431","DOI":"10.1016\/j.neucom.2008.12.017","volume":"72","author":"JA Lee","year":"2009","unstructured":"Lee, J.A., Verleysen, M.: Quality assessment of dimensionality reduction: rank-based criteria. Neurocomputing 72(7\u20139), 1431\u20131443 (2009)","journal-title":"Neurocomputing"},{"key":"8_CR9","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1038\/s41592-018-0308-4","volume":"16","author":"GC Linderman","year":"2019","unstructured":"Linderman, G.C., Rachh, M., Hoskins, J.G., Steinerberger, S., Kluger, Y.: Fast interpolation-based t-SNE for improved visualization of single-cell RNA-seq data. Nat. Methods 16, 243\u2013245 (2019)","journal-title":"Nat. Methods"},{"key":"8_CR10","unstructured":"van der Maaten, L.: Learning a parametric embedding by preserving local structure. In: International Conference on Artificial Intelligence and Statistics, pp. 384\u2013391 (2009)"},{"issue":"1","key":"8_CR11","first-page":"3221","volume":"15","author":"L van der Maaten","year":"2014","unstructured":"van der Maaten, L.: Accelerating t-SNE using tree-based algorithms. J. Mach. Learn. Res. 15(1), 3221\u20133245 (2014)","journal-title":"J. Mach. Learn. Res."},{"key":"8_CR12","first-page":"2579","volume":"9","author":"L van der Maaten","year":"2008","unstructured":"van der Maaten, L., Hinton, G.: Visualizing data using t-SNE. J. Mach. Learn. Res. 9, 2579\u20132605 (2008)","journal-title":"J. Mach. Learn. Res."},{"key":"8_CR13","doi-asserted-by":"crossref","unstructured":"McInnes, L., Healy, J., Melville, J.: UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv (2018)","DOI":"10.21105\/joss.00861"},{"key":"8_CR14","doi-asserted-by":"crossref","unstructured":"Schmidt, B.: Stable random projection: Lightweight, general-purpose dimensionality reduction for digitized libraries. J. Cult. Anal. (2018)","DOI":"10.31235\/osf.io\/36neu"},{"key":"8_CR15","doi-asserted-by":"crossref","unstructured":"Tang, J., Liu, J., Zhang, M., Mei, Q.: Visualizing large-scale and high-dimensional data. In: Proceedings of the 25th International Conference on World Wide Web, pp. 287\u2013297. International World Wide Web Conferences Steering Committee (2016)","DOI":"10.1145\/2872427.2883041"},{"issue":"7729","key":"8_CR16","doi-asserted-by":"publisher","first-page":"72","DOI":"10.1038\/s41586-018-0654-5","volume":"563","author":"B Tasic","year":"2018","unstructured":"Tasic, B., et al.: Shared and distinct transcriptomic cell types across neocortical areas. Nature 563(7729), 72 (2018)","journal-title":"Nature"},{"issue":"10","key":"8_CR17","doi-asserted-by":"publisher","first-page":"e2","DOI":"10.23915\/distill.00002","volume":"1","author":"M Wattenberg","year":"2016","unstructured":"Wattenberg, M., Vi\u00e9gas, F., Johnson, I.: How to use t-SNE effectively. Distill 1(10), e2 (2016)","journal-title":"Distill"},{"key":"8_CR18","unstructured":"Yang, Z., King, I., Xu, Z., Oja, E.: Heavy-tailed symmetric stochastic neighbor embedding. In: Advances in Neural Information Processing Systems, pp. 2169\u20132177 (2009)"},{"issue":"4","key":"8_CR19","doi-asserted-by":"publisher","first-page":"999","DOI":"10.1016\/j.cell.2018.06.021","volume":"174","author":"A Zeisel","year":"2018","unstructured":"Zeisel, A., et al.: Molecular architecture of the mouse nervous system. Cell 174(4), 999\u20131014 (2018)","journal-title":"Cell"}],"updated-by":[{"DOI":"10.1007\/978-3-030-46150-8_44","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2020,4,30]],"date-time":"2020-04-30T00:00:00Z","timestamp":1588204800000}}],"container-title":["Lecture Notes in Computer Science","Machine Learning and Knowledge Discovery in Databases"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-030-46150-8_8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,6]],"date-time":"2025-05-06T05:30:26Z","timestamp":1746509426000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-030-46150-8_8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020]]},"ISBN":["9783030461492","9783030461508"],"references-count":19,"URL":"https:\/\/doi.org\/10.1007\/978-3-030-46150-8_8","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020]]},"assertion":[{"value":"30 April 2020","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"30 April 2020","order":2,"name":"change_date","label":"Change Date","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"Correction","order":3,"name":"change_type","label":"Change Type","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"The chapter was inadvertently published without the supplementary file. The supplementary file and its ESM Hint have been added to the chapter.","order":4,"name":"change_details","label":"Change Details","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"ECML PKDD","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Joint European Conference on Machine Learning and Knowledge Discovery in Databases","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"W\u00fcrzburg","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Germany","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2019","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"16 September 2019","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"20 September 2019","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"ecml2019","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"http:\/\/ecmlpkdd2019.org\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Single-blind","order":1,"name":"type","label":"Type","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Microsoft CMT","order":2,"name":"conference_management_system","label":"Conference Management System","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"733","order":3,"name":"number_of_submissions_sent_for_review","label":"Number of Submissions Sent for Review","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"130","order":4,"name":"number_of_full_papers_accepted","label":"Number of Full Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"0","order":5,"name":"number_of_short_papers_accepted","label":"Number of Short Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"18% - The value is computed by the equation \"Number of Full Papers Accepted \/ Number of Submissions Sent for Review * 100\" and then rounded to a whole number.","order":6,"name":"acceptance_rate_of_full_papers","label":"Acceptance Rate of Full Papers","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"3.04","order":7,"name":"average_number_of_reviews_per_paper","label":"Average Number of Reviews per Paper","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"5.3","order":8,"name":"average_number_of_papers_per_reviewer","label":"Average Number of Papers per Reviewer","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Yes","order":9,"name":"external_reviewers_involved","label":"External Reviewers Involved","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"ECML PKDD Workshops Information: single-blind review, submissions: 200, full papers accepted: 70, short papers accepted: 46","order":10,"name":"additional_info_on_review_process","label":"Additional Info on Review Process","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}}]}}