{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T04:51:35Z","timestamp":1775191895682,"version":"3.50.1"},"reference-count":48,"publisher":"Emerald","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,3,19]]},"abstract":"<jats:p>We perform a large-scale analysis of third-party trackers on the World Wide Web. We extract third-party embeddings from more than 3.5 billion web pages of the CommonCrawl 2012 corpus, and aggregate those to a dataset containing more than 140 million third-party embeddings in over 41 million domains. We study this data on several levels and provide the following contributions:<\/jats:p>\n                  <jats:p>(1) Our work leverages the largest empirical web tracking dataset collected so far, and exceeds related studies by more than an order of magnitude in the number of domains and web pages analyzed. As our dataset also contains the link structure of the web, we are able to derive a ranking measure for tracker occurrences based on aggregated network centrality rather than simple domain counts. We make our extracted data and computed rankings available to the research community.<\/jats:p>\n                  <jats:p>(2) On a global level, we give a precise figure for the extent of tracking, give insights into the structural properties of the \u2018online tracking sphere' and analyse which trackers (and subsequently, which companies) are used by how many websites, leveraging our ranking measure derived from the link structure of the web.<\/jats:p>\n                  <jats:p>(3) On a country-specific level, we analyse which trackers are used by websites in different countries, and identify the countries in which websites choose significantly different trackers than in the rest of the world. In particular, the three tracking domains with the highest PageRank are all owned by Google. The only exception to this pattern are a handful of countries such as China and Russia. Our results suggest that this dominance is strongly associated with country-specific political factors such as freedom of the press.<\/jats:p>\n                  <jats:p>(4) We investigate whether the content of websites influences the choice of trackers they use, leveraging more than ninety thousand categorized domains. In particular, we analyse whether highly privacy-critical websites about health and addiction make different choices of trackers than other websites. Our findings indicate that websites with highly privacy-critical content are less likely to contain trackers (60% vs 90% for other websites), even though the majority of them still do contain trackers.<\/jats:p>","DOI":"10.1561\/106.00000014","type":"journal-article","created":{"date-parts":[[2018,3,19]],"date-time":"2018-03-19T06:50:56Z","timestamp":1521442256000},"page":"53-66","source":"Crossref","is-referenced-by-count":18,"title":["On the Ubiquity of Web Tracking: Insights from a Billion-Page Web\n                    Crawl"],"prefix":"10.1561","volume":"4","author":[{"given":"Sebastian","family":"Schelter","sequence":"first","affiliation":[{"name":"Technical University Berlin, Germany"}]},{"given":"J\u00e9r\u00f4","family":"Kunegis","sequence":"additional","affiliation":[{"name":"University of Namur, Belgium"}]}],"member":"140","published-online":{"date-parts":[[2018,3,19]]},"reference":[{"key":"2025122209033188600_ref001","first-page":"674","author":"Acar","year":"2014","journal-title":"ACM CCS"},{"key":"2025122209033188600_ref002","volume-title":"Proc. Web 2.0\n                        Security and Privacy","author":"Bau","year":"2013"},{"issue":"10","key":"2025122209033188600_ref003","doi-asserted-by":"crossref","DOI":"10.1088\/1742-5468\/2008\/10\/P10008","article-title":"Fast unfolding of communities in\n                        large networks","volume":"2008","author":"Blondel","year":"2008","journal-title":"Journal of Statistical\n                        Mechanics: Theory and Experiment"},{"issue":"1","key":"2025122209033188600_ref004","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1016\/S1389-1286(00)00083-9","article-title":"Graph structure in the\n                        Web","volume":"33","author":"Broder","year":"2000","journal-title":"Computer networks"},{"key":"2025122209033188600_ref005","unstructured":"CatchaDigital\n          . 2013.\n                        \u201cWorldwide ad spending forecast: Emerging markets,\n                        mobile provide opportunities for growth\u201d.\n                        URL: http:\/\/www.slideshare.net\/catchadigital\/emarketer-worldwideadspendingforecast."},{"key":"2025122209033188600_ref006","unstructured":"U.S. Census Bureau\n          . 2015.\n                        \u201cU.S. trade in goods by\n                    country\u201d. URL:https:\/\/www.census.gov\/foreign-trade\/balance\/country.xlsx."},{"key":"2025122209033188600_ref007","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1145\/2342549.2342552","volume-title":"2012 ACM Workshop on Online\n                        Social Networks","author":"Chaabane","year":"2012"},{"issue":"4","key":"2025122209033188600_ref008","doi-asserted-by":"crossref","first-page":"661","DOI":"10.1137\/070710111","article-title":"Powerlaw distributions in\n                        empirical data","volume":"51","author":"Clauset","year":"2009","journal-title":"SIAM"},{"issue":"1","key":"2025122209033188600_ref009","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1145\/1629175.1629198","article-title":"MapReduce: A flexible data\n                        processing tool","volume":"53","author":"Dean","year":"2010","journal-title":"Communications of the\n                        ACM"},{"key":"2025122209033188600_ref010","unstructured":"\u201cDisconnect\u201d.\n                        2016. https:\/\/disconnect.me\/."},{"key":"2025122209033188600_ref011","unstructured":"DMOZ\n          . 2016.\n                        \u201cDirectory of the Web\u201d.\n                        https:\/\/www.dmoz.org\/."},{"issue":"1","key":"2025122209033188600_ref012","first-page":"61","article-title":"Accurate methods for the statistics of\n                        surprise and coincidence","volume":"19","author":"Dunning","year":"1993","journal-title":"Comput.\n                        Linguist"},{"key":"2025122209033188600_ref013","first-page":"1","volume-title":"Privacy Enhancing\n                        Technologies","author":"Eckersley","year":"2010"},{"key":"2025122209033188600_ref014","unstructured":"The Economist Intelligence Unit\n          .\n                        2012.\n                    \u201cDemocracyindex\u201d.\n                        URL: http:\/\/www.eiu.com\/public\/topical_report.aspx?campaignid=\n                        DemocracyIndex12."},{"key":"2025122209033188600_ref015","unstructured":"Electronic Frontier Foundation\n          .\n                        2015. \u201cHealthCare.gov sends personal data\n                        to dozens of tracking websites\u201d.\n                        URL: https:\/\/www.eff.org\/de\/deeplinks\/2015\/01\/healthcare.gov-sends-personal-data."},{"key":"2025122209033188600_ref016","doi-asserted-by":"crossref","unstructured":"Englehardt,\n                                S. and\n                                A.Narayanan.\n                        2016. \u201cOnline tracking: A 1-million-site\n                        measurement and analysis\u201d. URL:\n                        http:\/\/randomwalker.info\/publications\/OpenWPM_1_ million_\n                        site_tracking_measurement.pdf.","DOI":"10.1145\/2976749.2978313"},{"key":"2025122209033188600_ref017","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1145\/2736277.2741679","volume-title":"WWW","author":"Englehardt","year":"2015"},{"key":"2025122209033188600_ref018","unstructured":"Freedom House\n          . 2012.\n                        \u201cFreedom of the press\u201d.\n                        URL: https:\/\/freedomhouse.org\/sites\/default\/files\/FOTP%5C%202012%5C%20Final%5C%20Full%5C%20Report.pdf."},{"key":"2025122209033188600_ref019","unstructured":"\u201cGhostery\u201d.\n                        2016. https:\/\/www.ghostery.com\/."},{"key":"2025122209033188600_ref020","volume-title":"PAMS","author":"Klavri","year":"2016"},{"key":"2025122209033188600_ref021","first-page":"65","volume-title":"ACM SIGCOMM","author":"Krishnamurthy","year":"2006"},{"key":"2025122209033188600_ref022","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1145\/1592665.1592668","volume-title":"ACM workshop on Online Social Networks","author":"Krishnamurthy","year":"2009"},{"key":"2025122209033188600_ref023","first-page":"541","volume-title":"WWW","author":"Krishnamurthy","year":"2009"},{"key":"2025122209033188600_ref024","first-page":"4","volume-title":"USENIX Conference on\n                        Online social networks","author":"Krishnamurthy","year":"2010"},{"key":"2025122209033188600_ref025","doi-asserted-by":"crossref","first-page":"1343","DOI":"10.1145\/2487788.2488173","volume-title":"Proceedings of the 22nd\n                        International Conference on World Wide Web Companion","author":"Kunegis","year":"2013"},{"issue":"3","key":"2025122209033188600_ref026","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1080\/15427951.2014.958250","article-title":"Exploiting the structure of bipartite\n                        graphs for algebraic and spectral graph theory\n                    applications","volume":"11","author":"Kunegis","year":"2015","journal-title":"Internet Math"},{"key":"2025122209033188600_ref027","volume-title":"ACM Web\n                        Science","author":"Lehmberg","year":"2014"},{"key":"2025122209033188600_ref028","volume-title":"International Journal of Communication","author":"Libert","year":"2015"},{"issue":"3","key":"2025122209033188600_ref029","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1145\/2658983","article-title":"Privacy implications of health\n                        information seeking on the Web","volume":"58","author":"Libert","year":"2015","journal-title":"Communications of the ACM"},{"key":"2025122209033188600_ref030","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1109\/SP.2012.47","volume-title":"Security and Privacy\n                        (SP), 2012 IEEE Symposium on","author":"Mayer","year":"2012"},{"key":"2025122209033188600_ref031","first-page":"427","volume-title":"WWW","author":"Meusel","year":"2014"},{"issue":"1","key":"2025122209033188600_ref032","doi-asserted-by":"crossref","DOI":"10.1561\/106.00000003","article-title":"The graph structure in the\n                        Web-Analyzed on different aggregation levels","volume":"1","author":"Meusel","year":"2015","journal-title":"The Journal of Web Science"},{"key":"2025122209033188600_ref033","volume-title":"Tech.\n                        rep","author":"Page","year":"1999"},{"key":"2025122209033188600_ref034","article-title":"The design and implementation of\n                        the Tor browser [draft]","author":"Perry","year":"2015"},{"key":"2025122209033188600_ref035","unstructured":"\u201cPrivacy Badger\u201d.\n                        2016. https:\/\/www.eff.org\/de\/node\/73969"},{"key":"2025122209033188600_ref036","first-page":"12","volume-title":"Proceedings of the 9th USENIX conference on Networked Systems Design\n                        and Implementation","author":"Roesner","year":"2012"},{"key":"2025122209033188600_ref037","volume-title":"Tilburg Law School Legal Studies\n                        Research Paper Series","author":"Roosendaal","year":"2011"},{"key":"2025122209033188600_ref038","volume-title":"ICWSM","author":"Schelter","year":"2016"},{"key":"2025122209033188600_ref039","volume-title":"Tech. rep","author":"Spiegler","year":"2013"},{"key":"2025122209033188600_ref040","unstructured":"The Guardian\n          . 2015a.\n                        \u201cBelgian court orders Facebook to stop tracking\n                        non-members\u201d. URL: http:\/\/www.theguardian.com\/technology\/2015\/nov\/10\/belgian-court-orders-facebook-to-stop-tracking-non-members."},{"key":"2025122209033188600_ref041","unstructured":"The Guardian\n          . 2015b.\n                        \u201cGoogle is returning to China? It never really\n                        left\u201d. URL: https:\/\/www.theguardian.com\/technology\/2015\/sep\/21\/google-is-returning-to-china-it-never-really-left."},{"key":"2025122209033188600_ref042","unstructured":"The Intercept\n          . 2015.\n                        \u201cFrom radio to porn, British spies track web\n                        users' online identities\u201d.\n                        URL: https:\/\/theintercept.com\/2015\/09\/25\/gchq-radio-porn-spies-track-web-users-online-identities\/."},{"key":"2025122209033188600_ref043","unstructured":"Trackography\n          . 2014.\n                        \u201cMeet the trackers; me and my\n                    shadow\u201d. URL: https:\/\/myshadow.org\/trackography-meet-the-trackers."},{"key":"2025122209033188600_ref044","unstructured":"Vsevolod,\n                                P.\n          \n          2014. \u201cRussia beyond the headlines: Yandex\n                        reacts to Putin comments about foreign influence as share price\n                        falls\u201d. URL: https:\/\/rbth.com\/business\/2014\/04\/28\/yandex_reacts_to_putin_comments_about_foreign_influence_as_share_pri_36283.html."},{"key":"2025122209033188600_ref045","unstructured":"\u201cWeb Data Commons - Hyperlink\n                        Graphs\u201d. http:\/\/webdatacommons.org\/hyperlinkgraph\/."},{"key":"2025122209033188600_ref046","unstructured":"Wikipedia\n          . 2015.\n                        \u201cList of countries by English-speaking\n                        population\u201d. URL: https:\/\/en.wikipedia.org\/wiki\/List%5C_of%5C_countries%5C_by%5C_English-speaking%5C_population."},{"key":"2025122209033188600_ref047","unstructured":"The World Bank\n          . 2015.\n                        \u201cPopulation, total\u201d.\n                        URL: http:\/\/data.worldbank.org\/indicator\/SP.POP.TOTL."},{"key":"2025122209033188600_ref048","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1145\/2872427.2883028","volume-title":"Proceedings of the 25th\n                        International Conference on World Wide Web","author":"Yu","year":"2016"}],"container-title":["The Journal of Web Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/jws\/article-pdf\/4\/4\/53\/11141468\/106.00000014en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/jws\/article-pdf\/4\/4\/53\/11141468\/106.00000014en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,22]],"date-time":"2025-12-22T14:03:44Z","timestamp":1766412224000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/jws\/article\/4\/4\/53\/1331689\/On-the-Ubiquity-of-Web-Tracking-Insights-from-a"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,3,19]]},"references-count":48,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2018,3,19]]}},"URL":"https:\/\/doi.org\/10.1561\/106.00000014","relation":{},"ISSN":["2332-4031"],"issn-type":[{"value":"2332-4031","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,3,19]]}}}