{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T17:01:31Z","timestamp":1782234091079,"version":"3.54.5"},"reference-count":12,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2024,11,1]],"date-time":"2024-11-01T00:00:00Z","timestamp":1730419200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Data Science"],"published-print":{"date-parts":[[2024,11,25]]},"abstract":"<jats:p>The human-readable simplicity with which the CSV format was devised, together with the absence of a standard that strictly defines this format, has allowed the proliferation of several variants in the dialects with which these files are written. The latter has meant that the exchange of information between data management systems, or between countries and regions, requires human intervention during the data mining and cleansing process. This has led to the development of various computational tools that aim to accurately determine the dialects of CSV files, in order to avoid data loss at data loading stage in a given system. However, the dialect detection is a complex problem and current systems have limitations or make assumptions that need to be improved and\/or extended. This paper proposes a method for determining CSV file dialects through table uniformity, a statistical approach based on table consistency and records dispersion measurement along with the detection of data type over each field. The new method has a 93.38% average accuracy on a dataset with 548 CSV files composed of samples coming from a data load testing framework, the test suite provided by the CSV on the Web Working Group (CSVW), curated experimental data set from similar tool development and some others CSV files added as verification of the parsing routines. In tests, the proposed solution outperforms the state-of-the-art tool by achieving an average improvement of 16.45%, resulting in an net increment of about 10% in the accuracy with which dialects are detected on truly messy data for this research dataset. Furthermore, the proposed method is accurate enough to determine dialects by reading only ten records, requiring more data to disambiguate those cases where the first records do not contain the necessary information to conclude with a dialect determination.<\/jats:p>","DOI":"10.3233\/ds-240062","type":"journal-article","created":{"date-parts":[[2024,7,26]],"date-time":"2024-07-26T10:51:41Z","timestamp":1721991101000},"page":"55-72","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Detecting CSV file dialects by table uniformity measurement and data type\u00a0inference"],"prefix":"10.1177","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9620-1119","authenticated-orcid":false,"given":"Wilfredo","family":"Garc\u00eda","sequence":"first","affiliation":[{"name":"CEO office, ECP Solutions, Santiago, Rep\u00fablica Dominicana"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2024,11,1]]},"reference":[{"key":"ref001","doi-asserted-by":"publisher","DOI":"10.17713\/ajs.v38i3.272"},{"key":"ref002","doi-asserted-by":"publisher","DOI":"10.1145\/2830508"},{"key":"ref003","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407810"},{"key":"ref004","doi-asserted-by":"publisher","DOI":"10.1145\/3085504.3085520"},{"key":"ref005","doi-asserted-by":"publisher","DOI":"10.18420\/BTW2023-20"},{"key":"ref006","unstructured":"S.\u00a0Idreos et al., Here are my data files. Here are my queries. Where are my results? in: Proceedings of 5th Biennial Conference on Innovative Data Systems Research. Biennial Conference on Innovative Data Systems Research (CIDR 2011), Asilomar, California, USA, 2011, pp.\u00a057\u201368. https:\/\/www.cidrdb.org\/cidr2011\/Papers\/CIDR11_Paper7.pdf (visited on 07\/24\/2021)."},{"key":"ref007","doi-asserted-by":"publisher","DOI":"10.14778\/2732977.2732986"},{"key":"ref008","doi-asserted-by":"publisher","DOI":"10.1109\/OBD.2016.18"},{"key":"ref009","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3222538.url"},{"key":"ref010","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220057"},{"key":"ref011","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-019-00646-y"},{"key":"ref012","doi-asserted-by":"publisher","DOI":"10.14778\/3594512.3594518"}],"container-title":["Data Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/DS-240062","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/DS-240062","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/DS-240062","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T18:10:11Z","timestamp":1777399811000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/DS-240062"}},"subtitle":[],"editor":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8596-222X","authenticated-orcid":false,"given":"Ruben","family":"Verborgh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1267-0234","authenticated-orcid":false,"given":"Tobias","family":"Kuhn","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2024,11,1]]},"references-count":12,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,11,25]]}},"alternative-id":["10.3233\/DS-240062"],"URL":"https:\/\/doi.org\/10.3233\/ds-240062","relation":{},"ISSN":["2451-8484","2451-8492"],"issn-type":[{"value":"2451-8484","type":"print"},{"value":"2451-8492","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,1]]}}}