{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"institution":[{"id":[{"id":"https:\/\/ror.org\/03mb6wj31","id-type":"ROR","asserted-by":"publisher"},{"id":"https:\/\/www.isni.org\/000000041937028X","id-type":"ISNI","asserted-by":"publisher"},{"id":"https:\/\/www.wikidata.org\/entity\/Q1640731","id-type":"wikidata","asserted-by":"publisher"}],"name":"Universitat Polit\u00e8cnica de Catalunya","acronym":["UPC"]}],"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T12:25:56Z","timestamp":1785414356680,"version":"3.56.0"},"reference-count":0,"publisher":"Universitat Polit\u00e8cnica de Catalunya","license":[{"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>Data have become number one assets of today's business world. Thus, its exploitation and analysis attracted the attention of people from different fields and having different technical backgrounds. Data-intensive flows are central processes in today's business intelligence (BI) systems, deploying different technologies to deliver data, from a multitude of data sources, in user-preferred and analysis-ready formats. However, designing and optimizing such data flows, to satisfy both users' information needs and agreed quality standards, have been known as a burdensome task, typically left to the manual efforts of a BI system designer. These tasks have become even more challenging  for next generation BI systems, where data flows typically need to combine data from in-house transactional storages, and data coming from external sources, in a variety of formats (e.g., social media, governmental data, news feeds). Moreover, for making an impact to business outcomes, data flows are expected to answer unanticipated analytical needs of a broader set of business users' and deliver valuable information in near real-time (i.e., at the right time).  These challenges largely indicate a need for boosting the automation of the design and optimization of data-intensive flows. \r\n\r\nThis PhD thesis aims at providing automatable means for managing the lifecycle of data-intensive flows. The study primarily analyzes the remaining challenges to be solved in the field of data-intensive flows, by performing a survey of current literature, and envisioning an architecture for managing the lifecycle of data-intensive flows. Following the proposed architecture, we further focus on providing automatic techniques for covering different phases of the data-intensive flows' lifecycle. In particular, the thesis first proposes an approach (CoAl) for incremental design of data-intensive flows, by means of multi-flow consolidation. CoAl not only facilitates the maintenance of data flow designs in front of changing information needs, but also supports the multi-flow optimization of data-intensive flows, by maximizing their reuse. Next, in the data warehousing (DW) context, we propose a complementary method (ORE) for incremental design of the target DW schema, along with systematically tracing the evolution metadata, which can further facilitate the design of back-end data-intensive flows (i.e., ETL processes). The thesis then studies the problem of implementing data-intensive flows into deployable formats of different execution engines, and proposes the BabbleFlow system for translating logical data-intensive flows into executable formats, spanning single or multiple execution engines. Lastly, the thesis focuses on managing the execution of data-intensive flows on distributed data processing platforms, and to this end, proposes an algorithm (H-WorD) for supporting the scheduling of data-intensive flows by workload-driven redistribution of data in computing clusters.  \r\n\r\nThe overall outcome of this thesis an end-to-end platform for managing the lifecycle of data-intensive flows, called Quarry. The techniques proposed in this thesis, plugged to the Quarry platform, largely facilitate the manual efforts, and assist users of different technical skills in their analytical tasks. Finally, the results of this thesis largely contribute to the field of data-intensive flows in today's BI systems, and advocate for further attention by both academia and industry to the problems of design and optimization of data-intensive flows.<\/jats:p>\n                <jats:p>Actualment, les dades han esdevingut el principal actiu del m\u00f3n empresarial. En conseq\u00fc\u00e8ncia, la seva explotaci\u00f3 i an\u00e0lisi ha atret l'atenci\u00f3 de gent provinent de diferents camps i experi\u00e8ncia t\u00e8cnica. Els fluxes de dades intensius s\u00f3n processos centrals en els actuals sistemes d'intelig\u00e8ncia de negoci (BI), desplegant diferents tecnologies per a proporcionar dades, provinents de diferents fonts i centrant-se en formats orientats a l'usuari. Tantmateix, el disseny i l'optimitzaci\u00f3 de tals fluxes, per tal de satisfer ambd\u00f3s usuaris de la informaci\u00f3 i els est\u00e0ndars de qualitat, resulta una tasca tediosa, normalment dirigida als esfor\u00e7os manuals del dissenyador del sistema BI. Aquestes tasques han esdevingut encara m\u00e9s complexes en el context dels sistemes BI de nova generaci\u00f3, on els fluxes de dades t\u00edpicament combinen dades internes de fonts transaccionals, amb dades externes representades amb diferents formats (xarxes socials, dades governamentals, not\u00edcies). A m\u00e9s a m\u00e9s, per tal de tenir un impacte en el negoci, s'espera que els fluxes de dades responguin a necessitats anal\u00edtiques no anticipades en un marge de temps proper a temps real. Aquests reptes clarament indiquen la necessitat de millora en l'automatitzaci\u00f3 del disseny i optimitzaci\u00f3 dels fluxes de dades intensius. L'objectiu d'aquesta tesi doctoral \u00e9s el de proporcionar mitjans autom\u00e0tics per tal de manegar el cicle de vida de fluxes de dades intensius. L'estudi primerament analitza els reptes pendents de resoldre en l'\u00e0rea de fluxes intensius de dades, mitjan\u00e7ant l'an\u00e0lisi de la literatura recent, i concebent una arquitectura per a la gesti\u00f3 del cicle de vida dels fluxes de dades intensius. A partir de l'arquitectura proposada, ens centrem en la proposta de t\u00e8cniques autom\u00e0tiques per tal de cobrir cadascuna de les fases del cicle de vida dels fluxes intensius de dades. Particularment, aquesta tesi inicialment proposa una t\u00e8cnica (CoAl) per el disseny incremental dels fluxes de dades intensius, mitjan\u00e7ant la consolidaci\u00f3 de multiples fluxes. CoAl no nom\u00e9s facilita el manteniment dels flux de dades davant de noves necessitats d'informaci\u00f3, sin\u00f3 que tamb\u00e9 permet la optimitzaci\u00f3 de m\u00faltiples fluxes mitjan\u00e7ant la maximitzaci\u00f3 de la reusabilitat. Posteriorment, en un contexte de magatzems de dades (DW), proposem un m\u00e8tode complementari (ORE) per el disseny incremental d'un esquema de DW objectiu, acompanyat per la tra\u00e7a sistem\u00e0tica de metadades d'evoluci\u00f3, les quals poden facilitar el disseny dels fluxes intensius de dades (processos ETL). A continuaci\u00f3, la tesi estudia el problema d'implementaci\u00f3 de fluxes de dades intensius a diferents sistemes d'execuci\u00f3, i proposa el sistema BabbleFlow per la traducci\u00f3 de fluxes de dades intensius l\u00f2gics a formats executables, a un o m\u00faltiples sistemes d'execuci\u00f3. Finalment, la tesi es centra en la gesti\u00f3 dels fluxes de dades intensius en plataformes distribu\u00efdes de processament de dades, amb aquest objectiu es proposa un algorisme (H-WorD) per donar suport a la planificaci\u00f3 de l'execuci\u00f3 de fluxes intensius de dades mitjan\u00e7ant la redistribuci\u00f3 de dades dirigides per la carga de treball. El resultat general d'aquesta tesi \u00e9s una plataforma d'inici a fi per tal de gestionar el cicle de vida dels fluxes intensius de dades, anomenada Quarry. Les t\u00e8cniques propostes en aquesta tesi, incorporades a la plataforma Quarry, en gran part simplifiquen els esfor\u00e7os manuals i assisteixen usuaris amb diferent experi\u00e8ncia t\u00e8cnica a les seves tasques anal\u00edtiques. Finalment, els resultats d'aquesta tesi contribueixen a l'\u00e0rea de fluxes intensius de dades en els sistemes de BI actuals. A m\u00e9s a m\u00e9s, reflecteixen la necessitat de m\u00e9s atenci\u00f3 per part dels mons acad\u00e8mic i industrial als problemes de disseny i optimitzaci\u00f3 de fluxes de dades intensius.<\/jats:p>","DOI":"10.5821\/dissertation-2117-100960","type":"dissertation","created":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T01:22:35Z","timestamp":1732152155000},"approved":{"date-parts":[[2016,9,26]]},"source":"Crossref","is-referenced-by-count":0,"title":["Requirement-driven design and optimization of data-intensive flows"],"prefix":"10.5821","author":[{"given":"Petar","family":"Jovanovic","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"3865","container-title":[],"original-title":[],"contributor":[{"sequence":"additional","affiliation":[],"role":[null]}],"deposited":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T06:26:32Z","timestamp":1773901592000},"score":1,"resource":{"primary":{"URL":"https:\/\/hdl.handle.net\/2117\/100960"}},"subtitle":[],"editor":[{"given":"Alberto","family":"Abell\u00f3 Gamazo","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]},{"given":"Toon","family":"Calders","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]},{"given":"\u00d3scar","family":"Romero Moral","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[null]]},"references-count":0,"URL":"https:\/\/doi.org\/10.5821\/dissertation-2117-100960","relation":{},"subject":[]}}