{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"institution":[{"id":[{"id":"https:\/\/ror.org\/03mb6wj31","id-type":"ROR","asserted-by":"publisher"},{"id":"https:\/\/www.isni.org\/000000041937028X","id-type":"ISNI","asserted-by":"publisher"},{"id":"https:\/\/www.wikidata.org\/entity\/Q1640731","id-type":"wikidata","asserted-by":"publisher"}],"name":"Universitat Polit\u00e8cnica de Catalunya","acronym":["UPC"]}],"indexed":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T18:12:55Z","timestamp":1770747175727,"version":"3.49.0"},"reference-count":0,"publisher":"Universitat Polit\u00e8cnica de Catalunya","license":[{"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>Query expansion techniques aim at improving the results achieved by a user's query by means of introducing new expansion terms, called expansion features. Expansion features introduce new concepts that are semantically related with the concepts in the user's query and that allow retrieving documents that otherwise would be not. Thus, the challenge is to select those expansion features that are capable of improving the results the most. A bad choice of expansion features may be counterproductive. In this thesis, we use an external source of information, a Knowledge Base (KB), as source expansion features.\r\n\r\nA knowledge base consists of a set of entries, each of which represent a concept and has, at least, a name, which can be used as expansion feature. The techniques framed in this family have become more popular due to the increase of available data, as, for example, Wikipedia. Particularly, we focus on exploiting those KB whose entries are linked to each other, conforming a graph of entries. To the best of our knowledge, most of the techniques framed on the KB family rely on some kind of text analysis, such as explicit semantic analysis, or are based on other existing query expansion techniques such as pseudo relevance feedback. However, the underlying net-work structure of KBs has been barely exploited.\r\n\r\nIn this thesis, we show that the structure can be used to identify reliable expansion feature for the query expansion process. Thus, we design a novel expansion technique, Structural Query Expansion (SQE). For SQE to benefit from the particular structures of KBs, we propose a methodology to identify the structural characteristics that, given a query, allow identifying those nodes in the KB that are good candidates to be used as source of expansion features, called from now on expansion nodes. The methodology consists in building a ground truth that connects each query from a query set with those nodes of the KB that when used to extract the expansion features allow achieving the best results in terms of precision, we call the set of those nodes, expansion query graph. Then, we compare the expansion query graph of each query to find shared characteristics. SQE materializes the revealed characteristics into a set of structural motifs. In the particular case of Wikipedia, we have found two motifs called triangular and square. In the former, the query node and the expansion node are doubly linked and the expansion node belongs to, at least, the same categories as the query node. In the latter, the query node and the expansion node also are doubly linked and their categories are connected somehow. These motifs are used to, given a query and its query nodes, identify all the expansion nodes which are used as source of expansion features. Notice that we have designed this technique to be orthogonal to others because is fully decoupled from the search process and does not depend on the particular collection of documents.\r\n\r\nWe have tested our techniques with three different datasets to avoid any kind of overfitting. The results are shown to be consistent among the three of them. Also, the results which are validated with statistical significance tests, show that SQE is capable to achieve up to 150% improvement in the precision. Finally, we show the performance of our technique which runs in sub-second times (358.23ms at maximum) which makes it feasible for a real query expansion system. This is especially relevant because, to the best of our knowledge, the performance is an aspect that is being ignored in most of the works and, thus, it is difficult to know whether they can be include in real systems or not.<\/jats:p>\n                <jats:p>Les t\u00e8cniques d'expansi\u00f3 de consultes tenen com a objecte millorar els resultats obtinguts per la consulta d'un usuari a partir de la introducci\u00f3 de termes d'expansi\u00f3, anomenat caracter\u00edstiques d'expansi\u00f3. Les caracter\u00edstiques d'expansi\u00f3 introdueixen nous conceptes que estan relacionats sem\u00e0nticament amb els conceptes de la consulta de l'usuari i que permeten obtenir documents que d'altra manera no es podrien obtenir. Per tant, el repte \u00e9s seleccionar les caracter\u00edstiques d'expansi\u00f3 que s\u00f3n capaces de millorar al m\u00e0xim els resultats, doncs una mala elecci\u00f3 pot ser contra-productiva. En aquesta tesis, utilitzem una font externa d'informaci\u00f3, una Base de Coneixement (KB), com a font de caracter\u00edstiques d'expansi\u00f3. Una KB \u00e9s un conjunt d'entrades, cadascuna de les quals representa un concepte i que t\u00e9, com a m\u00ednim, un nom, que \u00e9s susceptible de ser usat com a caracter\u00edstica d'expansi\u00f3. Les t\u00e8cniques emmarcades en aquesta fam\u00edlia han esdevingut populars degut al creixement de la informaci\u00f3 disponible, per exemple, Wikipedia. Particularment, nosaltres en centrem en utilitzar aquelles KB les entrades de les quals estan relacionades entre si, conformant d'aquesta manera, un graf d'entrades. Segons les nostres informacions, la majora de les t\u00e8cniques emmarcades en aquesta fam\u00edlia utilitzen algun tipus d'an\u00e0lisi ling\u00fc\u00edstic, o estan basades en d'altres t\u00e8cniques com relevance feedback. Ara b\u00e9, la estructura subjacent de la xarxa gaireb\u00e9 no s'ha utilitzat. En aquesta tesis, mostrem que la estructura es pot fer servir per identificar caracter\u00edstiques d'expansi\u00f3 fiables pel proc\u00e9s d'expansi\u00f3 de consultes. De fet, proposem una t\u00e8cnica d'expansi\u00f3 novell, Structural Query Expansion (SQE), que la explota. Perqu\u00e8 SQE pugui beneficiar-se de les particularitats estructurals de les KBs, hem proposat tamb\u00e9 una metodologia per revelar les caracter\u00edstiques estructurals que, donada una consulta, permeten identificar aquells nodes que s\u00f3n una bona font de caracter\u00edstiques d'expansi\u00f3, els anomenats, nodes d'expansi\u00f3. Aquesta metodologia consisteix en construir un ground truth que relaciona una conjunt de consultes amb el seu optimal expansion query graph. L'optimal expansion query graph \u00e9s el conjunt de nodes d'expansi\u00f3 que quan s'utilitzen com a font de caracter\u00edstiques d'expansi\u00f3, permeten obtenir els millors resultats en termes de precisi\u00f3. Un cop tenim els optimal expansion query graphs, els comparem entre si per a buscar caracter\u00edstiques compartides. SQE materialitza aquestes caracter\u00edstiques en un conjunt de motius estructurals. En el cas de Wikipedia hem trobat 2 motius: el triangular i el quadr\u00e0tic. En els dos casos el node de la consulta ha d'estar doblement lincat amb el node d'expansi\u00f3. En el triangular, les categories del node d'expansi\u00f3 ha de pert\u00e0nyer, com a m\u00ednim, a les mateixes categories que el node de la consulta, mentre que en el quadr\u00e0tic tan sols cal que les categories del node de la consulta i el d'expansi\u00f3 estiguin relacionades. Aquest motius s'utilitzen per, donada una consulta, identificar tots els seus nodes d'expansi\u00f3. Hem dissenyat aquesta t\u00e8cnica com una t\u00e8cnica ortogonal a d'altres ja que est\u00e0 desacoblada del proc\u00e9s de cerca i no dep\u00e8n de la col\u00b7lecci\u00f3 de documents. Hem provar la nostra t\u00e8cnica amb 3 jocs de dades diferents per a evitar qualsevol tipus d'especialitzaci\u00f3. Els resultats s\u00f3n consistents entre els tres. Hem validat els resultats amb testos de significan\u00e7a estad\u00edstica obtenint millores del 150% en la precisi\u00f3. Finalment, pel que fa el rendiment de la nostra proposta, mostrem que s'executa en mil\u00b7lisegons, i aix\u00f2 la fa susceptible de ser utilitzada en sistemes d'expansi\u00f3 reals. Aix\u00f2 \u00e9s especialment rellevant perqu\u00e8, segons les nostres informacions, aquest \u00e9s un aspecte que s'ignora en la literatura i, per tant, \u00e9s dif\u00edcil de saber la viabilitat de les propostes que existeixen en entorns reals.<\/jats:p>","DOI":"10.5821\/dissertation-2117-113286","type":"dissertation","created":{"date-parts":[[2023,7,19]],"date-time":"2023-07-19T02:17:06Z","timestamp":1689733026000},"approved":{"date-parts":[[2017,9,28]]},"source":"Crossref","is-referenced-by-count":0,"title":["Query expansion by relying on the structure of knowledge bases"],"prefix":"10.5821","author":[{"sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joan","family":"Guisado G\u00e1mez","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"3865","container-title":[],"original-title":[],"deposited":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T06:36:31Z","timestamp":1770705391000},"score":1,"resource":{"primary":{"URL":"https:\/\/hdl.handle.net\/2117\/113286"}},"subtitle":[],"editor":[{"given":"Josep","family":"Larriba Pey","sequence":"first","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[null]]},"references-count":0,"URL":"https:\/\/doi.org\/10.5821\/dissertation-2117-113286","relation":{},"subject":[]}}