{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T04:28:19Z","timestamp":1772252899215,"version":"3.50.1"},"reference-count":25,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2017,4,15]],"date-time":"2017-04-15T00:00:00Z","timestamp":1492214400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>This paper presents a novel CNN-based architecture, referred to as Q-Net, to learn local feature descriptors that are useful for matching image patches from two different spectral bands. Given correctly matched and non-matching cross-spectral image pairs, a quadruplet network is trained to map input image patches to a common Euclidean space, regardless of the input spectral band. Our approach is inspired by the recent success of triplet networks in the visible spectrum, but adapted for cross-spectral scenarios, where, for each matching pair, there are always two possible non-matching patches: one for each spectrum. Experimental evaluations on a public cross-spectral VIS-NIR dataset shows that the proposed approach improves the state-of-the-art. Moreover, the proposed technique can also be used in mono-spectral settings, obtaining a similar performance to triplet network descriptors, but requiring less training data.<\/jats:p>","DOI":"10.3390\/s17040873","type":"journal-article","created":{"date-parts":[[2017,4,18]],"date-time":"2017-04-18T11:22:04Z","timestamp":1492514524000},"page":"873","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":35,"title":["Cross-Spectral Local Descriptors via Quadruplet Network"],"prefix":"10.3390","volume":"17","author":[{"given":"Cristhian","family":"Aguilera","sequence":"first","affiliation":[{"name":"Computer Vision Center, Edifici O, Campus UAB, Bellaterra 08193, Barcelona, Spain"},{"name":"Computer Science Department, Universitat Aut\u00f2noma de Barcelona, Campus UAB, Bellaterra 08193,Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Angel","family":"Sappa","sequence":"additional","affiliation":[{"name":"Computer Vision Center, Edifici O, Campus UAB, Bellaterra 08193, Barcelona, Spain"},{"name":"Facultad de Ingenier\u00eda en Electricidad y Computaci\u00f3n, CIDIS, Escuela Superior Polit\u00e9cnica del Litoral,ESPOL, Campus Gustavo Galindo, Km 30.5 v\u00eda Perimetral, Guayaquil 09-01-5863, Ecuador"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cristhian","family":"Aguilera","sequence":"additional","affiliation":[{"name":"DIEE, University of B\u00edo-B\u00edo, Concepci\u00f3n 4051381, Concepci\u00f3n, Chile"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ricardo","family":"Toledo","sequence":"additional","affiliation":[{"name":"Computer Vision Center, Edifici O, Campus UAB, Bellaterra 08193, Barcelona, Spain"},{"name":"Computer Science Department, Universitat Aut\u00f2noma de Barcelona, Campus UAB, Bellaterra 08193,Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2017,4,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MMUL.2012.24","article-title":"Microsoft Kinect Sensor and Its Effect","volume":"19","author":"Zhang","year":"2012","journal-title":"IEEE MultiMedia"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"You, C.W., Lane, N.D., Chen, F., Wang, R., Chen, Z., Bao, T.J., Montes-de Oca, M., Cheng, Y., Lin, M., and Torresani, L. (2013, January 25\u201328). CarSafe app: Alerting drowsy and distracted drivers using dual cameras on smartphones. Proceeding of the 11th Annual International Conference on Mobile Systems, Applications, and Services (ACM 2013), Taipei, Taiwan.","DOI":"10.1145\/2462456.2466711"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive Image Features from Scale-Invariant Keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_4","unstructured":"Balntas, V., Johns, E., Tang, L., and Mikolajczyk, K. (arXiv, 2016). PN-Net: Conjoined Triple Deep Network for Learning Local Image Descriptors, arXiv."},{"key":"ref_5","unstructured":"Github (2017, April 15). Qnet. Available online: http:\/\/github.com\/ngunsu\/qnet."},{"key":"ref_6","unstructured":"Yi, D., Lei, Z., and Li, S.Z. (2015, January 4\u20138). Shared representation learning for heterogenous face recognition. Proceeding of the 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Ljubljana, Slovenia."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ring, E., and Ammer, K. (2015). The technique of infrared imaging in medicine. Infrared Imaging, IOP Publishing.","DOI":"10.1088\/978-0-7503-1143-4"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1410","DOI":"10.1109\/TPAMI.2012.229","article-title":"Heterogeneous face recognition using kernel prototype similarities","volume":"35","author":"Klare","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","unstructured":"Aguilera, C.A., Aguilera, F.J., Sappa, A.D., Aguilera, C., and Toledo, R. (July, January 26). Learning cross-spectral similarity measures with deep convolutional neural networks. Proceeding of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Las Vegas, NV, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Brown, M., and Susstrunk, S. (2011, January 20\u201325). Multi-spectral SIFT for scene category recognition. Proceeding of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995637"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Bay, H., Tuytelaars, T., and Van Gool, L. (2006). Surf: Speeded up robust features. European Conference on Computer Vision, Springer.","DOI":"10.1007\/11744023_32"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Firmenichy, D., Brown, M., and S\u00fcsstrunk, S. (2011, January 11\u201314). Multispectral interest points for RGB-NIR image registration. Proceeding of the 2011 18th IEEE International Conference on Image Processing (ICIP), Brussels, Belgium.","DOI":"10.1109\/ICIP.2011.6115818"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Pinggera, P., Breckon, T., and Bischof, H. (2012, January 3\u20137). On Cross-Spectral Stereo Matching using Dense Gradient Features. Proceeding of the British Machine Vision Conference, Surrey, UK.","DOI":"10.5244\/C.26.103"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Morris, N.J.W., Avidan, S., Matusik, W., and Pfister, H. (2007, January 18\u201323). Statistics of Infrared Images. Proceeding of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383003"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"12661","DOI":"10.3390\/s120912661","article-title":"Multispectral image feature points","volume":"12","author":"Aguilera","year":"2012","journal-title":"Sensors"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1210","DOI":"10.1109\/TITS.2014.2354731","article-title":"Multispectral Stereo Odometry","volume":"16","author":"Mouats","year":"2015","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Aguilera, C.A., Sappa, A.D., and Toledo, R. (2015, January 27\u201330). LGHD: A feature descriptor for matching across non-linear intensity variations. Proceeding of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec, Canada.","DOI":"10.1109\/ICIP.2015.7350783"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"772","DOI":"10.1016\/j.patcog.2014.09.005","article-title":"Non-rigid visible and infrared face registration via regularized Gaussian fields criterion","volume":"48","author":"Ma","year":"2015","journal-title":"Pattern Recognit."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Shen, X., Xu, L., Zhang, Q., and Jia, J. (2014). Multi-modal and Multi-spectral Registration for Natural Images. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10593-2_21"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zagoruyko, S., and Komodakis, N. (2015, January 7\u201312). Learning to compare image patches via convolutional neural networks. Proceeding of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299064"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Simo-Serra, E., Trulls, E., Ferraz, L., Kokkinos, I., Fua, P., and Moreno-Noguer, F. (2015, January 13\u201316). Discriminative Learning of Deep Convolutional Feature Point Descriptors. Proceedings of the International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.22"},{"key":"ref_22","unstructured":"Han, X., Leung, T., Jia, Y., Sukthankar, R., and Berg, A.C. (2015, January 7\u201312). MatchNet: Unifying Feature and Metric Learning for Patch-Based Matching. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_23","first-page":"1","article-title":"Stereo matching by training a convolutional neural network to compare image patches","volume":"17","author":"Zbontar","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Winder, S., Hua, G., and Brown, M. (2009, January 20\u201325). Picking the best DAISY. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2009), Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206839"},{"key":"ref_25","unstructured":"Collobert, R., Kavukcuoglu, K., and Farabet, C. (2011, January 12\u201317). Torch7: A matlab-like environment for machine learning. Proceedings of the BigLearn, NIPS Workshop, Granada, Spain. Number EPFL-CONF-192376."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/17\/4\/873\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:32:45Z","timestamp":1760207565000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/17\/4\/873"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,4,15]]},"references-count":25,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2017,4]]}},"alternative-id":["s17040873"],"URL":"https:\/\/doi.org\/10.3390\/s17040873","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints201703.0061.v1","asserted-by":"object"}]},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,4,15]]}}}