{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,3,31]],"date-time":"2023-03-31T15:12:30Z","timestamp":1680275550216},"reference-count":25,"publisher":"Walter de Gruyter GmbH","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,3,28]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>In previous works we discovered that rule-based systems severely suffer in performance when increasing the number of rules. In order to increase the amount of possible boolean relations while keeping the number of rules fixed, we employ ideas from well known Spatial Transformer Systems and Self-Attention Networks: here, our learned rules are not static but are dynamically adjusted to fit the input data by training a separate rule-prediction system, which is predicting parameter matrices used in Neural Logic Rule Layers. We show, that these networks, termed Adaptive Neural Logic Rule Layers, outperform their static counterpart both in terms of final performance, as well as training stability and excitability during early stages of training.<\/jats:p>","DOI":"10.1515\/auto-2021-0136","type":"journal-article","created":{"date-parts":[[2022,3,10]],"date-time":"2022-03-10T12:12:52Z","timestamp":1646914372000},"page":"257-266","source":"Crossref","is-referenced-by-count":0,"title":["Adopting attention-mechanisms for Neural Logic Rule Layers"],"prefix":"10.1515","volume":"70","author":[{"given":"Jan Niclas","family":"Reimann","sequence":"first","affiliation":[{"name":"Department of Automation Technology and Learning Systems , South Westphalia University of Applied Sciences , L\u00fcbecker Ring 2 , Soest , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andreas","family":"Schwung","sequence":"additional","affiliation":[{"name":"Department of Automation Technology and Learning Systems , South Westphalia University of Applied Sciences , L\u00fcbecker Ring 2 , Soest , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steven X.","family":"Ding","sequence":"additional","affiliation":[{"name":"Automatic Control and Complex Systems , University of Duisburg-Essen , Bismarckst. 81 , Duisburg , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2022,3,11]]},"reference":[{"key":"2023033111222636372_j_auto-2021-0136_ref_001","doi-asserted-by":"crossref","unstructured":"Fukushima, K. and S. Miyake. 1982. Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In: (S.-i.Amari and M.\u2009A. Arbib, eds) Competition and Cooperation in Neural Nets. Springer Berlin Heidelberg, Berlin, Heidelberg, pp.\u2009267\u2013285.","DOI":"10.1007\/978-3-642-46466-9_18"},{"key":"2023033111222636372_j_auto-2021-0136_ref_002","unstructured":"Engstrom, L., B. Tran, D. Tsipras, L. Schmidt and A. Madry. 2019. Exploring the landscape of spatial robustness. In: Proceedings of Machine Learning Research, vol.\u200997. PMLR, pp.\u20091802\u20131811. Available from: https:\/\/proceedings.mlr.press\/v97\/engstrom19a.html."},{"key":"2023033111222636372_j_auto-2021-0136_ref_003","unstructured":"Jo, J. and Y. Bengio. 2017. Measuring the tendency of CNNs to learn surface statistical regularities. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1711.11561."},{"key":"2023033111222636372_j_auto-2021-0136_ref_004","unstructured":"Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.\u2009N. Gomez, L. Kaiser and I. Polosukhin. 2017. Attention is all you need. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1706.03762."},{"key":"2023033111222636372_j_auto-2021-0136_ref_005","unstructured":"Jaderberg, M., K. Simonyan, A. Zisserman et al. 2015. Spatial transformer networks. Advances in neural information processing systems 28: 2017\u20132025."},{"key":"2023033111222636372_j_auto-2021-0136_ref_006","unstructured":"Lin, M., Q. Chen and S. Yan. 2014. Network in network. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1312.4400."},{"key":"2023033111222636372_j_auto-2021-0136_ref_007","unstructured":"M. Stollenga, J. Masci, F. Gomez and J. Schmidhuber (eds). 2014. Deep Networks with Internal Selective Attention through Feedback Connections. NIPS\u201914, vol.\u200927."},{"key":"2023033111222636372_j_auto-2021-0136_ref_008","doi-asserted-by":"crossref","unstructured":"Fukushima, K. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Available from: https:\/\/www.rctn.org\/bruno\/public\/papers\/Fukushima1980.pdf.","DOI":"10.1007\/BF00344251"},{"key":"2023033111222636372_j_auto-2021-0136_ref_009","unstructured":"Goodfellow, I., D. Warde-Farley, M. Mirza, A. Courville and Y. Bengio. 2013. Maxout Networks. In: Proceedings of Machine Learning Research, vol.\u200928. PMLR."},{"key":"2023033111222636372_j_auto-2021-0136_ref_010","doi-asserted-by":"crossref","unstructured":"Schaul, T., T. Glasmachers and J. Schmidhuber. 2011. High dimensions and heavy tails for natural evolution strategies. In: Proceedings of the 13th Annual Conference on Genetic and Evolutionary Computation. New York, NY, USA: Association for Computing Machinery, pp.\u2009845\u2013852.","DOI":"10.1145\/2001576.2001692"},{"key":"2023033111222636372_j_auto-2021-0136_ref_011","unstructured":"Reimann, J.\u2009N. and A. Schwung. 2019. Neural logic rule layers. ArXiv. Available from: http:\/\/arxiv.org\/abs\/1907.00878."},{"key":"2023033111222636372_j_auto-2021-0136_ref_012","doi-asserted-by":"crossref","unstructured":"Lecun, Y., L. Bottou, Y. Bengio and P. Haffner. 1998. Gradient-based learning applied to document recognition. Proceedings of the IEEE 86(11): 2278\u20132324.","DOI":"10.1109\/5.726791"},{"key":"2023033111222636372_j_auto-2021-0136_ref_013","unstructured":"Hasanpour, S.\u2009H., M. Rouhani, M. Fayyaz and M. Sabokrou. 2018. Lets keep it simple, using simple architectures to outperform deeper and more complex architectures. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1608.06037."},{"key":"2023033111222636372_j_auto-2021-0136_ref_014","doi-asserted-by":"crossref","unstructured":"He K., X. Zhang, S. Ren and J. Sun. 2016. Deep Residual Learning for Image Recognition. IEEE Computer Society. Available from: https:\/\/www.cv-foundation.org\/openaccess\/content_cvpr_2016\/papers\/He_Deep_Residual_Learning_CVPR_2016_paper.pdf.","DOI":"10.1109\/CVPR.2016.90"},{"key":"2023033111222636372_j_auto-2021-0136_ref_015","unstructured":"Krizhevsky, A., V. Nair and G. Hinton. Cifar-10 (canadian institute for advanced research). Available from: http:\/\/www.cs.toronto.edu\/~kriz\/cifar.html."},{"key":"2023033111222636372_j_auto-2021-0136_ref_016","doi-asserted-by":"crossref","unstructured":"Akiba, T., S. Sano, T. Yanase, T. Ohta and M. Koyama. 2019. Optuna: A next-generation hyperparameter optimization framework. In: Proceedings of the 25rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.","DOI":"10.1145\/3292500.3330701"},{"key":"2023033111222636372_j_auto-2021-0136_ref_017","unstructured":"Zeiler, M.\u2009D. 2012. Adadelta: An adaptive learning rate method. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1212.5701."},{"key":"2023033111222636372_j_auto-2021-0136_ref_018","doi-asserted-by":"crossref","unstructured":"Kiefer, J. and J. Wolfowitz. 1952. Stochastic estimation of the maximum of a regression function. Annals of Mathematical Statistics 23: 462\u2013466.","DOI":"10.1214\/aoms\/1177729392"},{"key":"2023033111222636372_j_auto-2021-0136_ref_019","unstructured":"Glorot, X. and Y. Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In: (Y.\u2009W. Teh and M. Titterington, eds) Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics. Proceedings of Machine Learning Research, vol.\u20099. Chia Laguna Resort, Sardinia, Italy: PMLR, pp.\u2009249\u2013256. Available from: https:\/\/proceedings.mlr.press\/v9\/glorot10a.html."},{"key":"2023033111222636372_j_auto-2021-0136_ref_020","unstructured":"Kingma, D.\u2009P. and J. Ba. 2017. Adam: A method for stochastic optimization. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1412.6980."},{"key":"2023033111222636372_j_auto-2021-0136_ref_021","unstructured":"Bergstra, J., R. Bardenet, Y. Bengio and B. K\u00e9gl. 2011. Algorithms for hyper-parameter optimization. In: Proceedings of the 24th International Conference on Neural Information Processing Systems. NIPS\u201911. Red Hook, NY, USA: Curran Associates Inc, pp.\u20092546\u20132554."},{"key":"2023033111222636372_j_auto-2021-0136_ref_022","unstructured":"Goodfellow, I., Y. Bengio and A. Courville. 2016. Deep Learning. MIT Press."},{"key":"2023033111222636372_j_auto-2021-0136_ref_023","doi-asserted-by":"crossref","unstructured":"Zeiler, M.\u2009D. and R. Fergus. 2014. Visualizing and understanding convolutional networks. In: In Computer Vision\u2013ECCV 2014. Springer, pp.\u2009818\u2013833.","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"2023033111222636372_j_auto-2021-0136_ref_024","unstructured":"Kokhlikyan, N., V. Miglani, M. Martin, E. Wang, B. Alsallakh, J. Reynolds, A. Melnikov, N. Kliushkina, C. Araya, S. Yan and O. Reblitz-Richardson. 2020. Captum: A unified and generic model interpretability library for pytorch. ArXiv. Available from: https:\/\/arxiv.org\/abs\/2009.07896."},{"key":"2023033111222636372_j_auto-2021-0136_ref_025","unstructured":"Bengio, Y., N. L\u00e9onard and A. Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. ArXiv. Available from: https:\/\/arxiv.org\/abs\/1308.3432."}],"container-title":["at - Automatisierungstechnik"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/auto-2021-0136\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/auto-2021-0136\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,31]],"date-time":"2023-03-31T14:33:58Z","timestamp":1680273238000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/auto-2021-0136\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,1]]},"references-count":25,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3,11]]},"published-print":{"date-parts":[[2022,3,28]]}},"alternative-id":["10.1515\/auto-2021-0136"],"URL":"https:\/\/doi.org\/10.1515\/auto-2021-0136","relation":{},"ISSN":["2196-677X","0178-2312"],"issn-type":[{"value":"2196-677X","type":"electronic"},{"value":"0178-2312","type":"print"}],"subject":[],"published":{"date-parts":[[2022,3,1]]}}}