{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T11:49:36Z","timestamp":1740138576731,"version":"3.37.3"},"reference-count":39,"publisher":"Walter de Gruyter GmbH","issue":"3-4","funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["LA-2971\/1-1"],"award-info":[{"award-number":["LA-2971\/1-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["GI-711\/5-1"],"award-info":[{"award-number":["GI-711\/5-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,5,27]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Mathematical optimization is at the algorithmic core of machine learning. Almost any known algorithm for solving mathematical optimization problems has been applied in machine learning and the machine learning community itself is actively designing and implementing new algorithms for specific problems. These implementations have to be made available to machine learning practitioners which is mostly accomplished by distributing them as standalone software. Successful well-engineered implementations are collected in machine learning toolboxes that provide a more uniform access to the different solvers. A disadvantage of the toolbox approach is a lack of flexibility as toolboxes only provide access to a fixed set of machine learning models that cannot be modified. This can be a problem for the typical machine learning workflow that iterates the process of modeling, solving and validating. If a model does not perform well on validation data, it needs to be modified. In most cases these modifications require a new solver for the entailed optimization problems. Optimization frameworks that combine a modeling language for specifying optimization problems with a solver are better suited to the iterative workflow since they allow to address large problem classes. Here, we provide examples of the use of optimization frameworks in machine learning. We also illustrate the use of one such framework in a case study that follows the typical machine learning workflow.<\/jats:p>","DOI":"10.1515\/itit-2019-0031","type":"journal-article","created":{"date-parts":[[2020,3,4]],"date-time":"2020-03-04T09:01:53Z","timestamp":1583312513000},"page":"169-180","source":"Crossref","is-referenced-by-count":0,"title":["Optimization frameworks for machine learning: Examples and case study"],"prefix":"10.1515","volume":"62","author":[{"given":"Joachim","family":"Giesen","sequence":"first","affiliation":[{"name":"Friedrich Schiller University Jena , Jena , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S\u00f6ren","family":"Laue","sequence":"additional","affiliation":[{"name":"Friedrich Schiller University Jena , Jena , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthias","family":"Mitterreiter","sequence":"additional","affiliation":[{"name":"Friedrich Schiller University Jena , Jena , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2020,3,4]]},"reference":[{"key":"2023033120254865211_j_itit-2019-0031_ref_001_w2aab3b7d151b1b6b1ab2ab1Aa","unstructured":"Christopher M. Bishop. Pattern Recognition and Machine Learning, 5th Edition. Information science and statistics. Springer, 2007."},{"key":"2023033120254865211_j_itit-2019-0031_ref_002_w2aab3b7d151b1b6b1ab2ab2Aa","doi-asserted-by":"crossref","unstructured":"Trevor Hastie, Robert Tibshirani and Jerome H. Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd Edition. Springer Series in Statistics. Springer, 2009.","DOI":"10.1007\/978-0-387-84858-7"},{"key":"2023033120254865211_j_itit-2019-0031_ref_003_w2aab3b7d151b1b6b1ab2ab3Aa","unstructured":"Kevin P. Murphy. Machine Learning \u2013 A Probabilistic Perspective. Adaptive computation and machine learning series. MIT Press, 2012."},{"key":"2023033120254865211_j_itit-2019-0031_ref_004_w2aab3b7d151b1b6b1ab2ab4Aa","doi-asserted-by":"crossref","unstructured":"David R. Cox. The regression analysis of binary sequences (with discussion). J. Roy. Stat. Soc. B, 20:215\u2013242, 1958.","DOI":"10.1111\/j.2517-6161.1958.tb00292.x"},{"key":"2023033120254865211_j_itit-2019-0031_ref_005_w2aab3b7d151b1b6b1ab2ab5Aa","doi-asserted-by":"crossref","unstructured":"Bernhard Scholkopf, Ralf Herbrich and Alex J. Smola. A generalized representer theorem. In International Conference on Computational Learning Theory (COLT), 2001.","DOI":"10.1007\/3-540-44581-1_27"},{"key":"2023033120254865211_j_itit-2019-0031_ref_006_w2aab3b7d151b1b6b1ab2ab6Aa","unstructured":"Suvrit Sra, Sebastian Nowozin and Stephen J. Wright Optimization for Machine Learning. MIT Press, 2012."},{"key":"2023033120254865211_j_itit-2019-0031_ref_007_w2aab3b7d151b1b6b1ab2ab7Aa","doi-asserted-by":"crossref","unstructured":"Herbert Robbins and Sutton Monro. A stochastic approximation method. Ann. Math. Statist., 22(3):400\u2013407, 1951.","DOI":"10.1214\/aoms\/1177729586"},{"key":"2023033120254865211_j_itit-2019-0031_ref_008_w2aab3b7d151b1b6b1ab2ab8Aa","doi-asserted-by":"crossref","unstructured":"L. Bottou, F. Curtis and J. Nocedal. Optimization Methods for Large-Scale Machine Learning. SIAM Review, 60(2):223\u2013311, 2018.","DOI":"10.1137\/16M1080173"},{"key":"2023033120254865211_j_itit-2019-0031_ref_009_w2aab3b7d151b1b6b1ab2ab9Aa","unstructured":"F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825\u20132830, 2011."},{"key":"2023033120254865211_j_itit-2019-0031_ref_010_w2aab3b7d151b1b6b1ab2ac10Aa","unstructured":"Eibe Frank, Mark A. Hall and Ian H. Witten. The WEKA Workbench. Online Appendix for \u201cData Mining: Practical Machine Learning Tools and Techniques\u201d. Morgan Kaufmann, fourth edition, 2016."},{"key":"2023033120254865211_j_itit-2019-0031_ref_011_w2aab3b7d151b1b6b1ab2ac11Aa","unstructured":"Xiangrui Meng, Joseph Bradley, Burak Yavuz, Evan Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, DB Tsai, Manish Amde, Sean Owen, Doris Xin, Reynold Xin, Michael J. Franklin, Reza Zadeh, Matei Zaharia and Ameet Talwalkar. Mllib: Machine learning in apache spark. Journal of Machine Learning Research, 17(1), January 2016."},{"key":"2023033120254865211_j_itit-2019-0031_ref_012_w2aab3b7d151b1b6b1ab2ac12Aa","doi-asserted-by":"crossref","unstructured":"Chih-Chung Chang and Chih-Jen Lin. LIBSVM: A library for support vector machines. ACM Transactions on Intelligent Systems and Technology, 2:27:1\u201327:27, 2011.","DOI":"10.1145\/1961189.1961199"},{"key":"2023033120254865211_j_itit-2019-0031_ref_013_w2aab3b7d151b1b6b1ab2ac13Aa","doi-asserted-by":"crossref","unstructured":"Jerome H. Friedman, Trevor Hastie and Rob Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1\u201322, 2010.","DOI":"10.18637\/jss.v033.i01"},{"key":"2023033120254865211_j_itit-2019-0031_ref_014_w2aab3b7d151b1b6b1ab2ac14Aa","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu and Xiaoqiang Zheng. TensorFlow: A system for large-scale machine learning. In USENIX Conference on Operating Systems Design and Implementation (OSDI), pages 265\u2013283, 2016."},{"key":"2023033120254865211_j_itit-2019-0031_ref_015_w2aab3b7d151b1b6b1ab2ac15Aa","unstructured":"Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga and Adam Lerer. Automatic differentiation in pytorch. In NIPS Autodiff workshop, 2017."},{"key":"2023033120254865211_j_itit-2019-0031_ref_016_w2aab3b7d151b1b6b1ab2ac16Aa","unstructured":"Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama and Trevor Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014."},{"key":"2023033120254865211_j_itit-2019-0031_ref_017_w2aab3b7d151b1b6b1ab2ac17Aa","doi-asserted-by":"crossref","unstructured":"Andreas Griewank and Andrea Walther. Evaluating derivatives \u2013 principles and techniques of algorithmic differentiation. SIAM, 2008.","DOI":"10.1137\/1.9780898717761"},{"key":"2023033120254865211_j_itit-2019-0031_ref_018_w2aab3b7d151b1b6b1ab2ac18Aa","doi-asserted-by":"crossref","unstructured":"B.\u2009T. Polyak. Some methods of speeding up the convergence of iteration methods. USSR Computational Mathematics and Mathematical Physics, 4(5):1\u201317, 1964.","DOI":"10.1016\/0041-5553(64)90137-5"},{"key":"2023033120254865211_j_itit-2019-0031_ref_019_w2aab3b7d151b1b6b1ab2ac19Aa","unstructured":"Yurii Nesterov. A method for unconstrained convex minimization problem with the rate of convergence O(1\/k2)O(1\/{k^{2}}). Doklady AN USSR (translated as Soviet Math. Docl.), 269, 1983."},{"key":"2023033120254865211_j_itit-2019-0031_ref_020_w2aab3b7d151b1b6b1ab2ac20Aa","unstructured":"Robert Fourer, David M. Gay and Brian W. Kernighan. AMPL: a modeling language for mathematical programming. Thomson\/Brooks\/Cole, 2003."},{"key":"2023033120254865211_j_itit-2019-0031_ref_021_w2aab3b7d151b1b6b1ab2ac21Aa","unstructured":"A. Brooke, D. Kendrick and A. Meeraus. GAMS: release 2.25: a user\u2019s guide. The Scientific press series. Scientific Press, 1992."},{"key":"2023033120254865211_j_itit-2019-0031_ref_022_w2aab3b7d151b1b6b1ab2ac22Aa","doi-asserted-by":"crossref","unstructured":"Iain Dunning, Joey Huchette and Miles Lubin. JuMP: A modeling language for mathematical optimization. SIAM Review, 59(2):295\u2013320, 2017.","DOI":"10.1137\/15M1020575"},{"key":"2023033120254865211_j_itit-2019-0031_ref_023_w2aab3b7d151b1b6b1ab2ac23Aa","unstructured":"CVX Research, Inc. CVX: Matlab software for disciplined convex programming, version 2.1. http:\/\/cvxr.com\/cvx, December 2018."},{"key":"2023033120254865211_j_itit-2019-0031_ref_024_w2aab3b7d151b1b6b1ab2ac24Aa","doi-asserted-by":"crossref","unstructured":"M. Grant and S. Boyd. Graph implementations for nonsmooth convex programs. In V. Blondel, S. Boyd and H. Kimura, editors, Recent Advances in Learning and Control, Lecture Notes in Control and Information Sciences, pages 95\u2013110. 2008.","DOI":"10.1007\/978-1-84800-155-8_7"},{"key":"2023033120254865211_j_itit-2019-0031_ref_025_w2aab3b7d151b1b6b1ab2ac25Aa","doi-asserted-by":"crossref","unstructured":"Akshay Agrawal, Robin Verschueren, Steven Diamond and Stephen Boyd. A rewriting system for convex optimization problems. Journal of Control and Decision, 5(1):42\u201360, 2018.","DOI":"10.1080\/23307706.2017.1397554"},{"key":"2023033120254865211_j_itit-2019-0031_ref_026_w2aab3b7d151b1b6b1ab2ac26Aa","unstructured":"Steven Diamond and Stephen Boyd. CVXPY: A Python-embedded modeling language for convex optimization. Journal of Machine Learning Research, 17(83):1\u20135, 2016."},{"key":"2023033120254865211_j_itit-2019-0031_ref_027_w2aab3b7d151b1b6b1ab2ac27Aa","doi-asserted-by":"crossref","unstructured":"Jacob Mattingley and Stephen Boyd. CVXGEN: A Code Generator for Embedded Convex Optimization. Optimization and Engineering, 13(1):1\u201327, 2012.","DOI":"10.1007\/s11081-011-9176-9"},{"key":"2023033120254865211_j_itit-2019-0031_ref_028_w2aab3b7d151b1b6b1ab2ac28Aa","doi-asserted-by":"crossref","unstructured":"P. Giselsson and S. Boyd. Linear convergence and metric selection for Douglas-Rachford splitting and ADMM. IEEE Transactions on Automatic Control, 62(2):532\u2013544, Feb. 2017.","DOI":"10.1109\/TAC.2016.2564160"},{"key":"2023033120254865211_j_itit-2019-0031_ref_029_w2aab3b7d151b1b6b1ab2ac29Aa","doi-asserted-by":"crossref","unstructured":"Goran Banjac, Bartolomeo Stellato, Nicholas Moehle, Paul Goulart, Alberto Bemporad and Stephen P. Boyd. Embedded code generation using the OSQP solver. In Conference on Decision and Control, (CDC), pages 1906\u20131911, 2017.","DOI":"10.1109\/CDC.2017.8263928"},{"key":"2023033120254865211_j_itit-2019-0031_ref_030_w2aab3b7d151b1b6b1ab2ac30Aa","unstructured":"S\u00f6ren Laue, Matthias Mitterreiter and Joachim Giesen. GENO \u2013 GENeric Optimization for Classical Machine Learning. In Advances in Neural Information Processing Systems (NeurIPS), 2019."},{"key":"2023033120254865211_j_itit-2019-0031_ref_031_w2aab3b7d151b1b6b1ab2ac31Aa","unstructured":"S\u00f6ren Laue, Matthias Mitterreiter and Joachim Giesen. Computing higher order derivatives of matrix and tensor expressions. In Advances in Neural Information Processing Systems (NeurIPS), 2018."},{"key":"2023033120254865211_j_itit-2019-0031_ref_032_w2aab3b7d151b1b6b1ab2ac32Aa","doi-asserted-by":"crossref","unstructured":"S\u00f6ren Laue, Matthias Mitterreiter and Joachim Giesen. A Simple and Efficient Tensor Calculus. In AAAI Conference on Artificial Intelligence (AAAI), 2020.","DOI":"10.1609\/aaai.v34i04.5881"},{"key":"2023033120254865211_j_itit-2019-0031_ref_033_w2aab3b7d151b1b6b1ab2ac33Aa","doi-asserted-by":"crossref","unstructured":"Yurii Nesterov. Smooth minimization of non-smooth functions. Math. Program., 103(1):127\u2013152, 2005.","DOI":"10.1007\/s10107-004-0552-5"},{"key":"2023033120254865211_j_itit-2019-0031_ref_034_w2aab3b7d151b1b6b1ab2ac34Aa","doi-asserted-by":"crossref","unstructured":"Richard H. Byrd, Peihuang Lu, Jorge Nocedal and Ciyou Zhu. A limited memory algorithm for bound constrained optimization. SIAM J. Scientific Computing, 16(5):1190\u20131208, 1995.","DOI":"10.1137\/0916069"},{"key":"2023033120254865211_j_itit-2019-0031_ref_035_w2aab3b7d151b1b6b1ab2ac35Aa","doi-asserted-by":"crossref","unstructured":"Ciyou Zhu, Richard H. Byrd, Peihuang Lu and Jorge Nocedal. Algorithm 778: L-BFGS-B: fortran subroutines for large-scale bound-constrained optimization. ACM Trans. Math. Softw., 23(4):550\u2013560, 1997.","DOI":"10.1145\/279232.279236"},{"key":"2023033120254865211_j_itit-2019-0031_ref_036_w2aab3b7d151b1b6b1ab2ac36Aa","doi-asserted-by":"crossref","unstructured":"Jos\u00e9 Luis Morales and Jorge Nocedal. Remark on \u201calgorithm 778: L-BFGS-B: fortran subroutines for large-scale bound constrained optimization\u201d. ACM Trans. Math. Softw., 38(1):7:1\u20137:4, 2011.","DOI":"10.1145\/2049662.2049669"},{"key":"2023033120254865211_j_itit-2019-0031_ref_037_w2aab3b7d151b1b6b1ab2ac37Aa","doi-asserted-by":"crossref","unstructured":"Magnus R. Hestenes. Multiplier and gradient methods. Journal of Optimization Theory and Applications, 4(5):303\u2013320, 1969.","DOI":"10.1007\/BF00927673"},{"key":"2023033120254865211_j_itit-2019-0031_ref_038_w2aab3b7d151b1b6b1ab2ac38Aa","doi-asserted-by":"crossref","unstructured":"M.\u2009J.\u2009D. Powell. Algorithms for nonlinear constraints that use Lagrangian functions. Mathematical Programming, 14(1):224\u2013248, 1969.","DOI":"10.1007\/BF01588967"},{"key":"2023033120254865211_j_itit-2019-0031_ref_039_w2aab3b7d151b1b6b1ab2ac39Aa","doi-asserted-by":"crossref","unstructured":"Ernesto G. Birgin and Jos\u00e9 Mario Mart\u00ednez. Practical augmented Lagrangian methods for constrained optimization, volume 10 of Fundamentals of Algorithms. SIAM, 2014.","DOI":"10.1137\/1.9781611973365"}],"container-title":["it - Information Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyter.com\/view\/journals\/itit\/62\/3-4\/article-p169.xml","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/itit-2019-0031\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/itit-2019-0031\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,1]],"date-time":"2023-04-01T09:26:59Z","timestamp":1680341219000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/itit-2019-0031\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,4]]},"references-count":39,"journal-issue":{"issue":"3-4","published-online":{"date-parts":[[2020,4,1]]},"published-print":{"date-parts":[[2020,5,27]]}},"alternative-id":["10.1515\/itit-2019-0031"],"URL":"https:\/\/doi.org\/10.1515\/itit-2019-0031","relation":{},"ISSN":["2196-7032","1611-2776"],"issn-type":[{"type":"electronic","value":"2196-7032"},{"type":"print","value":"1611-2776"}],"subject":[],"published":{"date-parts":[[2020,3,4]]}}}