{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T11:12:07Z","timestamp":1780571527627,"version":"3.54.1"},"reference-count":60,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2023,3,9]],"date-time":"2023-03-09T00:00:00Z","timestamp":1678320000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,3,9]],"date-time":"2023-03-09T00:00:00Z","timestamp":1678320000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","award":["01IS20019A"],"award-info":[{"award-number":["01IS20019A"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"publisher","award":["201806030269"],"award-info":[{"award-number":["201806030269"]}],"id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Manuf"],"published-print":{"date-parts":[[2024,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>As an essential scheduling problem with several practical applications, the parallel machine scheduling problem (PMSP) with family setups constraints is difficult to solve and proven to be NP-hard. To this end, we present a deep reinforcement learning (DRL) approach to solve a PMSP considering family setups, aiming at minimizing the total tardiness. The PMSP is first modeled as a Markov decision process, where we design a novel variable-length representation of states and actions, so that the DRL agent can calculate a comprehensive priority for each job at each decision time point and then select the next job directly according to these priorities. Meanwhile, the variable-length state matrix and action vector enable the trained agent to solve instances of any scales. To handle the variable-length sequence and simultaneously ensure the calculated priority is a global priority among all jobs, we employ a recurrent neural network, particular gated recurrent unit, to approximate the policy of the agent. The agent is trained based on Proximal Policy Optimization algorithm. Moreover, we develop a two-stage training strategy to enhance the training efficiency. In the numerical experiments, we first train the agent on a given instance and then employ it to solve instances with much larger scales. The experimental results demonstrate the strong generalization capability of the trained agent and the comparison with three dispatching rules and two metaheuristics further validates the superiority of this agent.<\/jats:p>","DOI":"10.1007\/s10845-023-02094-4","type":"journal-article","created":{"date-parts":[[2023,3,26]],"date-time":"2023-03-26T20:50:32Z","timestamp":1679863832000},"page":"1107-1140","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":38,"title":["A two-stage RNN-based deep reinforcement learning approach for solving the parallel machine scheduling problem with due dates and family setups"],"prefix":"10.1007","volume":"35","author":[{"given":"Funing","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3397-1551","authenticated-orcid":false,"given":"Sebastian","family":"Lang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bingyuan","family":"Hong","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tobias","family":"Reggelin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,3,9]]},"reference":[{"key":"2094_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2020.105162","volume":"128","author":"V Abu-Marrul","year":"2021","unstructured":"Abu-Marrul, V., Martinelli, R., Hamacher, S., & Gribkovskaia, I. (2021). Matheuristics for a parallel machine scheduling problem with nonanticipatory family setup times: Application in the offshore oil and gas industry. Computers & Operations Research, 128, 105162.","journal-title":"Computers & Operations Research"},{"issue":"2","key":"2094_CR2","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1007\/s10845-015-1117-6","volume":"29","author":"M Afzalirad","year":"2018","unstructured":"Afzalirad, M., & Shafipour, M. (2018). Design of an efficient genetic algorithm for resource-constrained unrelated parallel machine scheduling problem with machine eligibility restrictions. Journal of Intelligent Manufacturing, 29(2), 423\u2013437.","journal-title":"Journal of Intelligent Manufacturing"},{"issue":"11","key":"2094_CR3","doi-asserted-by":"publisher","first-page":"3471","DOI":"10.1016\/j.cor.2006.02.009","volume":"34","author":"D Anghinolfi","year":"2007","unstructured":"Anghinolfi, D., & Paolucci, M. (2007). Parallel machine total tardiness scheduling with a new hybrid metaheuristic approach. Computers & Operations Research, 34(11), 3471\u20133490.","journal-title":"Computers & Operations Research"},{"issue":"5","key":"2094_CR4","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1023\/A:1008918229511","volume":"11","author":"V Armentano","year":"2000","unstructured":"Armentano, V., Yamashita, D. S., et al. (2000). Tabu search for scheduling on identical parallel machines to minimize mean tardiness. Journal of Intelligent Manufacturing, 11(5), 453\u2013460.","journal-title":"Journal of Intelligent Manufacturing"},{"issue":"9","key":"2094_CR5","doi-asserted-by":"publisher","first-page":"1705","DOI":"10.1007\/s00170-014-6390-6","volume":"76","author":"O Avalos-Rosales","year":"2015","unstructured":"Avalos-Rosales, O., Angel-Bello, F., & Alvarez, A. (2015). Efficient metaheuristic algorithm and re-formulations for the unrelated parallel machine scheduling problem with sequence and machine-dependent setup times. The International Journal of Advanced Manufacturing Technology, 76(9), 1705\u20131718.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"issue":"2","key":"2094_CR6","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1016\/S0925-5273(98)00034-6","volume":"55","author":"M Azizoglu","year":"1998","unstructured":"Azizoglu, M., & Kirca, O. (1998). Tardiness minimization on parallel machines. International Journal of Production Economics, 55(2), 163\u2013168.","journal-title":"International Journal of Production Economics"},{"key":"2094_CR7","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1016\/j.cie.2019.03.051","volume":"131","author":"S B\u00e1ez","year":"2019","unstructured":"B\u00e1ez, S., Angel-Bello, F., Alvarez, A., & Meli\u00e1n-Batista, B. (2019). A hybrid metaheuristic algorithm for a parallel machine scheduling problem with dependent setup times. Computers & Industrial Engineering, 131, 295\u2013305.","journal-title":"Computers & Industrial Engineering"},{"issue":"6","key":"2094_CR8","doi-asserted-by":"publisher","first-page":"6814","DOI":"10.1016\/j.eswa.2010.12.064","volume":"38","author":"S Balin","year":"2011","unstructured":"Balin, S. (2011). Non-identical parallel machine scheduling using genetic algorithm. Expert Systems with Applications, 38(6), 6814\u20136821.","journal-title":"Expert Systems with Applications"},{"key":"2094_CR9","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009). Curriculum learning. Proceedings of the 26th annual international conference on machine learning, (pp. 41\u201348).","DOI":"10.1145\/1553374.1553380"},{"issue":"2","key":"2094_CR10","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1109\/72.279181","volume":"5","author":"Y Bengio","year":"1994","unstructured":"Bengio, Y., Simard, P., & Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 5(2), 157\u2013166.","journal-title":"IEEE Transactions on Neural Networks"},{"issue":"1","key":"2094_CR11","doi-asserted-by":"publisher","first-page":"134","DOI":"10.1016\/j.ijpe.2008.04.011","volume":"115","author":"D Biskup","year":"2008","unstructured":"Biskup, D., Herrmann, J., & Gupta, J. N. D. (2008). Scheduling identical parallel machines to minimize total tardiness. International Journal of Production Economics, 115(1), 134\u2013142.","journal-title":"International Journal of Production Economics"},{"key":"2094_CR12","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). Openai gym. arXiv:1606.01540."},{"key":"2094_CR13","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Bahdanau, D., & Bengio, Y. (2014). On the properties of neural machine translation: Encoder\u2013decoder approaches. arXiv:1409.1259.","DOI":"10.3115\/v1\/W14-4012"},{"issue":"7","key":"2094_CR14","doi-asserted-by":"publisher","first-page":"1087","DOI":"10.1016\/S0305-0548(02)00059-X","volume":"30","author":"JK Cochran","year":"2003","unstructured":"Cochran, J. K., Horng, S.-M., & Fowler, J. W. (2003). A multi-population genetic algorithm to solve multi-objective scheduling problems for parallel machines. Computers & Operations Research, 30(7), 1087\u20131102.","journal-title":"Computers & Operations Research"},{"issue":"2","key":"2094_CR15","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1207\/s15516709cog1402_1","volume":"14","author":"JL Elman","year":"1990","unstructured":"Elman, J. L. (1990). Finding structure in time. Cognitive Science, 14(2), 179\u2013211.","journal-title":"Cognitive Science"},{"issue":"1","key":"2094_CR16","doi-asserted-by":"publisher","first-page":"224","DOI":"10.1016\/j.cie.2012.10.002","volume":"64","author":"K-T Fang","year":"2013","unstructured":"Fang, K.-T., & Lin, B. M. (2013). Parallel-machine scheduling to minimize tardiness penalty and power cost. Computers & Industrial Engineering, 64(1), 224\u2013234.","journal-title":"Computers & Industrial Engineering"},{"issue":"8","key":"2094_CR17","doi-asserted-by":"publisher","first-page":"166","DOI":"10.1287\/mnsc.11.8.B166","volume":"11","author":"JW Gavett","year":"1965","unstructured":"Gavett, J. W. (1965). Three heuristic rules for sequencing jobs to a single production facility. Management Science, 11(8), 166\u2013176.","journal-title":"Management Science"},{"key":"2094_CR18","doi-asserted-by":"publisher","first-page":"287","DOI":"10.1016\/S0167-5060(08)70356-X","volume":"5","author":"RL Graham","year":"1979","unstructured":"Graham, R. L., Lawler, E. L., Lenstra, J. K., & Kan, A. R. (1979). Optimization and approximation in deterministic sequencing and scheduling: A survey. Annals of Discrete Mathematics, 5, 287\u2013326.","journal-title":"Annals of Discrete Mathematics"},{"key":"2094_CR19","doi-asserted-by":"crossref","unstructured":"Guo, L., Zhuang, Z., Huang, Z., & Qin, W. (2020). Optimization of dynamic multi-objective non-identical parallel machine scheduling with multistage reinforcement learning. 2020 IEEE 16th international conference on automation science and engineering (CASE), (pp. 1215\u20131219).","DOI":"10.1109\/CASE48305.2020.9216743"},{"key":"2094_CR20","first-page":"1","volume":"34","author":"BM Kayhan","year":"2021","unstructured":"Kayhan, B. M., & Yildiz, G. (2021). Reinforcement learning applications to machine scheduling problems: A comprehensive literature review. Journal of Intelligent Manufacturing, 34, 1\u201325.","journal-title":"Journal of Intelligent Manufacturing"},{"issue":"2","key":"2094_CR21","doi-asserted-by":"publisher","first-page":"246","DOI":"10.1109\/TSM.2010.2045666","volume":"23","author":"Y-D Kim","year":"2010","unstructured":"Kim, Y.-D., Joo, B.-J., & Choi, S.-Y. (2010). Scheduling wafer lots on diffusion machines in a semiconductor wafer fabrication facility. IEEE Transactions on Semiconductor Manufacturing, 23(2), 246\u2013254.","journal-title":"IEEE Transactions on Semiconductor Manufacturing"},{"key":"2094_CR22","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv:1412.6980 ."},{"key":"2094_CR23","first-page":"3057","volume":"2020","author":"S Lang","year":"2020","unstructured":"Lang, S., Behrendt, F., Lanzerath, N., Reggelin, T., & M\u00fcller, M. (2020). Integration of deep reinforcement learning and discrete-event simulation for real-time scheduling of a flexible job shop production. Winter Simulation Conference (WSC), 2020, 3057\u20133068.","journal-title":"Winter Simulation Conference (WSC)"},{"issue":"1","key":"2094_CR24","doi-asserted-by":"publisher","first-page":"793","DOI":"10.1016\/j.ifacol.2021.08.093","volume":"54","author":"S Lang","year":"2021","unstructured":"Lang, S., Kuetgens, M., Reichardt, P., & Reggelin, T. (2021). Modeling production scheduling problems as reinforcement learning environments based on discrete-event simulation and openai gym. IFAC-PapersOnLine, 54(1), 793\u2013798.","journal-title":"IFAC-PapersOnLine"},{"issue":"5","key":"2094_CR25","doi-asserted-by":"publisher","first-page":"773","DOI":"10.1007\/s00170-009-2203-8","volume":"47","author":"Z-J Lee","year":"2010","unstructured":"Lee, Z.-J., Lin, S.-W., & Ying, K.-C. (2010). Scheduling jobs on dynamic parallel machines with sequence-dependent setup times. The International Journal of Advanced Manufacturing Technology, 47(5), 773\u2013781.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"key":"2094_CR26","doi-asserted-by":"publisher","first-page":"71752","DOI":"10.1109\/ACCESS.2020.2987820","volume":"8","author":"C-L Liu","year":"2020","unstructured":"Liu, C.-L., Chang, C.-C., & Tseng, C.-J. (2020). Actor-critic deep reinforcement learning for solving job shop scheduling problems. IEEE Access, 8, 71752\u201371762.","journal-title":"IEEE Access"},{"key":"2094_CR27","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2020.106208","volume":"91","author":"S Luo","year":"2020","unstructured":"Luo, S. (2020). Dynamic scheduling for flexible job shop with new job insertions by deep reinforcement learning. Applied Soft Computing, 91, 106208.","journal-title":"Applied Soft Computing"},{"key":"2094_CR28","doi-asserted-by":"publisher","first-page":"101390","DOI":"10.1109\/ACCESS.2021.3097254","volume":"9","author":"B Paeng","year":"2021","unstructured":"Paeng, B., Park, I.-B., & Park, J. (2021). Deep reinforcement learning for minimizing tardiness in parallel machine scheduling with sequence dependent family setups. IEEE Access, 9, 101390\u2013101401.","journal-title":"IEEE Access"},{"issue":"20","key":"2094_CR29","doi-asserted-by":"publisher","first-page":"5823","DOI":"10.1080\/00207543.2011.629634","volume":"50","author":"CW Pickardt","year":"2012","unstructured":"Pickardt, C. W., & Branke, J. (2012). Setup-oriented dispatching rules-a survey. International Journal of Production Research, 50(20), 5823\u20135842.","journal-title":"International Journal of Production Research"},{"issue":"2","key":"2094_CR30","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1287\/opre.33.2.363","volume":"33","author":"CN Potts","year":"1985","unstructured":"Potts, C. N., & Van Wassenhove, L. N. (1985). A branch and bound algorithm for the total weighted tardiness problem. Operations Research, 33(2), 363\u2013377.","journal-title":"Operations Research"},{"issue":"1","key":"2094_CR31","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1016\/S0377-2217(98)00023-X","volume":"116","author":"C Rajendran","year":"1999","unstructured":"Rajendran, C., & Holthaus, O. (1999). A comparative study of dispatching rules in dynamic flowshops and jobshops. European Journal of Operational Research, 116(1), 156\u2013170.","journal-title":"European Journal of Operational Research"},{"key":"2094_CR32","doi-asserted-by":"publisher","DOI":"10.1016\/j.rcim.2022.102406","volume":"78","author":"MLR Rodr\u00edguez","year":"2022","unstructured":"Rodr\u00edguez, M. L. R., Kubler, S., de Giorgio, A., Cordy, M., Robert, J., & Le Traon, Y. (2022). Multi-agent deep reinforcement learning based predictive maintenance on parallel machines. Robotics and Computer-Integrated Manufacturing, 78, 102406.","journal-title":"Robotics and Computer-Integrated Manufacturing"},{"key":"2094_CR33","doi-asserted-by":"publisher","first-page":"442","DOI":"10.1016\/j.promfg.2020.02.051","volume":"42","author":"B Rolf","year":"2020","unstructured":"Rolf, B., Reggelin, T., Nahhas, A., Lang, S., & M\u00fcller, M. (2020). Assigning dispatching rules using a genetic algorithm to solve a hybrid flow shop scheduling problem. Procedia Manufacturing, 42, 442\u2013449.","journal-title":"Procedia Manufacturing"},{"key":"2094_CR34","doi-asserted-by":"publisher","first-page":"274","DOI":"10.1016\/j.cie.2014.04.001","volume":"72","author":"JE Schaller","year":"2014","unstructured":"Schaller, J. E. (2014). Minimizing total tardiness for scheduling identical parallel machines with family setups. Computers & Industrial Engineering, 72, 274\u2013281.","journal-title":"Computers & Industrial Engineering"},{"key":"2094_CR35","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015). Trust region policy optimization. International conference on machine learning, (pp. 1889\u20131897)."},{"key":"2094_CR36","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv:1707.06347 ."},{"issue":"20","key":"2094_CR37","doi-asserted-by":"publisher","first-page":"4235","DOI":"10.1080\/00207540410001708461","volume":"42","author":"HJ Shin","year":"2004","unstructured":"Shin, H. J., & Leon, V. J. (2004). Scheduling with product family set-up times: An application in TFT LCD manufacturing. International Journal of Production Research, 42(20), 4235\u20134248.","journal-title":"International Journal of Production Research"},{"key":"2094_CR38","unstructured":"Sigtia, S., Benetos, E., Cherla, S., Weyde, T., Garcez, A., & Dixon, S. (2014). RNN-based music language models for improving automatic music transcription. Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR), (pp. 53\u201358)."},{"issue":"7676","key":"2094_CR39","doi-asserted-by":"publisher","first-page":"354","DOI":"10.1038\/nature24270","volume":"550","author":"D Silver","year":"2017","unstructured":"Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., et al. (2017). Mastering the game of go without human knowledge. Nature, 550(7676), 354\u2013359.","journal-title":"Nature"},{"issue":"8","key":"2094_CR40","doi-asserted-by":"publisher","first-page":"3668","DOI":"10.1109\/TCYB.2019.2950779","volume":"50","author":"S Sun","year":"2019","unstructured":"Sun, S., Cao, Z., Zhu, H., & Zhao, J. (2019). A survey of optimization methods from a machine learning perspective. IEEE Transactions on Cybernetics, 50(8), 3668\u20133681.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"2094_CR41","volume-title":"Reinforcement learning: An introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. MIT Press."},{"key":"2094_CR42","unstructured":"Tassel, P., Gebser, M., & Schekotihin, K. (2021). A reinforcement learning environment for job-shop scheduling. arXiv:2104.03760."},{"issue":"27","key":"2094_CR43","doi-asserted-by":"publisher","first-page":"767","DOI":"10.21105\/joss.00767","volume":"3","author":"R van der Ham","year":"2018","unstructured":"van der Ham, R. (2018). salabim: Discrete event simulation and animation in python. Journal of Open Source Software, 3(27), 767.","journal-title":"Journal of Open Source Software"},{"issue":"19","key":"2094_CR44","doi-asserted-by":"publisher","first-page":"5837","DOI":"10.1080\/00207543.2015.1011289","volume":"53","author":"D-J van der Zee","year":"2015","unstructured":"van der Zee, D.-J. (2015). Family-based dispatching with parallel machines. International Journal of Production Research, 53(19), 5837\u20135856.","journal-title":"International Journal of Production Research"},{"issue":"7782","key":"2094_CR45","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","volume":"575","author":"O Vinyals","year":"2019","unstructured":"Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., et al. (2019). Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature, 575(7782), 350\u2013354.","journal-title":"Nature"},{"issue":"4","key":"2094_CR46","doi-asserted-by":"publisher","first-page":"257","DOI":"10.23919\/CSMS.2021.0027","volume":"1","author":"L Wang","year":"2021","unstructured":"Wang, L., Pan, Z., & Wang, J. (2021). A review of reinforcement learning based intelligent optimization for manufacturing scheduling. Complex System Modeling and Simulation, 1(4), 257\u2013270.","journal-title":"Complex System Modeling and Simulation"},{"issue":"10","key":"2094_CR47","doi-asserted-by":"publisher","first-page":"1550","DOI":"10.1109\/5.58337","volume":"78","author":"PJ Werbos","year":"1990","unstructured":"Werbos, P. J. (1990). Backpropagation through time: What it does and how to do it. Proceedings of the IEEE, 78(10), 1550\u20131560.","journal-title":"Proceedings of the IEEE"},{"key":"2094_CR48","doi-asserted-by":"crossref","unstructured":"Wilbrecht, J. K., & Prescott, W. B. (1969). The influence of setup time on job shop performance. Management Science, 16(4), 274\u2013280.","DOI":"10.1287\/mnsc.16.4.B274"},{"key":"2094_CR49","unstructured":"Wu, Y., & Tian, Y. (2016). Training agent for first-person shooter game with actor-critic curriculum learning."},{"key":"2094_CR50","unstructured":"Yin, W., Kann, K., Yu, M., & Sch\u00fctze, H. (2017). Comparative study of cnn and rnn for natural language processing. arXiv:1702.01923."},{"issue":"4","key":"2094_CR51","doi-asserted-by":"publisher","first-page":"2848","DOI":"10.1016\/j.eswa.2009.09.006","volume":"37","author":"K-C Ying","year":"2010","unstructured":"Ying, K.-C., & Cheng, H.-M. (2010). Dynamic parallel machine scheduling with sequence-dependent setup times using an iterated greedy heuristic. Expert Systems with Applications, 37(4), 2848\u20132852.","journal-title":"Expert Systems with Applications"},{"issue":"2","key":"2094_CR52","doi-asserted-by":"publisher","first-page":"94","DOI":"10.1504\/IJSOI.2016.080083","volume":"8","author":"B Yuan","year":"2016","unstructured":"Yuan, B., Jiang, Z., & Wang, L. (2016). Dynamic parallel machine scheduling with random breakdowns using the learning agent. International Journal of Services Operations and Informatics, 8(2), 94\u2013103.","journal-title":"International Journal of Services Operations and Informatics"},{"key":"2094_CR53","first-page":"1565","volume":"2013","author":"B Yuan","year":"2013","unstructured":"Yuan, B., Wang, L., & Jiang, Z. (2013). Dynamic parallel machine scheduling using the learning agent. IEEE International Conference on Industrial Engineering and Engineering management, 2013, 1565\u20131569.","journal-title":"IEEE International Conference on Industrial Engineering and Engineering management"},{"issue":"9","key":"2094_CR54","doi-asserted-by":"publisher","first-page":"1487","DOI":"10.1007\/s00170-015-7215-y","volume":"81","author":"JR Zeidi","year":"2015","unstructured":"Zeidi, J. R., & MohammadHosseini, S. (2015). Scheduling unrelated parallel machines with sequence-dependent setup times. The International Journal of Advanced Manufacturing Technology, 81(9), 1487\u20131496.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"issue":"1","key":"2094_CR55","doi-asserted-by":"publisher","first-page":"542","DOI":"10.1109\/TITS.2020.3002271","volume":"22","author":"C Zhang","year":"2020","unstructured":"Zhang, C., Liu, Y., Wu, F., Tang, B., & Fan, W. (2020). Effective charging planning based on deep reinforcement learning for electric vehicles. IEEE Transactions on Intelligent Transportation Systems, 22(1), 542\u2013554.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"issue":"2","key":"2094_CR56","doi-asserted-by":"publisher","first-page":"446","DOI":"10.1016\/j.ejor.2011.05.052","volume":"215","author":"Z Zhang","year":"2011","unstructured":"Zhang, Z., Zheng, L., Hou, F., & Li, N. (2011). Semiconductor final test scheduling with sarsa ($$\\lambda $$, k) algorithm. European Journal of Operational Research, 215(2), 446\u2013458.","journal-title":"European Journal of Operational Research"},{"issue":"7","key":"2094_CR57","doi-asserted-by":"publisher","first-page":"1315","DOI":"10.1016\/j.cor.2011.07.019","volume":"39","author":"Z Zhang","year":"2012","unstructured":"Zhang, Z., Zheng, L., Li, N., Wang, W., Zhong, S., & Hu, K. (2012). Minimizing mean weighted tardiness in unrelated parallel machine scheduling with reinforcement learning. Computers & Operations Research, 39(7), 1315\u20131324.","journal-title":"Computers & Operations Research"},{"issue":"9","key":"2094_CR58","doi-asserted-by":"publisher","first-page":"968","DOI":"10.1007\/s00170-006-0662-8","volume":"34","author":"Z Zhang","year":"2007","unstructured":"Zhang, Z., Zheng, L., & Weng, M. X. (2007). Dynamic parallel machine scheduling with mean weighted tardiness objective by q-learning. The International Journal of Advanced Manufacturing Technology, 34(9), 968\u2013980.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"key":"2094_CR59","doi-asserted-by":"crossref","unstructured":"Zhou, D., Jia, R., & Yao, H. (2021). Robotic arm motion planning based on curriculum reinforcement learning. 2021 6th International Conference on Control and Robotics Engineering (ICCRE), (pp. 44\u201349).","DOI":"10.1109\/ICCRE51898.2021.9435700"},{"key":"2094_CR60","doi-asserted-by":"publisher","first-page":"383","DOI":"10.1016\/j.procir.2020.05.163","volume":"93","author":"L Zhou","year":"2020","unstructured":"Zhou, L., Zhang, L., & Horn, B. K. (2020). Deep reinforcement learning-based dynamic scheduling in smart manufacturing. Procedia CIRP, 93, 383\u2013388.","journal-title":"Procedia CIRP"}],"container-title":["Journal of Intelligent Manufacturing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10845-023-02094-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10845-023-02094-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10845-023-02094-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,29]],"date-time":"2024-02-29T04:44:27Z","timestamp":1709181867000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10845-023-02094-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,9]]},"references-count":60,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,3]]}},"alternative-id":["2094"],"URL":"https:\/\/doi.org\/10.1007\/s10845-023-02094-4","relation":{},"ISSN":["0956-5515","1572-8145"],"issn-type":[{"value":"0956-5515","type":"print"},{"value":"1572-8145","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,9]]},"assertion":[{"value":"11 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 February 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 March 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}