{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T18:25:37Z","timestamp":1763749537988,"version":"build-2065373602"},"reference-count":44,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T00:00:00Z","timestamp":1760918400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Economies"],"abstract":"<jats:p>Shimer documents that the search-and-matching model driven by productivity shocks explains only a small share of the observed volatility of unemployment and vacancies, which is known as the Shimer puzzle. We revisit this evidence by replacing the representative firm\u2019s optimization with a deep reinforcement learning (DRL) agent that learns its vacancy-posting policy through interaction in a Diamond\u2013Mortensen\u2013Pissarides (DMP) model. Comparing the learning economy with a conventional log-linearized DSGE solution under the same parameters, we find that while both frameworks preserve a downward-sloping Beveridge curve, learning-based economy produces much higher volatility in key labor market variables and returns to a steady state more slowly after shocks. These results point to bounded rationality and endogenous learning as mechanisms for labor market fluctuations and suggest that reinforcement learning can serve as a useful complement to standard macroeconomic analysis.<\/jats:p>","DOI":"10.3390\/economies13100302","type":"journal-article","created":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T12:27:29Z","timestamp":1760963249000},"page":"302","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Deep Reinforcement Learning in a Search-Matching Model of Labor Market Fluctuations"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-6028-6011","authenticated-orcid":false,"given":"Ruxin","family":"Chen","sequence":"first","affiliation":[{"name":"Department of Economics, Nagoya University, Aichi 464-8601, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,10,20]]},"reference":[{"key":"ref_1","unstructured":"Albrecht, S. V., Christianos, F., and Sch\u00e4fer, L. (2024). Multi-agent reinforcement learning: Foundations and modern approaches, MIT Press."},{"key":"ref_2","first-page":"353","article-title":"Designing economic agents that act like human agents: A behavioral approach to bounded rationality","volume":"81","author":"Arthur","year":"1991","journal-title":"The American Economic Review"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","article-title":"Deep reinforcement learning: A brief survey","volume":"34","author":"Arulkumaran","year":"2017","journal-title":"IEEE Signal Processing Magazine"},{"key":"ref_4","unstructured":"Ashwin, J., Beaudry, P., and Ellison, M. (2021). The unattractiveness of indeterminate dynamic equilibria, Centre for Economic Policy Research."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Atashbar, T., and Shi, R. A. (2022). Deep reinforcement learning: Emerging trends in macroeconomics and future prospects (IMF Working Papers No. 2022\/259), International Monetary Fund. Available online: https:\/\/ideas.repec.org\/p\/imf\/imfwpa\/2022-259.html.","DOI":"10.5089\/9798400224713.001"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"358","DOI":"10.1016\/S0166-4115(97)80105-7","article-title":"Reinforcement learning in artificial intelligence","volume":"121","author":"Barto","year":"1991","journal-title":"Advances in Psychology"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3267","DOI":"10.1257\/aer.20190623","article-title":"Artificial intelligence, algorithmic pricing, and collusion","volume":"110","author":"Calvano","year":"2020","journal-title":"American Economic Review"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3691326","article-title":"Estimating policy functions in payment systems using reinforcement learning","volume":"13","author":"Castro","year":"2025","journal-title":"ACM Transactions on Economics and Computation"},{"key":"ref_9","unstructured":"Chen, M., Joseph, A., Kumhof, M., Pan, X., Shi, R., and Zhou, X. (2021). Deep reinforcement learning in a monetary model. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chen, R., and Zhang, Z. (, January March). Deep reinforcement learning in labor market simulations. 2025 IEEE Symposium on Computational Intelligence for Financial Engineering and Economics (CiFer), Trondheim, Norway.","DOI":"10.1109\/CiFer64978.2025.10975741"},{"key":"ref_11","unstructured":"Covarrubias, M. (2025, September 01). Dynamic oligopoly and monetary policy: A deep reinforcement learning approach, Available online: https:\/\/drive.google.com\/file\/d\/1ivRIzPRzMr_Hlqrq6WsiQNlxxr1q0ZE_\/view."},{"key":"ref_12","unstructured":"Curry, M., Trott, A., Phade, S., Bai, Y., and Zheng, S. (2022). Finding general equilibria in many-agent economic simulations using deep reinforcement learning. arXiv."},{"key":"ref_13","first-page":"1","article-title":"Marginal jobs, heterogeneous firms, and unemployment flows","volume":"5","author":"Elsby","year":"2013","journal-title":"American Economic Journal: Macroeconomics"},{"key":"ref_14","unstructured":"Evans, B. P., and Ganesh, S. (, January May). Learning and calibrating heterogeneous bounded rational market behaviour with multi-agent reinforcement learning. 23rd International Conference on Autonomous Agents and Multiagent Systems, Auckland, New Zealand."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1086\/597302","article-title":"Unemployment fluctuations with staggered Nash wage bargaining","volume":"117","author":"Gertler","year":"2009","journal-title":"Journal of Political Economy"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1692","DOI":"10.1257\/aer.98.4.1692","article-title":"The cyclical behavior of equilibrium unemployment and vacancies revisited","volume":"98","author":"Hagedorn","year":"2008","journal-title":"American Economic Review"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1257\/0002828053828482","article-title":"Employment fluctuations with equilibrium wage stickiness","volume":"95","author":"Hall","year":"2005","journal-title":"American Economic Review"},{"key":"ref_18","unstructured":"Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (, January May). A very condensed survey and critique of multiagent deep reinforcement learning. 19th International Conference on Autonomous Agents and Multiagent Systems, Auckland, New Zealand."},{"key":"ref_19","unstructured":"Hill, E., Bardoscia, M., and Turrell, A. (2021). Solving heterogeneous general equilibrium economic models with deep reinforcement learning. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hinterlang, N., and T\u00e4nzer, A. (2021). Optimal monetary policy using reinforcement learning (Deutsche Bundesbank Discussion Paper), Deutsche Bundesbank.","DOI":"10.2139\/ssrn.4025682"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"279","DOI":"10.2307\/2297382","article-title":"On the efficiency of matching and related models of search and unemployment","volume":"57","author":"Hosios","year":"1990","journal-title":"The Review of Economic Studies"},{"key":"ref_22","unstructured":"Jirnyi, A., and Lepetyuk, V. (2025, September 01). A reinforcement learning approach to solving incomplete market models with aggregate uncertainty, Available online: https:\/\/papers.ssrn.com\/sol3\/papers.cfm?abstract_id=1832745."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"7684","DOI":"10.1109\/LRA.2022.3184795","article-title":"Multi-agent reinforcement learning for real-time dynamic production scheduling in a robot assembly cell","volume":"7","author":"Johnson","year":"2022","journal-title":"IEEE Robotics and Automation Letters"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"3030","DOI":"10.1257\/aer.20131702","article-title":"Efficient firm dynamics in a frictional labor market","volume":"105","author":"Kaas","year":"2015","journal-title":"American Economic Review"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1016\/j.red.2018.10.002","article-title":"Employment and hours over the business cycle in a model with search frictions","volume":"31","author":"Kudoh","year":"2019","journal-title":"Review of Economic Dynamics"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1146\/annurev-neuro-062111-150512","article-title":"Neural basis of reinforcement learning and decision making","volume":"35","author":"Lee","year":"2012","journal-title":"Annual Review of Neuroscience"},{"key":"ref_27","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (, January May). Continuous control with deep reinforcement learning. 4th International Conference on Learning Representations, San Juan, Puerto Rico."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1016\/j.jjie.2011.09.004","article-title":"Gross worker flows and unemployment dynamics in Japan","volume":"26","author":"Lin","year":"2012","journal-title":"Journal of the Japanese and International Economies"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1016\/j.jjie.2014.03.001","article-title":"An estimated search and matching model of the Japanese labor market","volume":"32","author":"Lin","year":"2014","journal-title":"Journal of the Japanese and International Economies"},{"key":"ref_30","first-page":"4768","article-title":"A unified approach to interpreting model predictions","volume":"30","author":"Lundberg","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1016\/j.jmoneco.2021.07.004","article-title":"Deep learning for solving dynamic economic models","volume":"122","author":"Maliar","year":"2021","journal-title":"Journal of Monetary Economics"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3616864","article-title":"Explainable reinforcement learning: A survey and comparative review","volume":"56","author":"Milani","year":"2024","journal-title":"ACM Computing Surveys"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1016\/j.japwor.2011.04.001","article-title":"Cyclical behavior of unemployment and job vacancies in Japan","volume":"23","author":"Miyamoto","year":"2011","journal-title":"Japan and the World Economy"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1016\/j.red.2007.01.004","article-title":"More on unemployment and vacancy fluctuations","volume":"10","author":"Mortensen","year":"2007","journal-title":"Review of Economic Dynamics"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1257\/jep.11.3.55","article-title":"Unemployment and labor market rigidities: Europe versus North America","volume":"11","author":"Nickell","year":"1997","journal-title":"Journal of Economic Perspectives"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"611","DOI":"10.3982\/QE452","article-title":"Solving the diamond\u2013mortensen\u2013pissarides model accurately","volume":"8","author":"Zhang","year":"2017","journal-title":"Quantitative Economics"},{"key":"ref_37","unstructured":"Pissarides, C. A. (2000). Equilibrium unemployment theory, The MIT Press. [2nd ed.]."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1339","DOI":"10.3982\/ECTA7562","article-title":"The unemployment volatility puzzle: Is wage stickiness the answer?","volume":"77","author":"Pissarides","year":"2009","journal-title":"Econometrica"},{"key":"ref_39","unstructured":"Puterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming, John Wiley & Sons."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Sargent, T. (1993). Bounded rationality in macroeconomics, Oxford University Press.","DOI":"10.1093\/oso\/9780198288640.001.0001"},{"key":"ref_41","unstructured":"Shi, R. A. (2023). Deep reinforcement learning and macroeconomic modelling. [Ph.D. thesis, University of Warwick]."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1257\/0002828053828572","article-title":"The cyclical behavior of equilibrium unemployment and vacancies","volume":"95","author":"Shimer","year":"2005","journal-title":"American Economic Review"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"456","DOI":"10.1006\/redy.1998.0056","article-title":"Search, concave production, and optimal firm size","volume":"2","author":"Smith","year":"1999","journal-title":"Review of Economic Dynamics"},{"key":"ref_44","unstructured":"Zhang, Z., and Chen, R. (2025). From individual learning to market equilibrium: Correcting structural and parametric biases in RL simulations of economic models. arXiv."}],"container-title":["Economies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2227-7099\/13\/10\/302\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T13:14:15Z","timestamp":1760966055000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2227-7099\/13\/10\/302"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,20]]},"references-count":44,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["economies13100302"],"URL":"https:\/\/doi.org\/10.3390\/economies13100302","relation":{},"ISSN":["2227-7099"],"issn-type":[{"value":"2227-7099","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,20]]}}}