{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,1]],"date-time":"2026-08-01T09:39:19Z","timestamp":1785577159310,"version":"3.56.0"},"reference-count":56,"publisher":"American Institute of Aeronautics and Astronautics (AIAA)","issue":"3","license":[{"start":{"date-parts":[[2027,1,4]],"date-time":"2027-01-04T00:00:00Z","timestamp":1799020800000},"content-version":"am","delay-in-days":309,"URL":"https:\/\/www.aiaa.org\/userlicenses\/1.0\/#CompEndUserLicense"}],"funder":[{"DOI":"10.13039\/100006602","name":"Air Force Research Laboratory","doi-asserted-by":"publisher","award":["FA9453-22-2-0050"],"award-info":[{"award-number":["FA9453-22-2-0050"]}],"id":[{"id":"10.13039\/100006602","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["arc.aiaa.org"],"crossmark-restriction":true},"short-container-title":["Journal of Aerospace Information Systems"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:p>Recent research is developing autonomous policies for agile Earth-observing satellite (AEOS) tasking using deep reinforcement learning (DRL). Prior work uses a fixed training environment with enhanced safety challenges to allow DRL to train a policy over a short number of orbits. This paper investigates how training environment enhancements can improve these policies\u2019 robustness by varying the training environment and extending simulated mission times. DRL enables real-time, onboard decision-making while accounting for complex dynamics and constraints. However, policies often struggle to generalize to longer missions or conditions different from training, partly due to the short simulated mission times. Multi-spacecraft systems amplify these challenges, where battery capacity and solar panel efficiency diverge over time due to external factors. Although satellites might run copies of a single policy, its robustness is essential to accommodate agent variability and ensure mission success. This paper investigates training with two curricula, constantly degraded environments, and a domain randomization approach to improve AEOS policy robustness and performance. Policies are tested on episodes 125-times longer than training episodes under nominal and degraded conditions. The proposed training environment enhancements\u2019 applicability is demonstrated in a heterogeneous spacecraft system, achieving a 63% reduction in failures and a 2.8% increase in median cumulative reward.<\/jats:p>","DOI":"10.2514\/1.i011697","type":"journal-article","created":{"date-parts":[[2026,1,4]],"date-time":"2026-01-04T13:22:28Z","timestamp":1767532948000},"page":"293-304","update-policy":"https:\/\/doi.org\/10.2514\/aiaa_crossmarkpolicy","source":"Crossref","is-referenced-by-count":2,"title":["Improving Robustness of Autonomous Earth-Observing Spacecraft Using Training Environment Enhancements"],"prefix":"10.2514","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7244-7491","authenticated-orcid":false,"given":"Lorenzzo Quevedo","family":"Mantovani","sequence":"first","affiliation":[{"name":"University of Colorado Boulder"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0002-6035","authenticated-orcid":false,"given":"Hanspeter","family":"Schaub","sequence":"additional","affiliation":[{"name":"University of Colorado Boulder"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1387","reference":[{"key":"r1","doi-asserted-by":"publisher","DOI":"10.2514\/1.I011649"},{"key":"r2","doi-asserted-by":"publisher","DOI":"10.2514\/1.I010754"},{"key":"r3","doi-asserted-by":"publisher","DOI":"10.2514\/1.A35169"},{"issue":"5","key":"r4","first-page":"5235","volume":"59","author":"Herrmann A.","year":"2023","journal-title":"IEEE Transactions on Aerospace and Electronic Systems"},{"key":"r5","doi-asserted-by":"publisher","DOI":"10.3390\/math11194059"},{"key":"r6","doi-asserted-by":"publisher","DOI":"10.2514\/1.A35736"},{"key":"r7","doi-asserted-by":"publisher","DOI":"10.3390\/make4010013"},{"key":"r8","doi-asserted-by":"publisher","DOI":"10.1016\/S1270-9638(02)01173-2"},{"key":"r9","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2005.12.026"},{"key":"r10","doi-asserted-by":"publisher","DOI":"10.2514\/1.A34931"},{"key":"r11","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.2019.2947978"},{"key":"r12","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2021.107292"},{"key":"r13","doi-asserted-by":"publisher","DOI":"10.2514\/1.A36143"},{"key":"r14","doi-asserted-by":"publisher","DOI":"10.2514\/6.2025-1148"},{"key":"r15","unstructured":"HarrisA.TeilT.SchaubH. \u201cSpacecraft Decision-Making Autonomy Using Deep Reinforcement Learning,\u201d AAS Spaceflight Mechanics Meeting, AAS Paper 19-447, Springfield, VA, 2019."},{"key":"r16","unstructured":"HarrisA.SchaubH. \u201cDeep On-Board Scheduling for Autonomous Attitude Guidance Operations,\u201d AAS Guidance, Navigation and Control Conference, AAS Paper 12-117, Springfield, VA, 2020."},{"key":"r17","doi-asserted-by":"publisher","DOI":"10.2514\/1.I010992"},{"key":"r18","doi-asserted-by":"publisher","DOI":"10.3389\/frspt.2023.1263489"},{"key":"r19","doi-asserted-by":"publisher","DOI":"10.2514\/1.I011209"},{"key":"r20","unstructured":"StephensonM.MantovaniL. Q.SchaubH. \u201cIntent Sharing For Emergent Collaboration In Autonomous Earth Observing Constellations,\u201d AAS Guidance and Control Conference, AAS Paper 24-192, Springfield, VA, 2024."},{"key":"r21","unstructured":"Hadj-SalahA.VerdierR.CaronC.PicardM.CapelleM. \u201cSchedule Earth Observation Satellites with Deep Reinforcement Learning,\u201d Proceedings of the 11th International Workshop on Planning and Scheduling for Space (IWPSS), co-located with the International Conference on Automated Planning and Scheduling, edited by ChienS., Berkeley, CA, 2019, pp.\u00a061\u201364."},{"key":"r22","doi-asserted-by":"crossref","unstructured":"NaikK.ChangO.KotulakC. \u201cDeep Reinforcement Learning for Autonomous Satellite Responsiveness to Observed Events,\u201d 2024 IEEE Aerospace Conference, Inst. of Electrical and Electronics Engineers, New York, 2024, pp.\u00a01\u201310. 10.1109\/AERO58975.2024.10521008","DOI":"10.1109\/AERO58975.2024.10521008"},{"key":"r23","unstructured":"MantovaniL. Q.NaganoY.SchaubH. \u201cReinforcement Learning for Satellite Autonomy Under Different Cloud Coverage Probability Observations,\u201d AAS Astrodynamics Specialist Conference, AAS Paper 24-189, Springfield, VA, 2024."},{"key":"r24","doi-asserted-by":"publisher","DOI":"10.1016\/j.swevo.2025.101857"},{"key":"r25","doi-asserted-by":"crossref","unstructured":"NazmyI.HarrisA.LahijanianM.SchaubH. \u201cShielded Deep Reinforcement Learning for Multi-Sensor Spacecraft Imaging,\u201d 2022 American Control Conference (ACC), IEEE Publ., Piscataway, NJ, 2022, pp.\u00a01808\u20131813. 10.23919\/ACC53348.2022.9867762","DOI":"10.23919\/ACC53348.2022.9867762"},{"key":"r26","doi-asserted-by":"crossref","unstructured":"HerrmannA.CarneiroJ. V.SchaubH. \u201cReinforcement Learning For the Multi-Satellite Earth-Observing Scheduling Problem,\u201d Proceedings of the 44th Annual American Astronautical Society Guidance, Navigation, and Control Conference, 2022, edited by SandnasM.SpencerD. B., Springer International Publ., Cham, 2024, pp.\u00a01351\u20131368.","DOI":"10.1007\/978-3-031-51928-4_74"},{"key":"r27","volume-title":"AIAA Science and Technology Forum and Exposition (SciTech)","author":"Stephenson M. A.","year":"2024"},{"key":"r28","doi-asserted-by":"crossref","unstructured":"BengioY.LouradourJ.CollobertR.WestonJ. \u201cCurriculum Learning,\u201d Proceedings of the 26th Annual International Conference on Machine Learning, ACM, Montreal Quebec Canada, 2009, pp.\u00a041\u201348. 10.1145\/1553374.1553380","DOI":"10.1145\/1553374.1553380"},{"key":"r29","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01611-x"},{"key":"r30","unstructured":"FlorensaC.HeldD.WulfmeierM.ZhangM.AbbeelP. \u201cReverse Curriculum Generation for Reinforcement Learning,\u201d Proceedings of the 1st Annual Conference on Robot Learning, Proceedings of Machine Learning Research, edited by LevineS.VanhouckeV.GoldbergK., Vol.\u00a078, PMLR, Cambridge, MA, 2017, pp.\u00a0482\u2013495."},{"key":"r31","doi-asserted-by":"crossref","unstructured":"HermannL.ArgusM.EitelA.AmiranashviliA.BurgardW.BroxT. \u201cAdaptive Curriculum Generation from Demonstrations for Sim-to-Real Visuomotor Control,\u201d 2020 IEEE International Conference on Robotics and Automation (ICRA), Inst. of Electrical and Electronics Engineers, New York, 2020, pp.\u00a06498\u20136505. 10.1109\/ICRA40945.2020.9197108","DOI":"10.1109\/ICRA40945.2020.9197108"},{"key":"r32","unstructured":"RudinN.HoellerD.ReistP.HutterM. \u201cLearning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning,\u201d Proceedings of the 5th Conference on Robot Learning, Proceedings of Machine Learning Research, edited by FaustA.HsuD.NeumannG., Vol.\u00a0164, PMLR, Cambridge, MA, 2022, pp.\u00a091\u2013100."},{"key":"r33","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3107742"},{"key":"r34","doi-asserted-by":"publisher","DOI":"10.1177\/02783649231224053"},{"key":"r35","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3251193"},{"key":"r36","unstructured":"TurchettaM.KolobovA.ShahS.KrauseA.AgarwalA. \u201cSafe Reinforcement Learning via Curriculum Induction,\u201d Proceedings of the 34th International Conference on Neural Information Processing Systems, Curran Assoc., New York, 2020."},{"key":"r37","doi-asserted-by":"publisher","DOI":"10.1016\/j.actaastro.2022.08.047"},{"key":"r38","unstructured":"FedericiL.ZavoliA. \u201cA Curriculum Learning Approach for Improving Constraint Handling in Reinforcement Learning Applications to Spacecraft Guidance and Control,\u201d SPAICE2024: The First Joint European Space Agency\/IAA Conference on AI in and for Space, Zenodo, Genova, Switzerland, 2024, pp.\u00a0482\u2013486. 10.5281\/ZENODO.13885667"},{"key":"r39","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-024-10520-8"},{"issue":"1","key":"r40","first-page":"7382","volume":"21","author":"Narvekar S.","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"r41","doi-asserted-by":"crossref","unstructured":"SongY.SchneiderJ. \u201cRobust Reinforcement Learning via Genetic Curriculum,\u201d 2022 International Conference on Robotics and Automation (ICRA), Inst. of Electrical and Electronics Engineers, New York, 2022, pp.\u00a05560\u20135566. 10.1109\/ICRA46639.2022.9812420","DOI":"10.1109\/ICRA46639.2022.9812420"},{"key":"r42","doi-asserted-by":"crossref","unstructured":"LiY.TianY.TongE.NiuW.LiuJ. \u201cRobust Reinforcement Learning via Progressive Task Sequence,\u201d Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, International Joint Conferences on Artificial Intelligence Organization, 2023, pp.\u00a0455\u2013463. 10.24963\/ijcai.2023\/51","DOI":"10.24963\/ijcai.2023\/51"},{"key":"r43","doi-asserted-by":"crossref","unstructured":"PengX. B.AndrychowiczM.ZarembaW.AbbeelP. \u201cSim-to-Real Transfer of Robotic Control with Dynamics Randomization,\u201d 2018 IEEE International Conference on Robotics and Automation (ICRA), Inst. of Electrical and Electronics Engineers, New York, 2018, pp.\u00a03803\u20133810. 10.1109\/ICRA.2018.8460528","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"r44","doi-asserted-by":"crossref","unstructured":"ChebotarY.HandaA.MakoviychukV.MacklinM.IssacJ.RatliffN.FoxD. \u201cClosing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience,\u201d 2019 International Conference on Robotics and Automation (ICRA), Inst. of Electrical and Electronics Engineers, New York, 2019, pp.\u00a08973\u20138979. 10.1109\/ICRA.2019.8793789","DOI":"10.1109\/ICRA.2019.8793789"},{"key":"r45","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2019.2942989"},{"key":"r46","series-title":"Adaptive Computation and Machine Learning Series","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"2018","edition":"2"},{"key":"r47","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(98)00023-X"},{"key":"r48","doi-asserted-by":"crossref","unstructured":"StephensonM.SchaubH. \u201cBSK-RL: Modular, High-Fidelity Reinforcement Learning Environments for Spacecraft Tasking,\u201d International Astronautical Congress, International Astronautical Federation, Paris, France, 2024.","DOI":"10.52202\/078372-0120"},{"key":"r49","doi-asserted-by":"publisher","DOI":"10.2514\/1.I010762"},{"issue":"1","key":"r50","first-page":"1","volume":"44","author":"Schaub H.","year":"1996","journal-title":"Journal of the Astronautical Sciences"},{"key":"r51","doi-asserted-by":"publisher","DOI":"10.2514\/4.861550"},{"key":"r52","unstructured":"LuoM.YaoJ.LiawR.LiangE.StoicaI. \u201cIMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks,\u201d Jan. 2020. 10.48550\/arXiv.1912.00167"},{"key":"r53","unstructured":"LiangE.LiawR.MoritzP.NishiharaR.FoxR.GoldbergK.GonzalezJ. E.JordanM. I.StoicaI. \u201cRLlib: Abstractions for Distributed Reinforcement Learning,\u201d June 2018. 10.48550\/arXiv.1712.09381"},{"key":"r54","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113515"},{"key":"r55","doi-asserted-by":"crossref","unstructured":"ReedR.SchaubH.LahijanianM. \u201cShielded Deep Reinforcement Learning for Complex Spacecraft Tasking,\u201d 2024 American Control Conference (ACC), Inst. of Electrical and Electronics Engineers, New York, 2024, pp.\u00a02331\u20132337. 10.23919\/ACC60939.2024.10644855","DOI":"10.23919\/ACC60939.2024.10644855"},{"key":"r56","unstructured":"MantovaniL. Q.SchaubH. \u201cPerformance Evaluation of Shielded Neural Networks for Autonomous Agile Earth Observing Satellites in Long Term Scenarios,\u201d Proceedings of the 14th International Workshop on Planning and Scheduling for Space (IWPSS), edited by ArtiguesC.JaubertJ.PraletC.ChienS., Toulouse, France, 2025, pp.\u00a0112\u2013121."}],"container-title":["Journal of Aerospace Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/arc.aiaa.org\/doi\/am-pdf\/10.2514\/1.I011697","content-type":"application\/pdf","content-version":"am","intended-application":"unspecified"},{"URL":"https:\/\/arc.aiaa.org\/doi\/pdf\/10.2514\/1.I011697","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/arc.aiaa.org\/doi\/pdf\/10.2514\/1.I011697","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,5]],"date-time":"2026-03-05T08:05:27Z","timestamp":1772697927000},"score":1,"resource":{"primary":{"URL":"https:\/\/arc.aiaa.org\/doi\/10.2514\/1.I011697"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3]]},"references-count":56,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["10.2514\/1.I011697"],"URL":"https:\/\/doi.org\/10.2514\/1.i011697","relation":{},"ISSN":["1940-3151","2327-3097"],"issn-type":[{"value":"1940-3151","type":"print"},{"value":"2327-3097","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3]]},"assertion":[{"value":"2025-05-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}