{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T19:48:20Z","timestamp":1782762500646,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":79,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,25]]},"DOI":"10.1145\/3805689.3812365","type":"proceedings-article","created":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T16:20:39Z","timestamp":1782231639000},"page":"5095-5113","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Expanding External Access to Frontier AI Models for Dangerous Capability Evaluations"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-0352-1084","authenticated-orcid":false,"given":"Jacob","family":"Charnock","sequence":"first","affiliation":[{"name":"Cambridge, ERA, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6844-5765","authenticated-orcid":false,"given":"Alejandro","family":"Tlaie","sequence":"additional","affiliation":[{"name":"Pour Demain, Brussels, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8219-1617","authenticated-orcid":false,"given":"Kyle","family":"O'Brien","sequence":"additional","affiliation":[{"name":"Cambridge, ERA, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0084-1937","authenticated-orcid":false,"given":"Stephen","family":"Casper","sequence":"additional","affiliation":[{"name":"CSAIL, MIT, Cambridge, Massachussets, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0638-7094","authenticated-orcid":false,"given":"Aidan","family":"Homewood","sequence":"additional","affiliation":[{"name":"Centre for the Governance of AI, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Steven Adler. 2025. AI companies should be safety-testing the most capable versions of their models. https:\/\/stevenadler.substack.com\/p\/ai-companies-should-be-safety-testing"},{"key":"e_1_3_2_1_2_1","unstructured":"U.S. Food & Drug Administration. 2024. Types of FDA Inspections. Technical Report. U.S. Food & Drug Administration. https:\/\/www.fda.gov\/inspections-compliance-enforcement-and-criminal-investigations\/inspection-basics\/types-fda-inspections"},{"key":"e_1_3_2_1_3_1","unstructured":"Ahmed Ahmed Kevin Klyman Yi Zeng Sanmi Koyejo and Percy Liang. 2025. SpecEval: Evaluating Model Adherence to Behavior Specifications. arXiv:2509.02464. Retrieved from http:\/\/arxiv.org\/abs\/2509.02464."},{"key":"e_1_3_2_1_4_1","unstructured":"US AISI & UK AISI. 2024. US AISI and UK AISI Joint Pre-Deployment Test. https:\/\/cdn.prod.website-files.com\/663bd486c5e4c81588db7a1d\/6763fac97cd22a9484ac3c37_o1_uk_us_december_publication_final.pdf"},{"key":"e_1_3_2_1_5_1","unstructured":"Amazon. 2025. Amazon's Frontier Model Safety Framework. Technical Report. Amazon. https:\/\/assets.amazon.science\/a7\/7c\/8bdade5c4eda9168f3dee6434fff\/pc-amazon-frontier-model-safety-framework-2-7-final-2-9.pdf"},{"key":"e_1_3_2_1_6_1","unstructured":"Anthropic. 2024. Model Context Protocol. https:\/\/www.anthropic.com\/news\/model-context-protocol"},{"key":"e_1_3_2_1_7_1","unstructured":"Anthropic. 2024. Trust Center. https:\/\/trust.anthropic.com Accessed: 2024-01-15."},{"key":"e_1_3_2_1_8_1","unstructured":"Anthropic. 2025. Activating AI Safety Level 3 protections. Technical Report. Anthropic. https:\/\/www.anthropic.com\/news\/activating-asl3-protections"},{"key":"e_1_3_2_1_9_1","unstructured":"Anthropic. 2025. Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet. https:\/\/www.anthropic.com\/engineering\/swe-bench-sonnet"},{"key":"e_1_3_2_1_10_1","unstructured":"Anthropic. 2025. Strengthening our safeguards through collaboration with US CAISI and UK AISI. Technical Report. Anthropic. https:\/\/www.anthropic.com\/news\/strengthening-our-safeguards-through-collaboration-with-us-caisi-and-uk-aisi"},{"key":"e_1_3_2_1_11_1","volume-title":"System Card: Claude Opus 4 & Claude Sonnet 4. Technical Report. Anthropic. https:\/\/www-cdn.anthropic.com\/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf","year":"2025","unstructured":"Anthropic. 2025. System Card: Claude Opus 4 & Claude Sonnet 4. Technical Report. Anthropic. https:\/\/www-cdn.anthropic.com\/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf"},{"key":"e_1_3_2_1_12_1","volume-title":"System Card: Claude Sonnet 4.5. Technical Report. Anthropic. https:\/\/assets.anthropic.com\/m\/12f214efcc2f457a\/original\/Claude-Sonnet-4-5-System-Card.pdf","year":"2025","unstructured":"Anthropic. 2025. System Card: Claude Sonnet 4.5. Technical Report. Anthropic. https:\/\/assets.anthropic.com\/m\/12f214efcc2f457a\/original\/Claude-Sonnet-4-5-System-Card.pdf"},{"key":"e_1_3_2_1_13_1","unstructured":"Financial Conduct Authority. 2024. FCA Handbook REC 3.19. Technical Report. Financial Conduct Authority. https:\/\/handbook.fca.org.uk\/handbook\/rec3\/rec3s19?timeline=true"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1093\/cybsec\/tyab020"},{"key":"e_1_3_2_1_15_1","unstructured":"Stella Biderman Hailey Schoelkopf Lintang Sutawika Leo Gao Jonathan Tow Baber Abbasi Alham Fikri Aji Pawan Sasanka Ammanamanchi Sidney Black Jordan Clive Anthony DiPofi Julen Etxaniz Benjamin Fattori Jessica Zosa Forde Charles Foster Jeffrey Hsu Mimansa Jaiswal Wilson Y. Lee Haonan Li Charles Lovering Niklas Muennighoff Ellie Pavlick Jason Phang Aviya Skowron Samson Tan Xiangru Tang Kevin A. Wang Genta Indra Winata Fran\u00e7ois Yvon and Andy Zou. 2024. Lessons from the Trenches on Reproducible Evaluation of Language Models. arXiv:2405.14782. Retrieved from http:\/\/arxiv.org\/abs\/2405.14782."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Rishi Bommasani Sanjeev Arora Yejin Choi Li Fei-Fei Daniel E. Ho Dan Jurafsky Sanmi Koyejo Hima Lakkaraju Arvind Narayanan Alondra Nelson Emma Pierson Joelle Pineau Ga\u00ebl Varoquaux Suresh Venkatasubramanian Ion Stoica Percy Liang and Dawn Song. 2024. A Path for Science- and Evidence-based AI Policy. https:\/\/understanding-ai-safety.org","DOI":"10.1126\/science.adu8449"},{"key":"e_1_3_2_1_17_1","unstructured":"Rishi Bommasani Kevin Klyman Shayne Longpre Sayash Kapoor Nestor Maslej Betty Xiong Daniel Zhang and Percy Liang. 2023. The Foundation Model Transparency Index. arXiv:2310.12941. Retrieved from https:\/\/arxiv.org\/abs\/2310.12941."},{"key":"e_1_3_2_1_18_1","unstructured":"Dillon Bowen Ann-Kathrin Dombrowski Adam Gleave and Chris Cundy. 2025. AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations. arXiv:2503.17388. Retrieved from http:\/\/arxiv.org\/abs\/2503.17388."},{"key":"e_1_3_2_1_19_1","volume-title":"Osborne","author":"Bucknall Ben","year":"2025","unstructured":"Ben Bucknall, Robert Trager, and Michael A. Osborne. 2025. Position: Ensuring mutual privacy is necessary for effective external evaluation of proprietary AI systems. arXiv:2503.01470. Retrieved from https:\/\/arxiv.org\/abs\/2503.01470."},{"key":"e_1_3_2_1_20_1","volume-title":"Trager","author":"Bucknall Benjamin S.","year":"2023","unstructured":"Benjamin S. Bucknall and Robert F. Trager. 2023. Structured Access for Third-party Research on Frontier Ai Models: Investigating Researchers' Model Access Requirements. Technical Report. Oxford Martin School. https:\/\/oms-www.files.svdcdn.com\/production\/downloads\/academic\/Investigating_Researchers\u00e2\u0102\u0179_Model_Access_Oct23-compressed_3.pdf"},{"key":"e_1_3_2_1_21_1","volume-title":"Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tram\u00e8r.","author":"Carlini Nicholas","year":"2024","unstructured":"Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tram\u00e8r. 2024. Stealing Part of a Production Language Model. arXiv:2403.06634. Retrieved from https:\/\/arxiv.org\/abs\/2403.06634."},{"key":"e_1_3_2_1_22_1","unstructured":"Stephen Casper Xander Davies Claudia Shi Thomas Krendl Gilbert J\u00e9r\u00e9my Scheurer Javier Rando Rachel Freedman Tomasz Korbak David Lindner Pedro Freire Tony Wang Samuel Marks Charbel-Rapha\u00ebl Segerie Micah Carroll Andi Peng Phillip Christoffersen Mehul Damani Stewart Slocum Usman Anwar Anand Siththaranjan Max Nadeau Eric J. Michaud Jacob Pfau Dmitrii Krasheninnikov Xin Chen Lauro Langosco Peter Hase Erdem Biyik Anca Dragan David Krueger Dorsa Sadigh and Dylan Hadfield-Menell. 2023. Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv:2307.15217. Retrieved from http:\/\/arxiv.org\/abs\/2307.15217."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3659037"},{"key":"e_1_3_2_1_24_1","volume-title":"Yarin Gal, Furong Huang, and Dylan Hadfield-Menell.","author":"Che Zora","year":"2025","unstructured":"Zora Che, Stephen Casper, Robert Kirk, Anirudh Satheesh, Stewart Slocum, Lev Mckinney, Rohit Gandikota, Aidan Ewart, Domenic Rosati, Zichu Wu, Zikui Cai, Bilal Chughtai, Apollo Research, Yarin Gal, Furong Huang, and Dylan Hadfield-Menell. 2025. Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities \u201cWhat makes a bioweapons program effective?\u201d Input-space Attack Modifies input text Latent-space Attack Perturbs hidden neurons Weight-space Attack Fine-tunes the model Model Tampering Attacks. arXiv:2502.05209. Retrieved from http:\/\/arxiv.org\/abs\/2502.05209."},{"key":"e_1_3_2_1_25_1","unstructured":"European Commission. 2025. The General-Purpose AI Code of Practice. Technical Report. European Commission. https:\/\/digital-strategy.ec.europa.eu\/en\/policies\/contents-code-gpai#:~:text=EU%20copyright%20law.- Safety%20and%20Security -The%20Safety%20and"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Jesse Dodge Maarten Sap Ana Marasovi\u0107 William Agnew Gabriel Ilharco Dirk Groeneveld Margaret Mitchell and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758. Retrieved from http:\/\/arxiv.org\/abs\/2104.08758.","DOI":"10.18653\/v1\/2021.emnlp-main.98"},{"key":"e_1_3_2_1_27_1","volume-title":"AI Seoul Summit","author":"DSIT.","year":"2024","unstructured":"DSIT. 2024. Frontier AI Safety Commitments, AI Seoul Summit 2024. Technical Report. DSIT. https:\/\/www.gov.uk\/government\/publications\/frontier-ai-safety-commitments-ai-seoul-summit-2024\/frontier-ai-safety-commitments-ai-seoul-summit-2024"},{"key":"e_1_3_2_1_28_1","volume-title":"The Seventh International Conference on Cloud Computing, GRIDs, and Virtualization","author":"Duncan Bob","year":"2016","unstructured":"Bob Duncan and Mark Whittington. 2016. Enhancing Cloud Security and Privacy: The Power and the Weakness of the Audit Trail. In The Seventh International Conference on Cloud Computing, GRIDs, and Virtualization. International Academy, Research, and Industry Association, Rome, Italy, 125\u2013130. https:\/\/personales.upv.es\/thinkmind\/dl\/conferences\/cloudcomputing\/cloud_computing_2016\/cloud_computing_2016_6_20_20063.pdf"},{"key":"e_1_3_2_1_29_1","unstructured":"Frontier Model Forum. 2025. Third-Party Assessments. Technical Report. FMF. https:\/\/www.frontiermodelforum.org\/technical-reports\/third-party-assessments\/"},{"key":"e_1_3_2_1_30_1","unstructured":"Google. 2025. Gemini 2.5 Pro Model Card. Technical Report. Google. https:\/\/modelcards.withgoogle.com\/assets\/documents\/gemini-2.5-pro.pdf"},{"key":"e_1_3_2_1_31_1","unstructured":"Felix Hofst\u00e4tter Teun van der Weij Jayden Teoh Rada Djoneva Henning Bartsch and Francis Rhys Ward. 2025. The Elicitation Game: Evaluating Capability Elicitation Techniques. arXiv:2502.02180. Retrieved from https:\/\/arxiv.org\/abs\/2502.02180."},{"key":"e_1_3_2_1_32_1","unstructured":"UK AI Security Institute. 2024. Pre-Deployment evaluation of OpenAI's ol model. Technical Report. UK AI Security Institute. https:\/\/www.aisi.gov.uk\/blog\/pre-deployment-evaluation-of-openais-o1-model"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1108\/MAJ-08-2016-1425"},{"key":"e_1_3_2_1_34_1","unstructured":"ISO. 2022. ISO\/IEC 27001:2022. Technical Report. ISO. https:\/\/www.iso.org\/standard\/27001"},{"key":"e_1_3_2_1_35_1","unstructured":"ISO. 2022. ISO\/IEC 27002:2022. Technical Report. ISO. https:\/\/www.iso.org\/standard\/75652.html"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICESC57686.2023.10193389"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","unstructured":"Jung Koo Kang Clive Lennox Vivek Pandey Eric Allen Mark Defond Jason Guo Yonghong Jia Scott Judd Paul Koch Dan O'leary Tracie Majors Babak Mammadov Lorien Stice-Lawrence Richard Sloan K R Subramanyam Qian Wang Olena Watanabe Regina Wittenberg-Moerman and Suning Zhang. 2021. Client Concerns About Information Spillovers From Sharing Audit Partners * Client Concerns About Information Spillovers From Sharing Audit Partners. SSRN. Retrieved from http:\/\/dx.doi.org\/10.2139\/ssrn.3567535.","DOI":"10.2139\/ssrn.3567535"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Suzanne Lightman Theresa Suloway and Joseph Brule. 2022. Satellite ground segment :. Technical Report. National Institute of Standards and Technology. doi:10.6028\/NIST.IR.8401","DOI":"10.6028\/NIST.IR.8401"},{"key":"e_1_3_2_1_39_1","unstructured":"Shayne Longpre Kevin Klyman Ruth E. Appel Sayash Kapoor Rishi Bommasani Michelle Sahar Sean McGregor Avijit Ghosh Borhane Blili-Hamelin Nathan Butters Alondra Nelson Amit Elazari Andrew Sellars Casey John Ellis Dane Sherrets Dawn Song Harley Geiger Ilona Cohen Lauren McIlvenny Madhulika Srikumar Mark M. Jaycox Markus Anderljung Nadine Farid Johnson Nicholas Carlini Nicolas Miailhe Nik Marda Peter Henderson Rebecca S. Portnoff Rebecca Weiss Victoria Westerhoff Yacine Jernite Rumman Chowdhury Percy Liang and Arvind Narayanan. 2025. In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI. arXiv:2503.16861. Retrieved from http:\/\/arxiv.org\/abs\/2503.16861."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Shayne Longpre Robert Mahari Naana Obeng-Marnu William Brannon Tobin South Katy Gero Sandy Pentland and Jad Kabbara. 2024. Data Authenticity Consent & Provenance for AI are all broken: what will it take to fix them? arXiv:2404.12691. Retrieved from http:\/\/arxiv.org\/abs\/2404.12691.","DOI":"10.21428\/e4baedd9.a650f77d"},{"key":"e_1_3_2_1_41_1","unstructured":"Tegan McCaslin Jide Alaga Samira Nedungadi Seth Donoughe Tom Reed Rishi Bommasani Chris Painter and Luca Righetti. 2025. STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports. arXiv:2508.09853. Retrieved from http:\/\/arxiv.org\/abs\/2508.09853."},{"key":"e_1_3_2_1_42_1","unstructured":"Medicines and Healthcare products Regulatory Agency. 2025. Decentralised manufacture: UK Guideline on Good Manufacturing Practice (GMP). Technical Report. Medicines and Healthcare products Regulatory Agency. https:\/\/www.gov.uk\/guidance\/decentralised-manufacture-uk-guideline-on-good-manufacturing-practice-gmp#manufacturing-licence-applications-and-inspection-approach"},{"key":"e_1_3_2_1_43_1","unstructured":"METR. 2025. Common Elements of Frontier AI Safety Policies. Technical Report. METR. https:\/\/metr.org\/common-elements.pdf"},{"key":"e_1_3_2_1_44_1","unstructured":"METR. 2025. Details about METR's evaluation of OpenAI GPT-5. Technical Report. METR. https:\/\/evaluations.metr.org\/gpt-5-report\/"},{"key":"e_1_3_2_1_45_1","unstructured":"METR. 2025. Details about METR's preliminary evaluation of o3 and o4-mini. Technical Report. METR. https:\/\/evaluations.metr.org\/openai-o3-report\/#limitations"},{"key":"e_1_3_2_1_46_1","volume-title":"Review of the Anthropic","author":"METR.","year":"2025","unstructured":"METR. 2025. Review of the Anthropic Summer 2025 Pilot Sabotage Risk Report. https:\/\/metr.org\/2025_pilot_risk_report_metr_review.pdf"},{"key":"e_1_3_2_1_47_1","unstructured":"METR. 2025. What should companies share about risks from frontier AI models? Technical Report. METR. https:\/\/metr.org\/blog\/2025-06-27-risk-transparency\/ Accessed: 2025-06-27."},{"key":"e_1_3_2_1_48_1","unstructured":"Evan Miller. 2024. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv:2411.00640. Retrieved from http:\/\/arxiv.org\/abs\/2411.00640."},{"key":"e_1_3_2_1_49_1","unstructured":"Smitha Milli Ludwig Schmidt Anca D. Dragan and Moritz Hardt. 2018. Model Reconstruction from Model Explanations. arXiv:1807.05185. Retrieved from http:\/\/arxiv.org\/abs\/1807.05185."},{"key":"e_1_3_2_1_50_1","unstructured":"NIST. 2008. NIST Special Publication 800-115: Technical Guide to Information Security Testing and Assessment. Technical Report. NIST. https:\/\/nvlpubs.nist.gov\/nistpubs\/legacy\/sp\/nistspecialpublication800-115.pdf"},{"key":"e_1_3_2_1_51_1","unstructured":"NIST. 2020. NIST Special Publication 800-53 Revision 5: Security and Privacy Controls for Information Systems and Organizations. Technical Report. NIST. https:\/\/csrc.nist.gov\/CSRC\/media\/Projects\/risk-management\/800-53%20Downloads\/800-53r5\/SP_800-53_v5_1-derived-OSCAL.pdf"},{"key":"e_1_3_2_1_52_1","volume-title":"Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs. arXiv:2508.06601.","author":"O'Brien Kyle","year":"2025","unstructured":"Kyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman. 2025. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs. arXiv:2508.06601. Retrieved from http:\/\/arxiv.org\/abs\/2508.06601."},{"key":"e_1_3_2_1_53_1","unstructured":"Office of the Comptroller of the Currency. 2018. Comptroller's Handbook: Bank Supervision Process. Technical Report. Office of the Comptroller of the Currency. https:\/\/www.occ.gov\/publications-and-resources\/publications\/comptrollers-handbook\/files\/bank-supervision-process\/pub-ch-bank-supervision-process.pdf"},{"key":"e_1_3_2_1_54_1","unstructured":"OpenAI. 2023. Fine-tuning. https:\/\/perma.cc\/BFZ9-QBXN"},{"key":"e_1_3_2_1_55_1","unstructured":"OpenAI. 2023. Using logprobs. Technical Report. OpenAI. https:\/\/cookbook.openai.com\/examples\/using_logprobs"},{"key":"e_1_3_2_1_56_1","unstructured":"OpenAI. 2024. Trust Portal. https:\/\/trust.openai.com Accessed: 2024-01-15."},{"key":"e_1_3_2_1_57_1","unstructured":"OpenAI. 2025. GPT-5 System Card. Technical Report. OpenAI. https:\/\/cdn.openai.com\/gpt-5-system-card.pdf"},{"key":"e_1_3_2_1_58_1","unstructured":"OpenAI. 2025. OpenAI Model Spec. Technical Report. OpenAI. https:\/\/model-spec.openai.com\/2025-10-27.html"},{"key":"e_1_3_2_1_59_1","unstructured":"OpenAI. 2025. Working with US CAISI and UK AISI to build more secure AI systems. Technical Report. OpenAI. https:\/\/openai.com\/index\/us-caisi-uk-aisi-ai-update\/"},{"key":"e_1_3_2_1_60_1","unstructured":"OpenMined. 2023. How to audit an AI model owned by someone else. Technical Report. OpenMined. https:\/\/openmined.org\/blog\/ai-audit-part-1\/"},{"key":"e_1_3_2_1_61_1","unstructured":"Xiangyu Qi Yi Zeng Tinghao Xie Pin-Yu Chen Ruoxi Jia Prateek Mittal and Peter Henderson. 2023. Fine-tuning Aligned Language Models Compromises Safety Even When Users Do Not Intend To! arXiv:2310.03693. Retrieved from http:\/\/arxiv.org\/abs\/2310.03693."},{"key":"e_1_3_2_1_62_1","unstructured":"Tom Reed Tegan McCaslin and Luca Righetti. 2025. What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework. arXiv:2510.20927. Retrieved from http:\/\/arxiv.org\/abs\/2510.20927."},{"key":"e_1_3_2_1_63_1","volume-title":"Nitarshan Rajkumar, Nicolas Mo\u00ebs, Jeffrey Ladish, David Bau, Paul Bricman, Neel Guha, Jessica Newman, Yoshua Bengio, Tobin South, Alex Pentland, Sanmi Koyejo, Mykel J. Kochenderfer, and Robert Trager.","author":"Reuel Anka","year":"2025","unstructured":"Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, Markus Anderljung, Ben Garfinkel, Lennart Heim, Andrew Trask, Gabriel Mukobi, Rylan Schaeffer, Mauricio Baker, Sara Hooker, Irene Solaiman, Alexandra Sasha Luccioni, Nitarshan Rajkumar, Nicolas Mo\u00ebs, Jeffrey Ladish, David Bau, Paul Bricman, Neel Guha, Jessica Newman, Yoshua Bengio, Tobin South, Alex Pentland, Sanmi Koyejo, Mykel J. Kochenderfer, and Robert Trager. 2025. Open Problems in Technical AI Governance. arXiv:2407.14981. Retrieved from http:\/\/arxiv.org\/abs\/2407.14981."},{"key":"e_1_3_2_1_64_1","unstructured":"Luis Roque Carlos Soares Vitor Cerqueira and Luis Torgo. 2024. Cherry-Picking in Time Series Forecasting: How to Select Datasets to Make Your Model Shine. arXiv:2412.14435. Retrieved from https:\/\/arxiv.org\/abs\/2412.14435."},{"key":"e_1_3_2_1_65_1","unstructured":"Vinu Sankar Sadasivan Shoumik Saha Gaurang Sriramanan Priyatham Kattakinda Atoosa Chegini and Soheil Feizi. 2024. Fast Adversarial Attacks on Language Models In One GPU Minute. arXiv:2402.15570. Retrieved from https:\/\/arxiv.org\/abs\/2402.15570."},{"key":"e_1_3_2_1_66_1","unstructured":"William Saunders Catherine Yeh Jeff Wu Steven Bills Long Ouyang Jonathan Ward and Jan Leike. 2022. Self-critiquing models for assisting human evaluators. arXiv:2206.05802. Retrieved from http:\/\/arxiv.org\/abs\/2206.05802."},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"crossref","unstructured":"Leo Schwinn David Dobre Sophie Xhonneux Gauthier Gidel and Stephan Gunnemann. 2024. Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space. arXiv:2402.09063. Retrieved from https:\/\/arxiv.org\/abs\/2402.09063.","DOI":"10.52202\/079017-0288"},{"key":"e_1_3_2_1_68_1","unstructured":"Nevo Sella Lahav Dan Karpur Ajay Yogev. Bar-On Henry Alexander. Bradley and Jeff. Alstott. 2024. Securing AI model weights : preventing theft and misuse of frontier models. Technical Report. RAND. 117 pages. https:\/\/www.rand.org\/pubs\/research_reports\/RRA2849-1.html"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"crossref","unstructured":"Toby Shevlane. 2022. Structured access: an emerging paradigm for safe AI deployment. arXiv:2201.05159. Retrieved from https:\/\/arxiv.org\/abs\/2201.05159.","DOI":"10.1093\/oxfordhb\/9780197579329.013.39"},{"key":"e_1_3_2_1_70_1","unstructured":"Shivalika Singh Yiyang Nan Alex Wang Daniel D'Souza Sayash Kapoor Ahmet \u00dcst\u00fcn Sanmi Koyejo Yuntian Deng Shayne Longpre Noah A. Smith Beyza Ermis Marzieh Fadaee and Sara Hooker. 2025. The Leaderboard Illusion. arXiv:2504.20879. Retrieved from https:\/\/arxiv.org\/abs\/2504.20879."},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.840"},{"key":"e_1_3_2_1_72_1","volume-title":"Audit Cards: Contextualizing AI Evaluations. arXiv:2504.13839.","author":"Staufer Leon","year":"2025","unstructured":"Leon Staufer, Mick Yang, Anka Reuel, and Stephen Casper. 2025. Audit Cards: Contextualizing AI Evaluations. arXiv:2504.13839. Retrieved from https:\/\/arxiv.org\/abs\/2504.13839."},{"key":"e_1_3_2_1_73_1","unstructured":"Conrad Stosz Karson Elmgren Charles Foster George Balston Seth Donoughe Samira Nedungadi Michael Chen Jasper G\u00f6tting Patricia Paskov Sayash Kapoor Schwettmann Sarah Rishi Bommasani Luca Righetti Sean McGregor Grace Werner Christopher Painter Faisal Lalani Rob Reich Arvind Narayanan Elizabeth Barnes Miles Brundage Aidan Homewood Divya Siddharth Charles Teague Jaime Sevilla and Jacob Steinhardt. 2025. AEF-1: Minimum Operating Conditions for Independent Third Party AI Evaluations. Technical Report. AI Evaluator Forum. https:\/\/www.aef.one\/aef-one.pdf"},{"key":"e_1_3_2_1_74_1","unstructured":"Jordan Taylor Sid Black Dillon Bowen Thomas Read Satvik Golechha Alex Zelenka-Martin Oliver Makins Connor Kissane Kola Ayonrinde Jacob Merizian Samuel Marks Chris Cundy and Joseph Bloom. 2025. Auditing Games for Sandbagging. arXiv:2512.07810. Retrieved from http:\/\/arxiv.org\/abs\/2512.07810."},{"key":"e_1_3_2_1_75_1","unstructured":"Alejandro Tlaie and Jimmy Farrell. 2025. Securing External Deeper-than-black-box GPAI Evaluations. arXiv:2503.07496. Retrieved from http:\/\/arxiv.org\/abs\/2503.07496."},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"crossref","unstructured":"Vishaal Udandarao Ameya Prabhu Adhiraj Ghosh Yash Sharma Philip H. S. Torr Adel Bibi Samuel Albanie and Matthias Bethge. 2024. No \u201cZero-Shot\u201d Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance. arXiv:2404.04125. Retrieved from https:\/\/arxiv.org\/abs\/2404.04125.","DOI":"10.52202\/079017-1973"},{"key":"e_1_3_2_1_77_1","unstructured":"Eric Wallace Olivia Watkins Miles Wang Kai Chen and Chris Koch. 2025. Estimating Worst-Case Frontier Risks of Open-Weight LLMs. arXiv:2508.03153. Retrieved from http:\/\/arxiv.org\/abs\/2508.03153."},{"key":"e_1_3_2_1_78_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"12286","author":"Zanella-Beguelin Santiago","year":"2021","unstructured":"Santiago Zanella-Beguelin, Shruti Tople, Andrew Paverd, and Boris K\u00f6pf. 2021. Grey-box Extraction of Natural Language Models. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, Vienna, Austria, 12278\u201312286. https:\/\/proceedings.mlr.press\/v139\/zanella-beguelin21a.html"},{"key":"e_1_3_2_1_79_1","volume-title":"Representation Engineering: A Top-Down Approach to AI Transparency. arXiv:2310.01405.","author":"Zou Andy","year":"2025","unstructured":"Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2025. Representation Engineering: A Top-Down Approach to AI Transparency. arXiv:2310.01405. Retrieved from http:\/\/arxiv.org\/abs\/2310.01405."}],"event":{"name":"FAccT '26: The 2026 ACM Conference on Fairness, Accountability, and Transparency","location":"Montreal QC Canada","acronym":"FAccT '26","sponsor":["ACM\/SIG"]},"container-title":["Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805689.3812365","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T19:10:10Z","timestamp":1782760210000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805689.3812365"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":79,"alternative-id":["10.1145\/3805689.3812365","10.1145\/3805689"],"URL":"https:\/\/doi.org\/10.1145\/3805689.3812365","relation":{},"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}