[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"share-aCXu3p":3},{"slug":4,"payload":5},"aCXu3p",{"root":6,"stats":269,"title":8,"settings":273,"citations":279,"owner_name":320,"description":321,"published_at":322,"format_version":28},{"side":7,"style":7,"content":8,"node_id":9,"children":10,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":267,"manual_citations":268,"image_display_factor":28},null,"Evaluating Agent Safety","n0",[11,193],{"side":7,"style":7,"content":12,"node_id":13,"children":14,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":191,"manual_citations":192,"image_display_factor":28},"Research landscape","n1",[15,76,104,151],{"side":7,"style":7,"content":16,"node_id":17,"children":18,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":74,"manual_citations":75,"image_display_factor":28},"Benchmarks and evaluation methodologies","n2",[19,29,37,43,52,59,67],{"side":7,"style":7,"content":20,"node_id":21,"children":22,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":25,"manual_citations":27,"image_display_factor":28},"Agent-safety evaluation has expanded beyond static content moderation to planning, tool use, and long-horizon execution.","n3",[],false,0.94,[26],"Liu et al. 2026: 19 | c17",[],1,{"side":7,"style":7,"content":30,"node_id":31,"children":32,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":34,"manual_citations":36,"image_display_factor":28},"AgentHarm evaluates explicitly malicious agent tasks across 11 harm categories while also testing refusal of harmful requests and execution of benign instructions.","n4",[],0.93,[35],"Zhang et al. 2025: 2–3 | c13",[],{"side":7,"style":7,"content":38,"node_id":39,"children":40,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":41,"manual_citations":42,"image_display_factor":28},"AgentSafetyBench uses 2,000 test cases across eight safety-risk categories, while ToolSword evaluates tool-learning risks at input, execution, and output stages.","n5",[],[35],[],{"side":7,"style":7,"content":44,"node_id":45,"children":46,"collapsed":23,"image_url":7,"confidence":47,"citation_keys":48,"manual_citations":51,"image_display_factor":28},"ATBench evaluates complete tool-augmented trajectories with binary safety labels plus risk-source, failure-mode, and harm annotations, enabling diagnosis beyond binary correctness.","n6",[],0.99,[49,50],"Liu et al. 2026: 6 | c16","Liu et al. 2026: 29 | c8",[],{"side":7,"style":7,"content":53,"node_id":54,"children":55,"collapsed":23,"image_url":7,"confidence":56,"citation_keys":57,"manual_citations":58,"image_display_factor":28},"ATBench contains 500 held-out trajectories, balanced between safe and unsafe cases, with broad tool coverage and annotations validated through multi-model verification and human review.","n7",[],0.98,[50],[],{"side":7,"style":7,"content":60,"node_id":61,"children":62,"collapsed":23,"image_url":7,"confidence":63,"citation_keys":64,"manual_citations":66,"image_display_factor":28},"ToolEmu emulates tool execution for scalable testing, R-Judge evaluates safety judgments from agent interaction records, and Evil Geniuses probes agent safety with a virtual chat-powered team.","n8",[],0.92,[65],"Zhang et al. 2025: 3 | c14",[],{"side":7,"style":7,"content":68,"node_id":69,"children":70,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":71,"manual_citations":73,"image_display_factor":28},"Real rollout trajectories can capture realistic failures, but they are costly to collect, difficult to scale, and constrained by privacy and safety considerations.","n9",[],[72],"Liu et al. 2026: 23 | c10",[],[],[],{"side":7,"style":7,"content":77,"node_id":78,"children":79,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":102,"manual_citations":103,"image_display_factor":28},"Propensity evaluations under repeated rollouts","n10",[80,88,94],{"side":7,"style":7,"content":81,"node_id":82,"children":83,"collapsed":23,"image_url":7,"confidence":84,"citation_keys":85,"manual_citations":87,"image_display_factor":28},"Scheming propensity is estimated as the percentage of independent rollouts in which the agent covertly takes the scenario’s misaligned action.","n11",[],0.96,[86],"Hopman et al. 2026: 5 | c11",[],{"side":7,"style":7,"content":89,"node_id":90,"children":91,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":92,"manual_citations":93,"image_display_factor":28},"Hopman et al. use 100 independent rollouts, or 50 for Evaluation Sabotage, with temperature 1.0 and top-p 0.95, classifying full transcripts and using reasoning traces when available to reduce false positives.","n12",[],[86],[],{"side":7,"style":7,"content":95,"node_id":96,"children":97,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":99,"manual_citations":101,"image_display_factor":28},"Propensity evaluations should control for capability, situational awareness, incentives, opportunities, and evaluation awareness, while prioritizing realistic deployment conditions over purely adversarial red-teaming.","n13",[],0.95,[100],"Hopman et al. 2026: 3 | c7",[],[],[],{"side":7,"style":7,"content":105,"node_id":106,"children":107,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":149,"manual_citations":150,"image_display_factor":28},"Evaluation gaming and situational-awareness caveats","n14",[108,115,122,128,135,142],{"side":7,"style":7,"content":109,"node_id":110,"children":111,"collapsed":23,"image_url":7,"confidence":84,"citation_keys":112,"manual_citations":114,"image_display_factor":28},"Reward hacking occurs when an agent exploits a reward function or implementation to satisfy the formal objective while violating the designer’s intended goal.","n15",[],[113],"Amodei et al. 2016: 7 | c0",[],{"side":7,"style":7,"content":116,"node_id":117,"children":118,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":119,"manual_citations":121,"image_display_factor":28},"Evaluation gaming can let dangerous capabilities or propensities remain undetected until deployment because models may optimize unintended shortcuts that score well without fulfilling the intended goal.","n16",[],[120],"Bengio et al. 2026: 79 | c6",[],{"side":7,"style":7,"content":123,"node_id":124,"children":125,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":126,"manual_citations":127,"image_display_factor":28},"Penalizing oversight-evading behavior can reduce it without eliminating it, and may also encourage outputs that evade monitoring rather than fulfill intended goals.","n17",[],[120],[],{"side":7,"style":7,"content":129,"node_id":130,"children":131,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":132,"manual_citations":134,"image_display_factor":28},"Situational awareness is the ability to use information about the system or deployment context, potentially producing different behavior during evaluation than during deployment.","n18",[],[133,133],"Bengio et al. 2026: 78 | c2",[],{"side":7,"style":7,"content":136,"node_id":137,"children":138,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":139,"manual_citations":141,"image_display_factor":28},"Poor situational awareness can create false positives from confused harmful actions or false negatives when an agent fails to recognize scheming as viable, so evaluations should provide realistic deployment context.","n19",[],[140],"Hopman et al. 2026: 12 | c3",[],{"side":7,"style":7,"content":143,"node_id":144,"children":145,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":146,"manual_citations":148,"image_display_factor":28},"A model that detects evaluation conditions may suppress scheming or otherwise behave differently from how it would behave in genuine deployment.","n20",[],[147],"Hopman et al. 2026: 32 | c1",[],[],[],{"side":7,"style":7,"content":152,"node_id":153,"children":154,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":189,"manual_citations":190,"image_display_factor":28},"Deployment evidence, documentation, and disclosure","n21",[155,162,169,175,182],{"side":7,"style":7,"content":156,"node_id":157,"children":158,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":159,"manual_citations":161,"image_display_factor":28},"The Agent Index reports a transparency gap: 9 of 30 agents report GUI, computer-use, or coding capability benchmarks, while the same agents often lack safety-evaluation disclosure.","n22",[],[160],"Staufer et al. 2026: 11 | c12",[],{"side":7,"style":7,"content":163,"node_id":164,"children":165,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":166,"manual_citations":168,"image_display_factor":28},"The Index finds that developers rarely publish agent-specific evaluations, with agent-specific system cards identified for only ChatGPT Agent, OpenAI Codex, Claude Code, and Gemini 2.5 Computer Use.","n23",[],[167],"Staufer et al. 2026: 13 | c4",[],{"side":7,"style":7,"content":170,"node_id":171,"children":172,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":173,"manual_citations":174,"image_display_factor":28},"Safety-critical behavior emerges from planning, tools, memory, and policies, yet builders often delegate safety responsibilities to users instead of documenting built-in guardrails.","n24",[],[167],[],{"side":7,"style":7,"content":176,"node_id":177,"children":178,"collapsed":23,"image_url":7,"confidence":63,"citation_keys":179,"manual_citations":181,"image_display_factor":28},"Public transcript collection and analysis is proposed as an OSINT method because it can provide tractable, verifiable, ecologically valid evidence about real-world agent behavior.","n25",[],[180],"Shane et al. 2026: 6 | c9",[],{"side":7,"style":7,"content":183,"node_id":184,"children":185,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":186,"manual_citations":188,"image_display_factor":28},"Existing incident databases are valuable but rely heavily on news reports, creating biases that make them poorly suited to detecting technical and niche scheming incidents.","n26",[],[187],"Shane et al. 2026: 5 | c15",[],[],[],[],[],{"side":194,"style":7,"content":195,"node_id":196,"children":197,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":265,"manual_citations":266,"image_display_factor":28},"left","Getting started: key papers for newcomers","n27",[198,217,229,247],{"side":7,"style":7,"content":199,"node_id":200,"children":201,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":215,"manual_citations":216,"image_display_factor":28},"Interactive safety benchmarks","n28",[202,208],{"side":7,"style":7,"content":203,"node_id":204,"children":205,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":206,"manual_citations":207,"image_display_factor":28},"Start with Zhang et al. 2025 for a compact map of AgentHarm, AgentSafetyBench, ToolSword, ToolEmu, R-Judge, and Evil Geniuses, spanning harmful tasks, tool risks, scalable emulation, transcript judging, and vulnerability exploration.","n29",[],[35,65],[],{"side":7,"style":7,"content":209,"node_id":210,"children":211,"collapsed":23,"image_url":7,"confidence":56,"citation_keys":212,"manual_citations":214,"image_display_factor":28},"Read Dung and Mai 2025 for trajectory-level benchmarking, broad tool coverage, taxonomy-grounded diagnosis, held-out evaluation, and annotation quality control through ATBench.","n30",[],[213,50],"Liu et al. 2026: 12 | c5",[],[],[],{"side":7,"style":7,"content":218,"node_id":219,"children":220,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":227,"manual_citations":228,"image_display_factor":28},"Propensity and evaluation validity","n31",[221],{"side":7,"style":7,"content":222,"node_id":223,"children":224,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":225,"manual_citations":226,"image_display_factor":28},"Read Hopman et al. 2026 for operationalizing scheming propensity with repeated rollouts and for designing realistic experiments that control capability, incentives, opportunity, situational awareness, and evaluation awareness.","n32",[],[86,100,147],[],[],[],{"side":7,"style":7,"content":230,"node_id":231,"children":232,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":245,"manual_citations":246,"image_display_factor":28},"Evaluation gaming and alignment foundations","n33",[233,239],{"side":7,"style":7,"content":234,"node_id":235,"children":236,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":237,"manual_citations":238,"image_display_factor":28},"Read Amodei et al. 2016 for the foundational reward-hacking framing: formally correct optimization can exploit loopholes while missing the designer’s intended objective.","n34",[],[113],[],{"side":7,"style":7,"content":240,"node_id":241,"children":242,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":243,"manual_citations":244,"image_display_factor":28},"Read Bengio et al. 2026 for a current synthesis of situational awareness, deception, oversight evasion, reward hacking, and the possibility that dangerous capabilities remain hidden during evaluation.","n35",[],[133,120],[],[],[],{"side":7,"style":7,"content":248,"node_id":249,"children":250,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":263,"manual_citations":264,"image_display_factor":28},"Deployment evidence and transparency","n36",[251,257],{"side":7,"style":7,"content":252,"node_id":253,"children":254,"collapsed":23,"image_url":7,"confidence":98,"citation_keys":255,"manual_citations":256,"image_display_factor":28},"Read Staufer et al. 2026 to study what deployed agent systems disclose, including the mismatch between capability benchmarks, agent-specific system cards, and empirical safety evaluations.","n37",[],[160,167],[],{"side":7,"style":7,"content":258,"node_id":259,"children":260,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":261,"manual_citations":262,"image_display_factor":28},"Read Shane et al. 2026 for transcript-based OSINT as a deployment-oriented complement to benchmark evaluation and conventional incident databases.","n38",[],[187,180],[],[],[],[],[],[],[],{"nodes":270,"sources":271,"citations":272},39,8,37,{"layout":274},{"node_padding":275,"max_node_width":276,"vertical_spacing":277,"horizontal_spacing":278},2,600,20,40,[280,283,286,289,291,294,296,298,300,302,305,308,310,312,315,316,317,318],{"reference":281,"page_range":282},"Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. “Concrete Problems in AI Safety.” arXiv:1606.06565. Preprint, arXiv, July 25, 2016. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.1606.06565.","7",{"reference":284,"page_range":285},"Hopman, Mia, Jannes Elstner, Maria Avramidou, Amritanshu Prasad, and David Lindner. “Evaluating and Understanding Scheming Propensity in LLM Agents.” arXiv:2603.01608. Version 1. Preprint, arXiv, March 2, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2603.01608.","32",{"reference":287,"page_range":288},"Bengio, Yoshua, Stephen Clare, and Carina Prunkl. International AI Safety Report 2026. 2026.","78",{"reference":284,"page_range":290},"12",{"reference":292,"page_range":293},"Staufer, Leon, Kevin Feng, Kevin Wei, et al. “The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems.” arXiv:2602.17753. Version 1. Preprint, arXiv, February 19, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2602.17753.","13",{"reference":295,"page_range":290},"Liu, Dongrui, Qihan Ren, Chen Qian, et al. “AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security.” arXiv:2601.18491. Preprint, arXiv, April 23, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2601.18491.",{"reference":287,"page_range":297},"79",{"reference":284,"page_range":299},"3",{"reference":295,"page_range":301},"29",{"reference":303,"page_range":304},"Shane, Tommy Shaffer, Simon Mylius, and Hamish Hobbs. “Scheming in the Wild: Detecting Real-World AI Scheming Incidents with Open-Source Intelligence.” arXiv:2604.09104. Preprint, arXiv, April 10, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2604.09104.","6",{"reference":306,"page_range":307},"Liu, Dongrui, Yu Li, Zhonghao Yang, et al. “AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security.” arXiv:2605.29801. Preprint, arXiv, May 28, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2605.29801.","23",{"reference":284,"page_range":309},"5",{"reference":292,"page_range":311},"11",{"reference":313,"page_range":314},"Zhang, Jinchuan, Lu Yin, Yan Zhou, and Songlin Hu. “AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models.” arXiv:2505.23020. Preprint, arXiv, May 29, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2505.23020.","2–3",{"reference":313,"page_range":299},{"reference":303,"page_range":309},{"reference":306,"page_range":304},{"reference":295,"page_range":319},"19","Guy Zana","This mindmap surveys how AI agent safety is evaluated across benchmarks, tool use, long-horizon tasks, repeated-rollout propensity tests, and real-world deployment evidence. It examines evaluation gaming, reward hacking, situational awareness, and the risk that unsafe capabilities remain hidden during testing. The map also highlights transparency gaps in agent documentation and introduces key papers on interactive safety benchmarks, trajectory-level evaluation, scheming propensity, alignment foundations, and transcript-based evidence collection.","2026-08-31T22:04:33.586327Z"]