[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"share-Mx8P5V":3},{"slug":4,"payload":5},"Mx8P5V",{"root":6,"stats":316,"title":8,"settings":320,"citations":326,"owner_name":380,"description":381,"published_at":382,"format_version":25},{"side":7,"style":7,"content":8,"node_id":9,"children":10,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":314,"manual_citations":315,"image_display_factor":25},null,"Prompt Injection: Attacks and Defenses","n0",[11,51,123,180,209,236],{"side":7,"style":7,"content":12,"node_id":13,"children":14,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":49,"manual_citations":50,"image_display_factor":25},"Foundations and threat model","n1",[15,26,34,42],{"side":7,"style":7,"content":16,"node_id":17,"children":18,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":21,"manual_citations":24,"image_display_factor":25},"Prompt injection has two broad forms: direct injection inserts malicious instructions into the input prompt, whereas indirect injection places them in external data processed by the agent.","n2",[],false,0.96,[22,23],"Chhabra et al. 2026: 3 | c6","Chhabra et al. 2026: 4–5 | c14",[],1,{"side":7,"style":7,"content":27,"node_id":28,"children":29,"collapsed":19,"image_url":7,"confidence":30,"citation_keys":31,"manual_citations":33,"image_display_factor":25},"AgentDoG separates agentic risk source, failure mode, and real-world harm, distinguishing where risk originates, how it manifests, and what consequences result.","n3",[],0.98,[32],"Liu et al. 2026: 4 | c1",[],{"side":7,"style":7,"content":35,"node_id":36,"children":37,"collapsed":19,"image_url":7,"confidence":38,"citation_keys":39,"manual_citations":41,"image_display_factor":25},"Greshake et al. model indirect injection as adversarial prompts embedded in sources retrieved at inference time, allowing control of an LLM-integrated application without direct model access.","n4",[],0.97,[40],"Greshake, Abdelnabi et al. 2023: 1 | c3",[],{"side":7,"style":7,"content":43,"node_id":44,"children":45,"collapsed":19,"image_url":7,"confidence":38,"citation_keys":46,"manual_citations":48,"image_display_factor":25},"Greshake et al. distinguish passive retrieval, active channels such as email, user-driven, and hidden injection methods by how adversarial instructions reach or influence the LLM.","n5",[],[47],"Greshake, Abdelnabi et al. 2023: 3 | c5",[],[],[],{"side":7,"style":7,"content":52,"node_id":53,"children":54,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":121,"manual_citations":122,"image_display_factor":25},"Attack techniques and surfaces","n6",[55,62,71,77,84,92,98,105,112],{"side":7,"style":7,"content":56,"node_id":57,"children":58,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":60,"manual_citations":61,"image_display_factor":25},"Direct injection can make an agent ignore prior guidelines, extract sensitive internal data, send unauthorized emails, and expose confidential information.","n7",[],0.94,[22],[],{"side":7,"style":7,"content":63,"node_id":64,"children":65,"collapsed":19,"image_url":7,"confidence":66,"citation_keys":67,"manual_citations":70,"image_display_factor":25},"Indirect injection hides instructions in webpages, documents, databases, APIs, tool outputs, or retrieval results, allowing an external content supplier to steer the agent without directly addressing the user prompt.","n8",[],0.95,[23,68,69],"Chhabra et al. 2026: 5 | c13","Triedman et al. 2025: 16 | c4",[],{"side":7,"style":7,"content":72,"node_id":73,"children":74,"collapsed":19,"image_url":7,"confidence":30,"citation_keys":75,"manual_citations":76,"image_display_factor":25},"Agent risk sources include malicious user instructions or jailbreaks, direct prompt injection, unreliable observations, tool-description injection, malicious tool execution, corrupted tool feedback, and inherent agent or LLM failures.","n9",[],[32],[],{"side":7,"style":7,"content":78,"node_id":79,"children":80,"collapsed":19,"image_url":7,"confidence":81,"citation_keys":82,"manual_citations":83,"image_display_factor":25},"Web and computer agents are vulnerable when attackers manipulate HTML, accessibility trees, or interface interactions, exploiting the decoupling between user input and later tool calls.","n10",[],0.93,[68],[],{"side":7,"style":7,"content":85,"node_id":86,"children":87,"collapsed":19,"image_url":7,"confidence":88,"citation_keys":89,"manual_citations":91,"image_display_factor":25},"Multi-agent system control-flow hijacking is presented as a distinct attack class that strategically manipulates multi-agent control flow toward arbitrary code execution, with the user as an unwitting victim.","n11",[],0.91,[90],"Triedman et al. 2025: 2 | c0",[],{"side":7,"style":7,"content":93,"node_id":94,"children":95,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":96,"manual_citations":97,"image_display_factor":25},"Greshake et al. organize indirect-injection consequences into information gathering, fraud, intrusion, malware, manipulated content, and availability harms, including leakage, remote control, disinformation, and denial of service.","n12",[],[47],[],{"side":7,"style":7,"content":99,"node_id":100,"children":101,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":102,"manual_citations":104,"image_display_factor":25},"In a GPT-4 synthetic application, Greshake et al. demonstrate remote control by repeatedly retrieving attacker instructions through search or a URL, enabling bidirectional communication and a potential remotely accessible backdoor.","n13",[],[103],"Greshake, Abdelnabi et al. 2023: 7 | c18",[],{"side":7,"style":7,"content":106,"node_id":107,"children":108,"collapsed":19,"image_url":7,"confidence":38,"citation_keys":109,"manual_citations":111,"image_display_factor":25},"Zhan et al.’s InjecAgent evaluates external attackers embedding malicious prompts in retrieved content to make tool-equipped LLM agents perform harmful actions, using formal attack definitions and GPT-4-assisted test-case generation.","n14",[],[110,110],"Zhan, Liang et al. 2024: 1 | c22",[],{"side":7,"style":7,"content":113,"node_id":114,"children":115,"collapsed":19,"image_url":7,"confidence":38,"citation_keys":116,"manual_citations":120,"image_display_factor":25},"InjecAgent contains 17 user cases, 62 attacker cases, and 1,054 test cases, and evaluates 30 agents. Its results report attack success across direct-harm and data-stealing scenarios, with substantial variation across models and settings.","n15",[],[117,118,119],"Zhan, Liang et al. 2024: 4 | c20","Zhan, Liang et al. 2024: 9 | c17","Zhan, Liang et al. 2024: 5 | c16",[],[],[],{"side":7,"style":7,"content":124,"node_id":125,"children":126,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":178,"manual_citations":179,"image_display_factor":25},"Defenses and security architecture","n16",[127,134,142,149,156,164,170],{"side":7,"style":7,"content":128,"node_id":129,"children":130,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":131,"manual_citations":133,"image_display_factor":25},"Prompt-injection defenses can be organized as agent-focused, system-focused, and user-focused methods, with system-focused methods divided into training-based and training-free approaches.","n17",[],[132],"Chhabra et al. 2026: 11–12 | c9",[],{"side":7,"style":7,"content":135,"node_id":136,"children":137,"collapsed":19,"image_url":7,"confidence":81,"citation_keys":138,"manual_citations":141,"image_display_factor":25},"Prompt-level defenses include delimiters, explicit instructions to ignore embedded directives, prompt sandwiching, reminders about adversarial risk, and repeated task instructions after tool responses.","n18",[],[139,140],"Debenedetti et al. 2025: 32 | c11","Triedman et al. 2025: 17 | c15",[],{"side":7,"style":7,"content":143,"node_id":144,"children":145,"collapsed":19,"image_url":7,"confidence":81,"citation_keys":146,"manual_citations":148,"image_display_factor":25},"Isolation defenses limit the tools available while an agent processes untrusted input, including committing to a predefined tool set and disabling access to other tools; Signed-Prompt adds authorization for sensitive command segments.","n19",[],[147,140],"Chhabra et al. 2026: 13 | c19",[],{"side":7,"style":7,"content":150,"node_id":151,"children":152,"collapsed":19,"image_url":7,"confidence":153,"citation_keys":154,"manual_citations":155,"image_display_factor":25},"Stronger system designs separate instructions from data architecturally, track trusted versus untrusted tool-output flows, and monitor internal state or formal task representations for shifts in the agent’s task.","n20",[],0.89,[139,139],[],{"side":7,"style":7,"content":157,"node_id":158,"children":159,"collapsed":19,"image_url":7,"confidence":160,"citation_keys":161,"manual_citations":163,"image_display_factor":25},"Instruction hierarchy prioritizes trusted instructions and is proposed to improve resistance to known and novel prompt injections while preserving general performance with limited impact.","n21",[],0.87,[162],"Ji et al. 2025: 47–48 | c7",[],{"side":7,"style":7,"content":165,"node_id":166,"children":167,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":168,"manual_citations":169,"image_display_factor":25},"Prompt augmentation is inexpensive and useful in baseline settings, but adaptive attacks can bypass it, so it should not be treated as a standalone security guarantee.","n22",[],[147,23],[],{"side":7,"style":7,"content":171,"node_id":172,"children":173,"collapsed":19,"image_url":7,"confidence":174,"citation_keys":175,"manual_citations":177,"image_display_factor":25},"Risk-sensitive approval mechanisms require confirmation for sensitive CLI, file-editing, command-execution, authentication, or payment actions, while some agents provide live oversight modes.","n23",[],0.9,[176],"Staufer et al. 2026: 9 | c8",[],[],[],{"side":181,"style":7,"content":182,"node_id":183,"children":184,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":207,"manual_citations":208,"image_display_factor":25},"left","Evaluation and benchmarks","n24",[185,192,199],{"side":7,"style":7,"content":186,"node_id":187,"children":188,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":189,"manual_citations":191,"image_display_factor":25},"AgentDojo, introduced by Debenedetti et al. (2024) and described here via Liu et al., evaluates indirect injection in realistic tool-use environments with adversarial instructions embedded in emails, webpages, and documents, reporting benign utility, utility under attack, and attack success rate.","n25",[],[190],"Liu et al. 2026: 42 | c10",[],{"side":7,"style":7,"content":193,"node_id":194,"children":195,"collapsed":19,"image_url":7,"confidence":88,"citation_keys":196,"manual_citations":198,"image_display_factor":25},"Agent Security Bench measures attacks across many agents and tools, while AgentHarm covers 110 harmful tasks across eleven categories with and without jailbreaks; AgentSafetyBench adds interactive tool-use safety evaluation.","n26",[],[69,197],"Liu et al. 2026: 17 | c12",[],{"side":7,"style":7,"content":200,"node_id":201,"children":202,"collapsed":19,"image_url":7,"confidence":203,"citation_keys":204,"manual_citations":206,"image_display_factor":25},"The benchmark landscape spans tool-use, web, computer, embodied, MCP, multi-agent, and multi-domain settings, with differences in process awareness, multi-turn interaction, sandboxing, judges, and reliability metrics.","n27",[],0.92,[205],"Chhabra et al. 2026: 18 | c2",[],[],[],{"side":181,"style":7,"content":210,"node_id":211,"children":212,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":234,"manual_citations":235,"image_display_factor":25},"Research landscape and open problems","n28",[213,219,227],{"side":7,"style":7,"content":214,"node_id":215,"children":216,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":217,"manual_citations":218,"image_display_factor":25},"The field is separating traditional user-driven jailbreak research from agent-security research on indirect injection, tool use, retrieval, and multi-agent systems.","n29",[],[69],[],{"side":7,"style":7,"content":220,"node_id":221,"children":222,"collapsed":19,"image_url":7,"confidence":223,"citation_keys":224,"manual_citations":226,"image_display_factor":25},"Adaptive attacks remain a central challenge: one corpus review reports 50% success against eight indirect-injection defenses, while benchmark summaries report average attack success above 80% in a broader agent-security setting.","n30",[],0.88,[23,225],"Sha et al. 2025: 10 | c21",[],{"side":7,"style":7,"content":228,"node_id":229,"children":230,"collapsed":19,"image_url":7,"confidence":174,"citation_keys":231,"manual_citations":233,"image_display_factor":25},"Multi-agent security studies how prompt injection, poisoned tool outputs, and malformed intermediate results can propagate misleading outputs across collaborating agents.","n31",[],[232],"Raza et al. 2026: 15 | c23",[],[],[],{"side":181,"style":7,"content":237,"node_id":238,"children":239,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":312,"manual_citations":313,"image_display_factor":25},"Getting started: key papers for newcomers","n32",[240,258,276,294],{"side":7,"style":7,"content":241,"node_id":242,"children":243,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":256,"manual_citations":257,"image_display_factor":25},"Foundational attacks and taxonomies","n33",[244,250],{"side":7,"style":7,"content":245,"node_id":246,"children":247,"collapsed":19,"image_url":7,"confidence":30,"citation_keys":248,"manual_citations":249,"image_display_factor":25},"Read Greshake et al. (2023) for the original indirect-injection threat model, injection-method taxonomy, harm categories, and remote-control demonstration. Read Zhan et al. (2024) for InjecAgent’s benchmark design, 30-agent evaluation, and attack-success findings.","n34",[],[40,47,103,118,119],[],{"side":7,"style":7,"content":251,"node_id":252,"children":253,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":254,"manual_citations":255,"image_display_factor":25},"Read Chhabra et al. (2026) to learn the direct-versus-indirect distinction, agent attack surfaces, adaptive bypass evidence, and the defense taxonomy.","n35",[],[22,23,132],[],[],[],{"side":7,"style":7,"content":259,"node_id":260,"children":261,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":274,"manual_citations":275,"image_display_factor":25},"Agent and multi-agent security","n36",[262,268],{"side":7,"style":7,"content":263,"node_id":264,"children":265,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":266,"manual_citations":267,"image_display_factor":25},"Read Triedman et al. (2025) to study multi-agent system control-flow hijacking, its distinction from jailbreaks and ordinary indirect injection, and the security implications of arbitrary code execution.","n37",[],[90],[],{"side":7,"style":7,"content":269,"node_id":270,"children":271,"collapsed":19,"image_url":7,"confidence":30,"citation_keys":272,"manual_citations":273,"image_display_factor":25},"Read Liu et al. (2026), AgentDoG, to learn how agentic safety analysis separates risk source, failure mode, and real-world harm, including tool-description injection, corrupted tool feedback, and inherent agent or LLM failures.","n38",[],[32,32],[],[],[],{"side":7,"style":7,"content":277,"node_id":278,"children":279,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":292,"manual_citations":293,"image_display_factor":25},"Defenses and secure systems design","n39",[280,286],{"side":7,"style":7,"content":281,"node_id":282,"children":283,"collapsed":19,"image_url":7,"confidence":59,"citation_keys":284,"manual_citations":285,"image_display_factor":25},"Read Debenedetti et al. (2025) to compare delimiters, prompt sandwiching, structured queries, fine-tuning, internal-state monitoring, formal task languages, architectural separation, canaries, and information-flow control.","n40",[],[139,139],[],{"side":7,"style":7,"content":287,"node_id":288,"children":289,"collapsed":19,"image_url":7,"confidence":174,"citation_keys":290,"manual_citations":291,"image_display_factor":25},"Read Raza et al. (2026) to connect adversarial training, safety constraints, tool-output validation, and multi-agent propagation within a broader trust, risk, and security-management framework.","n41",[],[232],[],[],[],{"side":7,"style":7,"content":295,"node_id":296,"children":297,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":310,"manual_citations":311,"image_display_factor":25},"Evaluation and benchmarking","n42",[298,304],{"side":7,"style":7,"content":299,"node_id":300,"children":301,"collapsed":19,"image_url":7,"confidence":66,"citation_keys":302,"manual_citations":303,"image_display_factor":25},"Read Debenedetti et al. (2024), AgentDojo, to learn how realistic email, webpage, and document environments measure benign utility, utility under attack, and attack success rate. The corpus covers it via Liu et al.","n43",[],[190],[],{"side":7,"style":7,"content":305,"node_id":306,"children":307,"collapsed":19,"image_url":7,"confidence":81,"citation_keys":308,"manual_citations":309,"image_display_factor":25},"Read Liu et al. (2026) and Chhabra et al. (2026) to compare AgentHarm, AgentSafetyBench, AgentSecurityBench, AgentDojo, AgentDyn, and broader security benchmarks across metrics and environments.","n44",[],[197,205],[],[],[],[],[],[],[],{"nodes":317,"sources":318,"citations":319},45,11,55,{"layout":321},{"node_padding":322,"max_node_width":323,"vertical_spacing":324,"horizontal_spacing":325},2,600,20,40,[327,330,333,336,339,341,343,344,347,350,352,355,358,360,362,364,365,367,368,370,372,373,376,377],{"reference":328,"page_range":329},"Triedman, Harold, Rishi Jha, and Vitaly Shmatikov. “Multi-Agent Systems Execute Arbitrary Malicious Code.” arXiv:2503.12188. Preprint, arXiv, September 12, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2503.12188.","2",{"reference":331,"page_range":332},"Liu, Dongrui, Qihan Ren, Chen Qian, et al. “AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security.” arXiv:2601.18491. Preprint, arXiv, April 23, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2601.18491.","4",{"reference":334,"page_range":335},"Chhabra, Anshuman, Shrestha Datta, Shahriar Kabir Nahin, and Prasant Mohapatra. “Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges.” IEEE Access 14 (2026): 49455–82. https:\u002F\u002Fdoi.org\u002F10.1109\u002FACCESS.2026.3675554.","18",{"reference":337,"page_range":338},"Greshake, Kai, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. “Not What You'Ve Signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” Version 2. Preprint, ArXiv. https:\u002F\u002Fdoi.org\u002F10.48550\u002FARXIV.2302.12173.","1",{"reference":328,"page_range":340},"16",{"reference":337,"page_range":342},"3",{"reference":334,"page_range":342},{"reference":345,"page_range":346},"Ji, Jiaming, Tianyi Qiu, Boyuan Chen, et al. “AI Alignment: A Comprehensive Survey.” arXiv:2310.19852. Preprint, arXiv, April 4, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2310.19852.","47–48",{"reference":348,"page_range":349},"Staufer, Leon, Kevin Feng, Kevin Wei, et al. “The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems.” arXiv:2602.17753. Version 1. Preprint, arXiv, February 19, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2602.17753.","9",{"reference":334,"page_range":351},"11–12",{"reference":353,"page_range":354},"Liu, Dongrui, Yu Li, Zhonghao Yang, et al. “AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security.” arXiv:2605.29801. Preprint, arXiv, May 28, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2605.29801.","42",{"reference":356,"page_range":357},"Debenedetti, Edoardo, Ilia Shumailov, Tianqi Fan, et al. “Defeating Prompt Injections by Design.” arXiv:2503.18813. Preprint, arXiv, June 24, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2503.18813.","32",{"reference":353,"page_range":359},"17",{"reference":334,"page_range":361},"5",{"reference":334,"page_range":363},"4–5",{"reference":328,"page_range":359},{"reference":366,"page_range":361},"Zhan, Qiusi, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024. “InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.” Version 3. Preprint, ArXiv. https:\u002F\u002Fdoi.org\u002F10.48550\u002FARXIV.2403.02691.",{"reference":366,"page_range":349},{"reference":337,"page_range":369},"7",{"reference":334,"page_range":371},"13",{"reference":366,"page_range":332},{"reference":374,"page_range":375},"Sha, Zeyang, Hanling Tian, Zhuoer Xu, Shiwen Cui, Changhua Meng, and Weiqiang Wang. “Agent Safety Alignment via Reinforcement Learning.” arXiv:2507.08270. Preprint, arXiv, July 11, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2507.08270.","10",{"reference":366,"page_range":338},{"reference":378,"page_range":379},"Raza, Shaina, Ranjan Sapkota, Manoj Karkee, and Christos Emmanouilidis. “TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-Based Agentic Multi-Agent Systems.” AI Open 7 (January 2026): 71–95. https:\u002F\u002Fdoi.org\u002F10.1016\u002Fj.aiopen.2026.02.006.","15","Guy Zana","The attack works because your agent reads. Anything it reads, a webpage, a tool output, a PDF, can carry instructions someone else wrote. This map traces the field from the original indirect-injection paper through today's attack surfaces, then organizes the defenses: prompt-level tricks and their known bypass rates, isolation, instruction hierarchy, and architectural separation. Benchmarks and a grouped reading list included, with every claim linked to its source passage.","2026-08-25T15:25:38.895350Z"]