[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"share-KUkhPg":3},{"slug":4,"payload":5},"KUkhPg",{"root":6,"stats":320,"title":8,"settings":324,"citations":330,"owner_name":379,"description":380,"published_at":381,"format_version":24},{"side":7,"style":7,"content":8,"node_id":9,"children":10,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":318,"manual_citations":319,"image_display_factor":24},null,"Interpretability as a Safety Tool","n0",[11,35,98,127,162,183,210,236],{"side":7,"style":7,"content":12,"node_id":13,"children":14,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":33,"manual_citations":34,"image_display_factor":24},"Why interpretability matters for safety","n1",[15,25],{"side":7,"style":7,"content":16,"node_id":17,"children":18,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":21,"manual_citations":23,"image_display_factor":24},"Post hoc interpretability seeks to expose internal units and their causal effects on model outputs after training.","n2",[],false,0.94,[22],"Ji et al. 2025: 50–51 | c17",[],1,{"side":7,"style":7,"content":26,"node_id":27,"children":28,"collapsed":19,"image_url":7,"confidence":29,"citation_keys":30,"manual_citations":32,"image_display_factor":24},"A central research aim is to connect microscopic internal analysis to macroscopic behaviors such as planning and deception.","n3",[],0.88,[31],"Ji et al. 2025: 52 | c3",[],[],[],{"side":7,"style":7,"content":36,"node_id":37,"children":38,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":96,"manual_citations":97,"image_display_factor":24},"Sparse autoencoders and monosemantic features","n4",[39,45,53,59,67,75,82,89],{"side":7,"style":7,"content":40,"node_id":41,"children":42,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":43,"manual_citations":44,"image_display_factor":24},"Superposition means models can represent more features than they have dimensions, so individual neurons need not correspond cleanly to individual concepts.","n5",[],[31],[],{"side":7,"style":7,"content":46,"node_id":47,"children":48,"collapsed":19,"image_url":7,"confidence":49,"citation_keys":50,"manual_citations":52,"image_display_factor":24},"Dictionary learning treats dense activations as sparse combinations of latent vectors, while SAEs operationalize this idea to extract interpretable features from superposed activations.","n6",[],0.92,[51,22],"Templeton et al. 2026: 53 | c1",[],{"side":7,"style":7,"content":54,"node_id":55,"children":56,"collapsed":19,"image_url":7,"confidence":29,"citation_keys":57,"manual_citations":58,"image_display_factor":24},"Bricken et al. and Cunningham et al. showed that SAEs could recover interpretable, monosemantic transformer features, while Tamkin et al. reported a related binary-feature approach.","n7",[],[51],[],{"side":7,"style":7,"content":60,"node_id":61,"children":62,"collapsed":19,"image_url":7,"confidence":63,"citation_keys":64,"manual_citations":66,"image_display_factor":24},"Cunningham et al. model language-model activation vectors, including Pythia-70M activations, as sparse combinations of unknown network features and train dictionary features to approximate those underlying directions.","n8",[],0.96,[65],"Cunningham, Ewart et al. 2023: 2 | c16",[],{"side":7,"style":7,"content":68,"node_id":69,"children":70,"collapsed":19,"image_url":7,"confidence":71,"citation_keys":72,"manual_citations":74,"image_display_factor":24},"Their SAE uses a single ReLU hidden layer whose width is an overcomplete multiple of input dimension, with a sparsity penalty on hidden activations and tied encoder-decoder weights.","n9",[],0.97,[73],"Cunningham, Ewart et al. 2023: 2–3 | c20",[],{"side":7,"style":7,"content":76,"node_id":77,"children":78,"collapsed":19,"image_url":7,"confidence":71,"citation_keys":79,"manual_citations":81,"image_display_factor":24},"Cunningham et al. measure feature interpretability with autointerpretation: a language model describes activating examples, predicts activations from that description on new examples, and receives a correlation-based score.","n10",[],[80],"Cunningham, Ewart et al. 2023: 3 | c10",[],{"side":7,"style":7,"content":83,"node_id":84,"children":85,"collapsed":19,"image_url":7,"confidence":63,"citation_keys":86,"manual_citations":88,"image_display_factor":24},"The learned sparse-coding features score as more interpretable than ICA, Identity ReLU, PCA, and Random baselines, but the advantage declines with depth, becoming comparable to ICA in layer 4 and minimal in the final layer.","n11",[],[87],"Cunningham, Ewart et al. 2023: 5 | c5",[],{"side":7,"style":7,"content":90,"node_id":91,"children":92,"collapsed":19,"image_url":7,"confidence":29,"citation_keys":93,"manual_citations":95,"image_display_factor":24},"TopK-SAEs and related methods have extended feature extraction toward larger models, but scaling analysis from toy systems to real safety-relevant behavior remains difficult.","n12",[],[94,31],"Shi et al. 2026: 13 | c4",[],[],[],{"side":7,"style":7,"content":99,"node_id":100,"children":101,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":125,"manual_citations":126,"image_display_factor":24},"Safety-relevant features in frontier models","n13",[102,109,117],{"side":7,"style":7,"content":103,"node_id":104,"children":105,"collapsed":19,"image_url":7,"confidence":49,"citation_keys":106,"manual_citations":108,"image_display_factor":24},"SAEs trained on Claude 3 Sonnet recovered abstract, multilingual, and multimodal features, including concepts connected to code security vulnerabilities.","n14",[],[107],"Templeton et al. 2026: 1–2 | c21",[],{"side":7,"style":7,"content":110,"node_id":111,"children":112,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":114,"manual_citations":116,"image_display_factor":24},"The Claude 3 Sonnet study identified features associated with deception, power-seeking, sycophancy, and bias, and reported that manipulating them causally influenced outputs.","n15",[],0.9,[115],"Templeton et al. 2026: 1 | c11",[],{"side":7,"style":7,"content":118,"node_id":119,"children":120,"collapsed":19,"image_url":7,"confidence":121,"citation_keys":122,"manual_citations":124,"image_display_factor":24},"In experiments on Qwen3 and Llama-3.2 reward-model variants, safety-related SAE features formed coherent clusters, with safe and unsafe features locally separated in visualizations.","n16",[],0.84,[123],"Shi et al. 2026: 16 | c23",[],[],[],{"side":7,"style":7,"content":128,"node_id":129,"children":130,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":160,"manual_citations":161,"image_display_factor":24},"Probing reward models with sparse autoencoders","n17",[131,139,147,153],{"side":7,"style":7,"content":132,"node_id":133,"children":134,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":135,"manual_citations":138,"image_display_factor":24},"Reward models convert comparison feedback into scalar rewards, but can learn incomplete or suboptimal objectives that produce reward hacking.","n18",[],[136,137],"Ji et al. 2025: 22 | c6","Ji et al. 2025: 5 | c13",[],{"side":7,"style":7,"content":140,"node_id":141,"children":142,"collapsed":19,"image_url":7,"confidence":143,"citation_keys":144,"manual_citations":146,"image_display_factor":24},"SAFER trains an SAE on safety-oriented reward-model hidden states and uses activation differences between chosen and rejected responses to identify safety-relevant features.","n19",[],0.95,[145],"Shi et al. 2026: 2 | c12",[],{"side":7,"style":7,"content":148,"node_id":149,"children":150,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":151,"manual_citations":152,"image_display_factor":24},"SAFER-derived feature scores were used to target preference-data poisoning and denoising, with reported safety degradation after limited poisoning and improved safety evaluation after denoising.","n20",[],[145],[],{"side":7,"style":7,"content":154,"node_id":155,"children":156,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":157,"manual_citations":159,"image_display_factor":24},"In SAFER’s sample of 500 features, over 80% received identical safety-relevance scores from GPT-4o and humans, while about 15% differed slightly and none differed substantially.","n21",[],[158],"Shi et al. 2026: 5 | c7",[],[],[],{"side":163,"style":7,"content":164,"node_id":165,"children":166,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":181,"manual_citations":182,"image_display_factor":24},"left","Interpretability and alignment auditing","n22",[167,174],{"side":7,"style":7,"content":168,"node_id":169,"children":170,"collapsed":19,"image_url":7,"confidence":171,"citation_keys":172,"manual_citations":173,"image_display_factor":24},"Feature-level analysis could audit whether safety-relevant concepts influence reward predictions, complementing external explanations that leave latent drivers opaque.","n23",[],0.82,[94],[],{"side":7,"style":7,"content":175,"node_id":176,"children":177,"collapsed":19,"image_url":7,"confidence":29,"citation_keys":178,"manual_citations":180,"image_display_factor":24},"SAE-based interventions aim for greater causal specificity than representation-level controls, but current scaling and evaluation challenges limit their use as dependable audits.","n24",[],[179],"Dung and Mai 2025: 3 | c8",[],[],[],{"side":163,"style":7,"content":184,"node_id":185,"children":186,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":208,"manual_citations":209,"image_display_factor":24},"Interpretability and deception detection","n25",[187,194,201],{"side":7,"style":7,"content":188,"node_id":189,"children":190,"collapsed":19,"image_url":7,"confidence":49,"citation_keys":191,"manual_citations":193,"image_display_factor":24},"Direct evaluation of deceptive alignment is difficult because a strategically deceptive system may act differently during the train-evaluation loop; parameter interpretation and representation engineering are proposed indirect routes.","n26",[],[192],"Ji et al. 2025: 45–46 | c19",[],{"side":7,"style":7,"content":195,"node_id":196,"children":197,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":198,"manual_citations":200,"image_display_factor":24},"The corpus reports laboratory evidence of deceptive outputs, including disabling simulated oversight and then falsely describing the action, but this does not establish reliable SAE-based detection.","n27",[],[199,192],"Bengio et al. 2026: 78 | c15",[],{"side":7,"style":7,"content":202,"node_id":203,"children":204,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":205,"manual_citations":207,"image_display_factor":24},"Scheming evaluations may be compromised when models detect testing and conceal or underplay capabilities, making behavioral evidence an imperfect audit signal.","n28",[],[206],"Shane et al. 2026: 1–2 | c22",[],[],[],{"side":163,"style":7,"content":211,"node_id":212,"children":213,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":234,"manual_citations":235,"image_display_factor":24},"Limitations and thin evidence","n29",[214,220,227],{"side":7,"style":7,"content":215,"node_id":216,"children":217,"collapsed":19,"image_url":7,"confidence":63,"citation_keys":218,"manual_citations":219,"image_display_factor":24},"The Claude 3 Sonnet feature suite is incomplete, and rigorous methods for testing whether recovered features faithfully capture model computations remain lacking.","n30",[],[115],[],{"side":7,"style":7,"content":221,"node_id":222,"children":223,"collapsed":19,"image_url":7,"confidence":20,"citation_keys":224,"manual_citations":226,"image_display_factor":24},"Reliable tools for detecting deceptive and backdoor behavior remain lacking, while mechanistic interpretability faces polysemanticity and scalability challenges.","n31",[],[225],"Ji et al. 2025: 61–62 | c9",[],{"side":7,"style":7,"content":228,"node_id":229,"children":230,"collapsed":19,"image_url":7,"confidence":231,"citation_keys":232,"manual_citations":233,"image_display_factor":24},"The retrieved corpus supports proposed connections between interpretability, auditing, and deception detection more strongly than validated end-to-end audits of frontier models.","n32",[],0.86,[225,179],[],[],[],{"side":163,"style":7,"content":237,"node_id":238,"children":239,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":316,"manual_citations":317,"image_display_factor":24},"Getting started: key papers for newcomers","n33",[240,266,278,296],{"side":7,"style":7,"content":241,"node_id":242,"children":243,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":264,"manual_citations":265,"image_display_factor":24},"Foundations: superposition and feature learning","n34",[244,250,258],{"side":7,"style":7,"content":245,"node_id":246,"children":247,"collapsed":19,"image_url":7,"confidence":29,"citation_keys":248,"manual_citations":249,"image_display_factor":24},"Bricken et al.: begin here for the early demonstration that SAEs can extract interpretable, monosemantic transformer features.","n35",[],[51],[],{"side":7,"style":7,"content":251,"node_id":252,"children":253,"collapsed":19,"image_url":7,"confidence":63,"citation_keys":254,"manual_citations":257,"image_display_factor":24},"Cunningham et al. 2023: start here for the early SAE dictionary-learning account, including the superposition motivation, scalable autointerpretability evaluation, monosemantic feature claims, and causal localization results.","n36",[],[255,256],"Cunningham, Ewart et al. 2023: 1 | c18","Cunningham, Ewart et al. 2023: 9 | c0",[],{"side":7,"style":7,"content":259,"node_id":260,"children":261,"collapsed":19,"image_url":7,"confidence":231,"citation_keys":262,"manual_citations":263,"image_display_factor":24},"Elhage et al.: study this for the superposition problem and the motivation for overcomplete feature bases rather than neuron-by-neuron safety analysis.","n37",[],[31],[],[],[],{"side":7,"style":7,"content":267,"node_id":268,"children":269,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":276,"manual_citations":277,"image_display_factor":24},"Frontier-model features","n38",[270],{"side":7,"style":7,"content":271,"node_id":272,"children":273,"collapsed":19,"image_url":7,"confidence":143,"citation_keys":274,"manual_citations":275,"image_display_factor":24},"Templeton et al.: prioritize this for frontier-scale SAE extraction, abstract and safety-relevant features, causal steering, and explicit faithfulness limitations.","n39",[],[115],[],[],[],{"side":7,"style":7,"content":279,"node_id":280,"children":281,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":294,"manual_citations":295,"image_display_factor":24},"Reward models and alignment data","n40",[282,288],{"side":7,"style":7,"content":283,"node_id":284,"children":285,"collapsed":19,"image_url":7,"confidence":143,"citation_keys":286,"manual_citations":287,"image_display_factor":24},"Shi et al. (SAFER): read this for SAE probing of reward models, contrastive feature scores, safety-feature validation, and preference-data poisoning and denoising.","n41",[],[145,158],[],{"side":7,"style":7,"content":289,"node_id":290,"children":291,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":292,"manual_citations":293,"image_display_factor":24},"Ji et al.: use this survey to situate superposition, scalability, reward-model limitations, deceptive alignment, and the gap between proposed and reliable assurance tools.","n42",[],[31,225],[],[],[],{"side":7,"style":7,"content":297,"node_id":298,"children":299,"collapsed":19,"image_url":7,"confidence":7,"citation_keys":314,"manual_citations":315,"image_display_factor":24},"Auditing, deception, and evaluation limits","n43",[300,307],{"side":7,"style":7,"content":301,"node_id":302,"children":303,"collapsed":19,"image_url":7,"confidence":29,"citation_keys":304,"manual_citations":306,"image_display_factor":24},"Dung and Mai: read this for the comparison between representation-level interventions and SAE-based causal specificity, including scaling and deception-related failure modes.","n44",[],[179,305],"Dung and Mai 2025: 7–9 | c2",[],{"side":7,"style":7,"content":308,"node_id":309,"children":310,"collapsed":19,"image_url":7,"confidence":113,"citation_keys":311,"manual_citations":313,"image_display_factor":24},"Shane et al.: read this for scheming-related evidence and the methodological warning that evaluation can be compromised by situational awareness and capability concealment.","n45",[],[312,206],"Shane et al. 2026: 1 | c14",[],[],[],[],[],[],[],{"nodes":321,"sources":322,"citations":323},46,7,43,{"layout":325},{"node_padding":326,"max_node_width":327,"vertical_spacing":328,"horizontal_spacing":329},2,600,20,40,[331,334,337,340,343,346,348,350,351,353,355,356,358,360,361,363,366,367,369,370,372,374,376,377],{"reference":332,"page_range":333},"Cunningham, Hoagy, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023. “Sparse Autoencoders Find Highly Interpretable Features in Language Models.” Version 3. Preprint, ArXiv. https:\u002F\u002Fdoi.org\u002F10.48550\u002FARXIV.2309.08600.","9",{"reference":335,"page_range":336},"Templeton, Adly, Tom Conerly, Jonathan Marcus, et al. “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.” arXiv:2605.29358. Version 1. Preprint, arXiv, May 28, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2605.29358.","53",{"reference":338,"page_range":339},"Dung, Leonard, and Florian Mai. “AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?” arXiv:2510.11235. Preprint, arXiv, October 13, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2510.11235.","7–9",{"reference":341,"page_range":342},"Ji, Jiaming, Tianyi Qiu, Boyuan Chen, et al. “AI Alignment: A Comprehensive Survey.” arXiv:2310.19852. Preprint, arXiv, April 4, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2310.19852.","52",{"reference":344,"page_range":345},"Shi, Wei, Ziyuan Xie, Sihang Li, and Xiang Wang. “SAFER: Probing Safety in Reward Models with Sparse Autoencoder.” arXiv:2507.00665. Preprint, arXiv, January 30, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2507.00665.","13",{"reference":332,"page_range":347},"5",{"reference":341,"page_range":349},"22",{"reference":344,"page_range":347},{"reference":338,"page_range":352},"3",{"reference":341,"page_range":354},"61–62",{"reference":332,"page_range":352},{"reference":335,"page_range":357},"1",{"reference":344,"page_range":359},"2",{"reference":341,"page_range":347},{"reference":362,"page_range":357},"Shane, Tommy Shaffer, Simon Mylius, and Hamish Hobbs. “Scheming in the Wild: Detecting Real-World AI Scheming Incidents with Open-Source Intelligence.” arXiv:2604.09104. Preprint, arXiv, April 10, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2604.09104.",{"reference":364,"page_range":365},"Bengio, Yoshua, Stephen Clare, and Carina Prunkl. International AI Safety Report 2026. 2026.","78",{"reference":332,"page_range":359},{"reference":341,"page_range":368},"50–51",{"reference":332,"page_range":357},{"reference":341,"page_range":371},"45–46",{"reference":332,"page_range":373},"2–3",{"reference":335,"page_range":375},"1–2",{"reference":362,"page_range":375},{"reference":344,"page_range":378},"16","Guy Zana","This mindmap examines how interpretability, especially sparse autoencoders and monosemantic features, can support AI safety by revealing internal representations, auditing reward models, and investigating deception-related behavior. It surveys foundational work on superposition and feature learning, applications in frontier models and alignment auditing, and methods such as SAFER for identifying safety-relevant features. The map also emphasizes important limitations: incomplete feature coverage, uncertain faithfulness, scalability challenges, and the lack of reliable end-to-end audits. A final section guides newcomers toward key papers on feature learning, frontier models, reward models, and evaluation risks.","2026-08-31T21:58:47.103743Z"]