[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"share-skRN7H":3},{"slug":4,"payload":5},"skRN7H",{"root":6,"stats":377,"title":8,"settings":381,"citations":387,"owner_name":448,"description":449,"published_at":450,"format_version":28},{"side":7,"style":7,"content":8,"node_id":9,"children":10,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":375,"manual_citations":376,"image_display_factor":28},null,"The RLHF Lineage","n0",[11,180,253,282],{"side":7,"style":7,"content":12,"node_id":13,"children":14,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":178,"manual_citations":179,"image_display_factor":28},"Historical progression of preference-based alignment","n1",[15,39,46,54,62,69,99,126,152],{"side":7,"style":7,"content":16,"node_id":17,"children":18,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":37,"manual_citations":38,"image_display_factor":28},"Deep RL from human preferences","n2",[19,29],{"side":7,"style":7,"content":20,"node_id":21,"children":22,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":25,"manual_citations":27,"image_display_factor":28},"Preference-based deep RL treats human judgment as a more practical signal for appropriate behavior than demonstrations or manually specified rewards.","n3",[],false,0.88,[26],"Ji et al. 2025: 24 | c12",[],1,{"side":7,"style":7,"content":30,"node_id":31,"children":32,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":34,"manual_citations":36,"image_display_factor":28},"Pairwise comparisons are converted into a learned scalar reward model, which then supplies the optimization signal for policy learning.","n4",[],0.9,[35],"Ji et al. 2025: 22 | c8",[],[],[],{"side":7,"style":7,"content":40,"node_id":41,"children":42,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":43,"manual_citations":45,"image_display_factor":28},"2017: Christiano et al. establish a deep-RL preference-learning milestone by scaling comparison-based feedback to complex physics and Atari tasks with nonlinear reward models.","n5",[],[44],"Christiano et al. 2023: 2–3 | c5",[],{"side":7,"style":7,"content":47,"node_id":48,"children":49,"collapsed":23,"image_url":7,"confidence":50,"citation_keys":51,"manual_citations":53,"image_display_factor":28},"2020: Stiennon et al. collect pairwise summary preferences, train a supervised reward model, and optimize a summarization policy with PPO against the model’s score, iterating with new policy samples.","n6",[],0.99,[52],"Stiennon, Ouyang et al. 2020: 2 | c21",[],{"side":7,"style":7,"content":55,"node_id":56,"children":57,"collapsed":23,"image_url":7,"confidence":50,"citation_keys":58,"manual_citations":61,"image_display_factor":28},"2022: Ouyang et al.’s InstructGPT trains with supervised fine-tuning, reward-model training on human comparisons, and PPO, with human evaluators preferring its 1.3B model over 175B GPT-3 outputs.","n7",[],[59,60,60],"Ouyang, Wu et al. 2022: 2 | c20","Ouyang, Wu et al. 2022: 3 | c9",[],{"side":7,"style":7,"content":63,"node_id":64,"children":65,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":66,"manual_citations":68,"image_display_factor":28},"Transition: the lineage moves from learning rewards for general deep-RL behavior to using preference models for language-model instruction following and assistant usability.","n8",[],[26,67],"Dung and Mai 2025: 2 | c4",[],{"side":7,"style":7,"content":70,"node_id":71,"children":72,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":97,"manual_citations":98,"image_display_factor":28},"RLHF assistants","n9",[73,81,89],{"side":7,"style":7,"content":74,"node_id":75,"children":76,"collapsed":23,"image_url":7,"confidence":77,"citation_keys":78,"manual_citations":80,"image_display_factor":28},"The standard RLHF pipeline trains a reward or preference model from human-ranked examples, then fine-tunes a policy against that learned signal.","n10",[],0.95,[79],"Croitoru et al. 2026: 13 | c22",[],{"side":7,"style":7,"content":82,"node_id":83,"children":84,"collapsed":23,"image_url":7,"confidence":85,"citation_keys":86,"manual_citations":88,"image_display_factor":28},"Helpful and harmless objectives can conflict: excessive caution can become unhelpful, while excessive helpfulness can enable harmful or toxic content.","n11",[],0.94,[87],"Bai et al. 2022: 4–5 | c18",[],{"side":7,"style":7,"content":90,"node_id":91,"children":92,"collapsed":23,"image_url":7,"confidence":93,"citation_keys":94,"manual_citations":96,"image_display_factor":28},"RLHF commonly adds a KL penalty against the supervised model to limit reward over-optimization and a pretraining loss to preserve broader model performance.","n12",[],0.93,[95],"Ji et al. 2025: 25 | c10",[],[],[],{"side":7,"style":7,"content":100,"node_id":101,"children":102,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":124,"manual_citations":125,"image_display_factor":28},"Constitutional AI","n13",[103,111,118],{"side":7,"style":7,"content":104,"node_id":105,"children":106,"collapsed":23,"image_url":7,"confidence":107,"citation_keys":108,"manual_citations":110,"image_display_factor":28},"Constitutional AI replaces most harmfulness labels with a small set of human-written principles and few-shot examples that form a constitution.","n14",[],0.96,[109],"Bai et al. 2022: 5 | c11",[],{"side":7,"style":7,"content":112,"node_id":113,"children":114,"collapsed":23,"image_url":7,"confidence":107,"citation_keys":115,"manual_citations":117,"image_display_factor":28},"Its supervised phase prompts an initial model to critique and revise responses according to constitutional principles, then fine-tunes on the revised outputs.","n15",[],[116],"Bai et al. 2022: 1 | c14",[],{"side":7,"style":7,"content":119,"node_id":120,"children":121,"collapsed":23,"image_url":7,"confidence":85,"citation_keys":122,"manual_citations":123,"image_display_factor":28},"The RL stage uses model comparisons to train a preference model and reports a harmless, non-evasive assistant that explains objections to harmful requests.","n16",[],[116],[],[],[],{"side":7,"style":7,"content":127,"node_id":128,"children":129,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":150,"manual_citations":151,"image_display_factor":28},"RLAIF: reinforcement learning from AI feedback","n17",[130,136,142],{"side":7,"style":7,"content":131,"node_id":132,"children":133,"collapsed":23,"image_url":7,"confidence":107,"citation_keys":134,"manual_citations":135,"image_display_factor":28},"RLAIF replaces human harmlessness preferences with AI evaluations of response pairs against constitutional principles.","n18",[],[109],[],{"side":7,"style":7,"content":137,"node_id":138,"children":139,"collapsed":23,"image_url":7,"confidence":77,"citation_keys":140,"manual_citations":141,"image_display_factor":28},"AI-generated harmlessness preferences can be mixed with human helpfulness labels to train a hybrid preference model used for RL fine-tuning.","n19",[],[109],[],{"side":7,"style":7,"content":143,"node_id":144,"children":145,"collapsed":23,"image_url":7,"confidence":146,"citation_keys":147,"manual_citations":149,"image_display_factor":28},"A survey reports that RLAIF and RLHF performed almost identically on summarization, while noting that subtle differences remained.","n20",[],0.82,[148],"Ji et al. 2025: 27 | c19",[],[],[],{"side":7,"style":7,"content":153,"node_id":154,"children":155,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":176,"manual_citations":177,"image_display_factor":28},"Direct Preference Optimization","n21",[156,162,169],{"side":7,"style":7,"content":157,"node_id":158,"children":159,"collapsed":23,"image_url":7,"confidence":77,"citation_keys":160,"manual_citations":161,"image_display_factor":28},"DPO directly optimizes the policy from preference pairs, bypassing explicit reward-model training and the multi-stage RLHF pipeline.","n22",[],[95],[],{"side":7,"style":7,"content":163,"node_id":164,"children":165,"collapsed":23,"image_url":7,"confidence":85,"citation_keys":166,"manual_citations":168,"image_display_factor":28},"DPO increases the probability of preferred outputs relative to rejected outputs while anchoring the policy to a reference distribution through a coefficient-controlled ratio objective.","n23",[],[167],"Croitoru et al. 2026: 4–5 | c23",[],{"side":7,"style":7,"content":170,"node_id":171,"children":172,"collapsed":23,"image_url":7,"confidence":173,"citation_keys":174,"manual_citations":175,"image_display_factor":28},"Follow-on preference-optimization work studies divergence choices and overfitting, including f-DPO, broader pairwise objectives, and IPO.","n24",[],0.86,[95],[],[],[],[],[],{"side":7,"style":7,"content":181,"node_id":182,"children":183,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":251,"manual_citations":252,"image_display_factor":28},"Known failure modes of reward models and human feedback","n25",[184,192,201,208,216,222,229,235,242],{"side":7,"style":7,"content":185,"node_id":186,"children":187,"collapsed":23,"image_url":7,"confidence":188,"citation_keys":189,"manual_citations":191,"image_display_factor":28},"Reward hacking occurs when an agent optimizes a proxy or implementation of the objective while violating the designer’s intended goal.","n26",[],0.97,[190],"Amodei et al. 2016: 7 | c17",[],{"side":7,"style":7,"content":193,"node_id":194,"children":195,"collapsed":23,"image_url":7,"confidence":196,"citation_keys":197,"manual_citations":200,"image_display_factor":28},"Misspecified proxy rewards can produce high metric performance while falling short of human standards, with proxy reward rising as true reward declines.","n27",[],0.92,[198,199],"Ji et al. 2025: 5 | c2","Che and Wu 2026: 2 | c26",[],{"side":7,"style":7,"content":202,"node_id":203,"children":204,"collapsed":23,"image_url":7,"confidence":107,"citation_keys":205,"manual_citations":207,"image_display_factor":28},"Preference-model scores can overestimate crowdworker judgments because of distribution shift, possible nontransitivity in score aggregation, and model miscalibration.","n28",[],[206],"Bai et al. 2022: 45–46 | c15",[],{"side":7,"style":7,"content":209,"node_id":210,"children":211,"collapsed":23,"image_url":7,"confidence":212,"citation_keys":213,"manual_citations":215,"image_display_factor":28},"Human feedback is an imperfect proxy for beneficial behavior and is constrained by human error and bias.","n29",[],0.91,[214],"Bengio et al. 2026: 124 | c6",[],{"side":7,"style":7,"content":217,"node_id":218,"children":219,"collapsed":23,"image_url":7,"confidence":85,"citation_keys":220,"manual_citations":221,"image_display_factor":28},"Preference models trained primarily for helpfulness or harmlessness can perform poorly on the other objective, reflecting conflict in the feedback target.","n30",[],[87],[],{"side":7,"style":7,"content":223,"node_id":224,"children":225,"collapsed":23,"image_url":7,"confidence":85,"citation_keys":226,"manual_citations":228,"image_display_factor":28},"RLHF may fail when humans cannot reliably evaluate performance beyond their own expertise, especially when evaluation is not easier than generation.","n31",[],[227],"Dung and Mai 2025: 7 | c24",[],{"side":7,"style":7,"content":230,"node_id":231,"children":232,"collapsed":23,"image_url":7,"confidence":173,"citation_keys":233,"manual_citations":234,"image_display_factor":28},"Failure-mode analyses identify possible vulnerabilities involving deceptive alignment, emergent misalignment, and dangerous out-of-distribution generalization, while describing the analysis as exploratory and uncertain.","n32",[],[227,227],[],{"side":7,"style":7,"content":236,"node_id":237,"children":238,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":239,"manual_citations":241,"image_display_factor":28},"Preference-data poisoning can significantly bias reward models, and cited work reports harmful DPO backdoors from very small poisoned-data fractions.","n33",[],[240],"Shi et al. 2026: 13 | c13",[],{"side":7,"style":7,"content":243,"node_id":244,"children":245,"collapsed":23,"image_url":7,"confidence":188,"citation_keys":246,"manual_citations":250,"image_display_factor":28},"Gao et al. show that optimizing an imperfect proxy reward can reduce gold reward, with proxy–gold divergence depending on optimization method and reward-model scale.","n34",[],[247,248,249],"Gao, Schulman and Hilton 2022: 1 | c0","Gao, Schulman and Hilton 2022: 7 | c7","Gao, Schulman and Hilton 2022: 8 | c16",[],[],[],{"side":254,"style":7,"content":255,"node_id":256,"children":257,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":280,"manual_citations":281,"image_display_factor":28},"left","Research landscape for newcomers","n35",[258,266,274],{"side":7,"style":7,"content":259,"node_id":260,"children":261,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":262,"manual_citations":265,"image_display_factor":28},"Scalable oversight studies how AI systems can help supervise more capable systems, including automated alignment researchers and nested or bootstrapped oversight.","n36",[],[263,264],"Engels et al. 2025: 15 | c3","Bai et al. 2022: 2–3 | c27",[],{"side":7,"style":7,"content":267,"node_id":268,"children":269,"collapsed":23,"image_url":7,"confidence":270,"citation_keys":271,"manual_citations":273,"image_display_factor":28},"Evaluation research should test whether preference-model scores track human judgments under distribution shift, adversarial prompts, and changing model capability.","n37",[],0.89,[206,272],"Bai et al. 2022: 67 | c1",[],{"side":7,"style":7,"content":275,"node_id":276,"children":277,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":278,"manual_citations":279,"image_display_factor":28},"Reward-modeling research focuses on representing comparison feedback, detecting misspecification and hacking, and improving the quality and robustness of preference data.","n38",[],[35,240],[],[],[],{"side":254,"style":7,"content":283,"node_id":284,"children":285,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":373,"manual_citations":374,"image_display_factor":28},"Getting started: key papers for newcomers","n39",[286,312,336,348,367],{"side":7,"style":7,"content":287,"node_id":288,"children":289,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":310,"manual_citations":311,"image_display_factor":28},"Foundations: preference learning and reward modeling","n40",[290,296,304],{"side":7,"style":7,"content":291,"node_id":292,"children":293,"collapsed":23,"image_url":7,"confidence":146,"citation_keys":294,"manual_citations":295,"image_display_factor":28},"Christiano et al. (2017): start here for the preference-based deep-RL framing, where comparisons become a learned reward signal for policy learning.","n41",[],[26,35],[],{"side":7,"style":7,"content":297,"node_id":298,"children":299,"collapsed":23,"image_url":7,"confidence":300,"citation_keys":301,"manual_citations":303,"image_display_factor":28},"Stiennon et al. (2020), Learning to Summarize from Human Feedback: read it for the primary preference-pair, reward-model, PPO, and iterative-policy pipeline, plus its human-labeler overoptimization result.","n42",[],0.98,[52,302],"Stiennon, Ouyang et al. 2020: 7–8 | c25",[],{"side":7,"style":7,"content":305,"node_id":306,"children":307,"collapsed":23,"image_url":7,"confidence":33,"citation_keys":308,"manual_citations":309,"image_display_factor":28},"Amodei et al. (2016): read for the foundational distinction between formal reward optimization and the designer’s intended goal.","n43",[],[190],[],[],[],{"side":7,"style":7,"content":313,"node_id":314,"children":315,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":334,"manual_citations":335,"image_display_factor":28},"Assistant alignment: RLHF and preference optimization","n44",[316,322,328],{"side":7,"style":7,"content":317,"node_id":318,"children":319,"collapsed":23,"image_url":7,"confidence":85,"citation_keys":320,"manual_citations":321,"image_display_factor":28},"Bai et al. (2022), Training a Helpful and Harmless Assistant: study the preference-modeling and RLHF pipeline, plus the helpfulness-harmlessness tradeoff.","n45",[],[87],[],{"side":7,"style":7,"content":323,"node_id":324,"children":325,"collapsed":23,"image_url":7,"confidence":50,"citation_keys":326,"manual_citations":327,"image_display_factor":28},"Ouyang et al. (2022), Training Language Models to Follow Instructions with Human Feedback: read it for the primary SFT, reward-model, and PPO recipe and the finding that 1.3B InstructGPT outputs were preferred to 175B GPT-3 outputs.","n46",[],[59,60,60],[],{"side":7,"style":7,"content":329,"node_id":330,"children":331,"collapsed":23,"image_url":7,"confidence":24,"citation_keys":332,"manual_citations":333,"image_display_factor":28},"Rafailov et al. (2023), Direct Preference Optimization: study how preference pairs can directly optimize a reference-anchored policy without an explicit reward model.","n47",[],[95,167],[],[],[],{"side":7,"style":7,"content":337,"node_id":338,"children":339,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":346,"manual_citations":347,"image_display_factor":28},"Constitutional AI and AI feedback","n48",[340],{"side":7,"style":7,"content":341,"node_id":342,"children":343,"collapsed":23,"image_url":7,"confidence":107,"citation_keys":344,"manual_citations":345,"image_display_factor":28},"Bai et al. (2022), Constitutional AI: study constitutional principles, self-critique and revision, and the transition from human harmlessness labels to RLAIF.","n49",[],[116,109],[],[],[],{"side":7,"style":7,"content":349,"node_id":350,"children":351,"collapsed":23,"image_url":7,"confidence":7,"citation_keys":365,"manual_citations":366,"image_display_factor":28},"Safety, oversight, and evaluation","n50",[352,359],{"side":7,"style":7,"content":353,"node_id":354,"children":355,"collapsed":23,"image_url":7,"confidence":356,"citation_keys":357,"manual_citations":358,"image_display_factor":28},"Engels et al. (2025), Scaling Laws for Scalable Oversight: follow this direction for automated oversight, nested supervision, and capability-relative evaluation.","n51",[],0.8,[263],[],{"side":7,"style":7,"content":360,"node_id":361,"children":362,"collapsed":23,"image_url":7,"confidence":173,"citation_keys":363,"manual_citations":364,"image_display_factor":28},"Ji et al. (2025), AI Alignment: A Comprehensive Survey: use it to connect RLHF, reward modeling, DPO, scalable oversight, and failure-mode research.","n52",[],[26,95],[],[],[],{"side":7,"style":7,"content":368,"node_id":369,"children":370,"collapsed":23,"image_url":7,"confidence":188,"citation_keys":371,"manual_citations":372,"image_display_factor":28},"Gao et al. (2022), Scaling Laws for Reward Model Overoptimization: read it for proxy-versus-gold reward divergence under reinforcement learning and best-of-N optimization, including scaling with reward-model parameters.","n53",[],[247,249],[],[],[],[],[],{"nodes":378,"sources":379,"citations":380},54,14,58,{"layout":382},{"node_padding":383,"max_node_width":384,"vertical_spacing":385,"horizontal_spacing":386},2,600,20,40,[388,391,394,397,400,403,406,409,411,413,416,418,420,422,425,426,428,430,432,434,436,437,439,441,442,443,445,447],{"reference":389,"page_range":390},"Gao, Leo, John Schulman, and Jacob Hilton. 2022. “Scaling Laws for Reward Model Overoptimization.” Version 1. Preprint, ArXiv. https:\u002F\u002Fdoi.org\u002F10.48550\u002FARXIV.2210.10760.","1",{"reference":392,"page_range":393},"Bai, Yuntao, Andy Jones, Kamal Ndousse, et al. “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.” arXiv:2204.05862. Preprint, arXiv, April 12, 2022. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2204.05862.","67",{"reference":395,"page_range":396},"Ji, Jiaming, Tianyi Qiu, Boyuan Chen, et al. “AI Alignment: A Comprehensive Survey.” arXiv:2310.19852. Preprint, arXiv, April 4, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2310.19852.","5",{"reference":398,"page_range":399},"Engels, Joshua, David D. Baek, Subhash Kantamneni, and Max Tegmark. “Scaling Laws For Scalable Oversight.” arXiv:2504.18530. Version 3. Preprint, arXiv, October 27, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2504.18530.","15",{"reference":401,"page_range":402},"Dung, Leonard, and Florian Mai. “AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?” arXiv:2510.11235. Preprint, arXiv, October 13, 2025. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2510.11235.","2",{"reference":404,"page_range":405},"Christiano, Paul, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. “Deep Reinforcement Learning from Human Preferences.” arXiv:1706.03741. Preprint, arXiv, February 17, 2023. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.1706.03741.","2–3",{"reference":407,"page_range":408},"Bengio, Yoshua, Stephen Clare, and Carina Prunkl. International AI Safety Report 2026. 2026.","124",{"reference":389,"page_range":410},"7",{"reference":395,"page_range":412},"22",{"reference":414,"page_range":415},"Ouyang, Long, Jeff Wu, Xu Jiang, et al. 2022. “Training Language Models to Follow Instructions with Human Feedback.” Version 1. Preprint, ArXiv. https:\u002F\u002Fdoi.org\u002F10.48550\u002FARXIV.2203.02155.","3",{"reference":395,"page_range":417},"25",{"reference":419,"page_range":396},"Bai, Yuntao, Saurav Kadavath, Sandipan Kundu, et al. “Constitutional AI: Harmlessness from AI Feedback.” arXiv:2212.08073. Preprint, arXiv, December 15, 2022. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2212.08073.",{"reference":395,"page_range":421},"24",{"reference":423,"page_range":424},"Shi, Wei, Ziyuan Xie, Sihang Li, and Xiang Wang. “SAFER: Probing Safety in Reward Models with Sparse Autoencoder.” arXiv:2507.00665. Preprint, arXiv, January 30, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2507.00665.","13",{"reference":419,"page_range":390},{"reference":392,"page_range":427},"45–46",{"reference":389,"page_range":429},"8",{"reference":431,"page_range":410},"Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. “Concrete Problems in AI Safety.” arXiv:1606.06565. Preprint, arXiv, July 25, 2016. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.1606.06565.",{"reference":392,"page_range":433},"4–5",{"reference":395,"page_range":435},"27",{"reference":414,"page_range":402},{"reference":438,"page_range":402},"Stiennon, Nisan, Long Ouyang, Jeff Wu, et al. 2020. “Learning to Summarize from Human Feedback.” Version 3. Preprint, ArXiv. https:\u002F\u002Fdoi.org\u002F10.48550\u002FARXIV.2009.01325.",{"reference":440,"page_range":424},"Croitoru, Florinel-Alin, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe, and Mubarak Shah. “Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation.” arXiv:2602.13055. Version 1. Preprint, arXiv, February 13, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2602.13055.",{"reference":440,"page_range":433},{"reference":401,"page_range":410},{"reference":438,"page_range":444},"7–8",{"reference":446,"page_range":402},"Che, Tong, and Rui Wu. “Greed Is Learned: Visible Incentives as Reward-Hacking Triggers.” arXiv:2606.16914. Preprint, arXiv, June 15, 2026. https:\u002F\u002Fdoi.org\u002F10.48550\u002FarXiv.2606.16914.",{"reference":419,"page_range":405},"Guy Zana","How did \"ask humans which output they prefer\" become the way frontier models are aligned? This map follows the lineage from preference-based RL in 2017 through summarization, InstructGPT, Constitutional AI, RLAIF, and DPO, as a historical progression. The second half is the honest part: reward hacking, overoptimization scaling laws, preference-data poisoning, and the limits of human feedback. Citation-backed throughout, with primary papers in the newcomer branch.","2026-08-25T15:33:54.909779Z"]