<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Center for Human-Compatible Artificial Intelligence – News</title><description>Center for Human-Compatible AI is building exceptional AI for humanity</description><link>https://humancompatible.ai/</link><item><title>Learning to Coordinate with Experts</title><link>https://humancompatible.ai/news/2025/03/07/learning-to-coordinate-with-experts/</link><guid isPermaLink="true">https://humancompatible.ai/news/2025/03/07/learning-to-coordinate-with-experts/</guid><description>Khanh Nguyen, Benjamin Plaut, Tu Trinh, and Mohamad Danesh introduce a fundamental coordination problem called Learning to Yield and Request Control (YRC), where the objective is to learn a strategy that determines when to act autonomously and when to seek expert assistance. They build an open-source benchmark featuring diverse domains, propose a novel validation approach, and investigate the performance of various learning methods across diverse environments, yielding insights that can guide future research.</description><pubDate>Fri, 07 Mar 2025 14:16:30 GMT</pubDate></item><item><title>Computational Frameworks for Human Care</title><link>https://humancompatible.ai/news/2025/02/20/computational-frameworksfor-human-care/</link><guid isPermaLink="true">https://humancompatible.ai/news/2025/02/20/computational-frameworksfor-human-care/</guid><description>Brian Christian, CHAI Affiliate, has published an article titled “&lt;a href=&quot;https://www.amacad.org/sites/default/files/publication/downloads/daedalus_wi25_12_christian.pdf&quot; data-type=&quot;link&quot; data-id=&quot;https://www.amacad.org/sites/default/files/publication/downloads/daedalus_wi25_12_christian.pdf&quot;&gt;Computational Frameworks for Human Care&lt;/a&gt;” in the most recent issue of Daedalus, the journal of the American Academy of Arts and Sciences. In it, Christian traces how AI alignment has progressed from simple reward mechanisms toward care-like relationships, revealing both the potential and limitations of machine caregiving while deepening our understanding of human care itself. The issue is titled “The Social Science of Caregiving” and was co-edited by CHAI Affiliate Alison Gopnik.</description><pubDate>Thu, 20 Feb 2025 14:08:22 GMT</pubDate></item><item><title>A Practical Definition of Political Neutrality for AI</title><link>https://humancompatible.ai/news/2025/02/04/a-practical-definition-of-political-neutrality-for-ai/</link><guid isPermaLink="true">https://humancompatible.ai/news/2025/02/04/a-practical-definition-of-political-neutrality-for-ai/</guid><description>NEW: Our current research project to build &lt;a href=&quot;https://docs.google.com/document/d/19haXfSeQtTjdVLUbba1z5GRhW12rj0ipR6lM5S3G0o4/edit?usp=sharing&quot;&gt;political neutrality evaluations&lt;/a&gt;.</description><pubDate>Tue, 04 Feb 2025 15:00:26 GMT</pubDate></item><item><title>RvS: What is Essential for Offline RL via Supervised Learning?</title><link>https://humancompatible.ai/news/2025/01/18/rvs-what-is-essential-for-offline-rl-via-supervised-learning/</link><guid isPermaLink="true">https://humancompatible.ai/news/2025/01/18/rvs-what-is-essential-for-offline-rl-via-supervised-learning/</guid><description>Scott Emmons, PhD student, was an author on “RvS: What is Essential for Offline RL via Supervised Learning?”</description><pubDate>Sat, 18 Jan 2025 09:55:12 GMT</pubDate></item><item><title>Getting By Goal Misgeneralization With a Little Help From a Mentor</title><link>https://humancompatible.ai/news/2024/12/25/getting-by-goal-misgeneralization-with-a-little-help-from-a-mentor-2/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/12/25/getting-by-goal-misgeneralization-with-a-little-help-from-a-mentor-2/</guid><description>“Tu Trinh, Ben Plaut, Khanh Nguyen, and Mohamad Danesh wrote the paper, “Getting By Goal Misgeneralization With a Little Help From a Mentor.” This paper explores whether goal misgeneralization can be mitigated by allowing an agent to ask for help when it is uncertain. The answer is mostly yes, although our current methods have substantial weaknesses and there are lots of interesting avenues for future work.”Tu Trinh, Ben Plaut, Khanh Nguyen, and Mohamad Danesh wrote the paper, This paper explores whether goal misgeneralization can be mitigated by allowing an agent to ask for help when it is uncertain. The answer is mostly yes, although our current methods have substantial weaknesses and there are lots of interesting avenues for future work.</description><pubDate>Wed, 25 Dec 2024 09:36:25 GMT</pubDate></item><item><title>Linear Probe Penalties Reduce LLM Sycophancy</title><link>https://humancompatible.ai/news/2024/12/14/linear-probe-penalties-reduce-llm-sycophancy/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/12/14/linear-probe-penalties-reduce-llm-sycophancy/</guid><description>Visiting ETH MsC student Henry Papadatos and supervising CHAI PhD student Rachel Freedman publish an article “Linear Probe Penalties Reduce LLM Sycophancy” at the NeurIPS SoLaR workshop. The paper demonstrates a generalizable methodology for reducing unwanted LLM behaviors that are not sufficiently disincentivized by RLHF fine-tuning</description><pubDate>Sat, 14 Dec 2024 09:34:04 GMT</pubDate></item><item><title>Rachel Freedman selected as inaugural Cooperative AI Fellow</title><link>https://humancompatible.ai/news/2024/11/30/rachel-freedman-selected-as-inaugural-cooperative-ai-fellow/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/11/30/rachel-freedman-selected-as-inaugural-cooperative-ai-fellow/</guid><description>Rachel Freedman, PhD Student, has been selected as one of the fellows for &lt;a href=&quot;https://www.cooperativeai.com/phd-fellowship/2025&quot; data-type=&quot;link&quot; data-id=&quot;https://www.cooperativeai.com/phd-fellowship/2025&quot;&gt;Cooperative AI’s PhD Fellow Program&lt;/a&gt;.</description><pubDate>Sat, 30 Nov 2024 09:30:08 GMT</pubDate></item><item><title>Representative Social Choice: From Learning Theory to AI Alignment</title><link>https://humancompatible.ai/news/2024/11/12/representative-social-choice-from-learning-theory-to-ai-alignment/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/11/12/representative-social-choice-from-learning-theory-to-ai-alignment/</guid><description>Tianyi Qiu, CHAI Intern, wrote this paper which was accepted by NeurIPS 2024 Pluralistic Alignment Workshop. &lt;a href=&quot;https://arxiv.org/pdf/2410.23953&quot; data-type=&quot;link&quot; data-id=&quot;https://arxiv.org/pdf/2410.23953&quot;&gt;Here&lt;/a&gt; is the link to the paper.</description><pubDate>Tue, 12 Nov 2024 12:55:26 GMT</pubDate></item><item><title>Getting By Goal Misgeneralization With a Little Help From a Mentor</title><link>https://humancompatible.ai/news/2024/10/10/getting-by-goal-misgeneralization-with-a-little-help-from-a-mentor/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/10/10/getting-by-goal-misgeneralization-with-a-little-help-from-a-mentor/</guid><description>Khanh Nguyen, Mohamad Danesh, Ben Plaut, and Alina Trinh wrote &lt;a href=&quot;https://arxiv.org/pdf/2410.21052&quot; data-type=&quot;link&quot; data-id=&quot;https://arxiv.org/pdf/2410.21052&quot;&gt;this paper&lt;/a&gt; which was presented at Towards Safe &amp;amp; Trustworthy Agents Workshop at NeurIPS 2024.</description><pubDate>Thu, 10 Oct 2024 09:19:23 GMT</pubDate></item><item><title>Language-Guided World Models: A Model-Based Approach to AI Control</title><link>https://humancompatible.ai/news/2024/09/18/language-guided-world-models-a-model-based-approach-to-ai-control/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/09/18/language-guided-world-models-a-model-based-approach-to-ai-control/</guid><description>Khanh Nguyen, CHAI Postdoctoral Fellow, published a paper at the Fourth International Combined Workshop on Spatial Language Understanding and Grounded Communication for Robotics (ACL 2024).</description><pubDate>Wed, 18 Sep 2024 14:32:50 GMT</pubDate></item><item><title>“The Alignment Problem” Wins Xingdu Book Award</title><link>https://humancompatible.ai/news/2024/08/29/the-alignment-problem-wins-xingdu-book-award/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/08/29/the-alignment-problem-wins-xingdu-book-award/</guid><description>Brian Christian’s book “The Alignment Problem” was announced as the sole winner in the New Knowledge Category for Imported Editions at the Xingdu Book Award ceremony in China. The Chinese translation was published this past year by Hunan Science &amp;amp; Technology Press.</description><pubDate>Thu, 29 Aug 2024 08:50:23 GMT</pubDate></item><item><title>Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback</title><link>https://humancompatible.ai/news/2024/08/07/social-choice-should-guide-ai-alignment-in-dealing-with-diverse-human-feedback/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/08/07/social-choice-should-guide-ai-alignment-in-dealing-with-diverse-human-feedback/</guid><description>Rachel Feedman, CHAI Phd Student, and Wes Holliday, CHAI Affiliate, published a paper at the International Conference on Machine Learning</description><pubDate>Wed, 07 Aug 2024 12:08:47 GMT</pubDate></item><item><title>AI Alignment with Changing and Influenceable Reward Functions</title><link>https://humancompatible.ai/news/2024/07/23/ai-alignment-with-changing-and-influenceable-reward-functions/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/07/23/ai-alignment-with-changing-and-influenceable-reward-functions/</guid><description>CHAI Researchers, Micah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell, and Anca Dragan, wrote the paper, “AI Alignment with Changing and Influenceable Reward Functions” which was accepted to ICML.</description><pubDate>Tue, 23 Jul 2024 12:04:59 GMT</pubDate></item><item><title>Forget deepfake videos. Text and voice are this election’s true AI threat.</title><link>https://humancompatible.ai/news/2024/07/08/forget-deepfake-videos-text-and-voice-are-this-elections-true-ai-threat/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/07/08/forget-deepfake-videos-text-and-voice-are-this-elections-true-ai-threat/</guid><description>Jonathan Stray, Senior Scientist at CHAI, and Jessica Alter, tech entrepreneur and co-founder of &lt;a href=&quot;https://www.techforcampaigns.org/&quot;&gt;Tech for Campaigns&lt;/a&gt;, wrote an op-ed for The Hill regarding the risks posed by AI in this current election cycle.</description><pubDate>Mon, 08 Jul 2024 11:59:38 GMT</pubDate></item><item><title>Mitigating Partial Observability in Decision Processes via the Lambda Discrepancy</title><link>https://humancompatible.ai/news/2024/06/28/mitigating-partial-observability-in-decision-processes-via-the-lambda-discrepancy/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/06/28/mitigating-partial-observability-in-decision-processes-via-the-lambda-discrepancy/</guid><description>This paper investigates fundamental concepts related to detecting and mitigating partial observability by measuring misalignment between value function estimates. The paper was presented at the “Finding the Frame” workshop at RLC 2024 and the “Foundations of Reinforcement Learning and Control” workshop at ICML 2024.</description><pubDate>Fri, 28 Jun 2024 10:50:36 GMT</pubDate></item><item><title>8th Annual CHAI Workshop</title><link>https://humancompatible.ai/news/2024/06/18/8th-annual-chai-workshop/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/06/18/8th-annual-chai-workshop/</guid><description>CHAI held its 8th annual workshop at Asilomar Conference Grounds from June 13th to June 16th in Pacific Grove. The workshop had over 200 attendees which was the highest attendance to date. The workshop featured over 60 speakers and panelists and covered a wide array of topics from Societal Effects of AI to Adversarial Robustness.</description><pubDate>Tue, 18 Jun 2024 10:42:38 GMT</pubDate></item><item><title>When Code Isn’t Law: Rethinking Regulation for Artificial Intelligence</title><link>https://humancompatible.ai/news/2024/06/05/when-code-isnt-law-rethinking-regulation-for-artificial-intelligence/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/06/05/when-code-isnt-law-rethinking-regulation-for-artificial-intelligence/</guid><description>Brian Judge, Mark Nitzberg, and Stuart Russell wrote an &lt;a href=&quot;https://academic.oup.com/policyandsociety/advance-article/doi/10.1093/polsoc/puae020/7684910&quot; data-type=&quot;link&quot; data-id=&quot;https://academic.oup.com/policyandsociety/advance-article/doi/10.1093/polsoc/puae020/7684910&quot;&gt;article&lt;/a&gt; that was featured in Oxford Academic’s Policy and Society.</description><pubDate>Wed, 05 Jun 2024 11:46:09 GMT</pubDate></item><item><title>Committing to the wrong artificial delegate in a collective-risk dilemma is better than directly committing mistakes</title><link>https://humancompatible.ai/news/2024/05/13/committing-to-the-wrong-artificial-delegate-in-a-collective-risk-dilemma-is-better-than-directly-committing-mistakes/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/05/13/committing-to-the-wrong-artificial-delegate-in-a-collective-risk-dilemma-is-better-than-directly-committing-mistakes/</guid><description>New research from computer scientists Inês Terrucha, Elias Fernández Domingos, Pieter Simoens, and Tom Lenaerts at the Vrije Universiteit Brussel, Université Libre de Bruxelles, and UC Berkeley’s Center for Human-Compatible AI</description><pubDate>Mon, 13 May 2024 09:00:00 GMT</pubDate></item><item><title>Reinforcement Learning with Human Feedback and Active Teacher Selection (RLHF and ATS)</title><link>https://humancompatible.ai/news/2024/04/30/reinforcement-learning-with-human-feedback-and-active-teacher-selection-rlhf-and-ats/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/04/30/reinforcement-learning-with-human-feedback-and-active-teacher-selection-rlhf-and-ats/</guid><description>CHAI PhD graduate student, Rachel Freedman gave a presentation at Stanford University on critical new developments in AI safety, focusing on problems and potential solutions with Reinforcement Learning from Human Feedback (RLHF).</description><pubDate>Tue, 30 Apr 2024 09:00:00 GMT</pubDate></item><item><title>Reinforcement Learning Safety Workshop (RLSW) @ RLC 2024</title><link>https://humancompatible.ai/news/2024/04/15/reinforcement-learning-safety-workshop-rlsw-rlc-2024/</link><guid isPermaLink="true">https://humancompatible.ai/news/2024/04/15/reinforcement-learning-safety-workshop-rlsw-rlc-2024/</guid><description>&lt;strong&gt;Important Dates&lt;/strong&gt;&lt;br&gt;Paper submission deadline: &lt;strong&gt;May 10, 2024 (AoE)&lt;/strong&gt;&lt;br&gt;Paper acceptance notification: &lt;strong&gt;May 23, 2024&lt;/strong&gt;</description><pubDate>Mon, 15 Apr 2024 09:00:00 GMT</pubDate></item></channel></rss>