DeepMind Safety Research·May 29Testing Gemini models for scheming tendenciesBy Victoria Krakovna, David Lindner, Sebastian Farquhar and Rohin Shah
DeepMind Safety Research·Apr 1Predicting When RL Training Breaks Chain-of-Thought MonitorabilityWhy RL training teaches models to hide their reasoning, and a conceptual framework to predict when it happens.
DeepMind Safety Research·Nov 3, 2025Consistency Training Could Help Limit Sycophancy and JailbreaksAuthors: Alex Irpan* and Alex Turner*, Mark Kurzeja, David Elson, and Rohin ShahA response icon1A response icon1
DeepMind Safety Research·Jul 8, 2025Evaluating and monitoring for AI schemingBy Victoria Krakovna, Scott Emmons, Erik Jenner, Mary Phuong, Lewis Ho, and Rohin Shah
DeepMind Safety Research·Apr 8, 2025An Approach to Technical AGI Safety and SecurityWe have written a paper on our approach to technical AGI safety and security. This post is a copy of the extended abstract, which…
DeepMind Safety Research·Mar 26, 2025Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research…Lewis Smith*, Sen Rajamanoharan*, Arthur Conmy, Callum McDougall, Janos Kramar, Tom Lieberum, Rohin Shah, Neel Nanda
DeepMind Safety Research·Feb 14, 2025Introducing our short course on AGI safetyWe are excited to release a short course on AGI safety for students, researchers and professionals interested in this topic. The course…A response icon4A response icon4
DeepMind Safety Research·Jan 31, 2025Steering Gemini using BIDPO vectorsBy Alex Turner and Mark KurzejaA response icon1A response icon1
DeepMind Safety Research·Jan 23, 2025MONA: A method for addressing multi-step reward hackingMONA enhances safety when we train an AI system to perform some task that takes multiple steps. Training an AI with MONA reduces its …
DeepMind Safety Research·Dec 23, 2024Human-AI Complementarity: A Goal for Amplified OversightHow do we ensure humans can continue to oversee increasingly powerful AI systems? We argue that achieving human-AI complementarity is key.