# Hi I'm Lev! <span class="rightimg"><span class="tinyimg"> ![[headshot.jpg]] </span></span> I'm a Member of Techincal Staff at METR. I was previously an Astra Fellow with Owain Evans and a graduate student at the Univesity of Toronto with supervised by [Sheila McIlraith](https://www.cs.toronto.edu/~sheila/) and [Roger Grosse](https://www.cs.toronto.edu/~rgrosse/). --- > [!example] Papers & Preprints > >* Jorio Cocola, **Lev E McKinney**, Harry Mayne, Jan Betley, and Owain Evans. Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble. arXiv preprint arXiv:2609.10883, 2026. URL https://arxiv.org/abs/2609.10883 >* Jason R Brown, Patrick Leask, and **Lev E McKinney**. Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment. arXiv preprint arXiv:2606.31591, 2026. URL https://arxiv.org/abs/2606.31591 >* Harry Mayne\*, **Lev E McKinney**\*, Jan Dubiński, Adam Karvonen, James Chua, and Owain Evans. Negation Neglect: When Models Fail to Learn Negations in Training. arXiv preprint arXiv:2605.13829, 2026. URL https://arxiv.org/abs/2605.13829 >* Zheng-Xin Yong, Parv Mahajan, Andy Wang, Ida Caspary, Yernat Yestekov, Zora Che, Mosh Levy, Elle Najt, Dennis Murphy, Prashant Kulkarni, **Lev E McKinney**, Kei Nishimura-Gasparian, Ram Potham, Aengus Lynch, and Michael L Chen. An Independent Safety Evaluation of Kimi K2.5. arXiv preprint arXiv:2604.03121, 2026. URL https://arxiv.org/abs/2604.03121 >* Matthew Kowal, Gonçalo Paulo, Louis Jaburi, Tom Tseng, **Lev E McKinney**, Stefan Heimersheim, Aaron David Tucker, Adam Gleave, and Kellin Pelrine. Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution. arXiv preprint arXiv:2602.14869, 2026. URL https://arxiv.org/abs/2602.14869 >* **Lev E McKinney**, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A McIlraith, and Roger Baker Grosse. Gauss-Newton Unlearning for the LLM Era. arXiv preprint arXiv:2602.10568, 2026. URL https://arxiv.org/abs/2602.10568. An earlier version appeared in the _ICML 2025 Workshop on Machine Unlearning for Generative AI_: https://openreview.net/forum?id=VFfttnDvW6 >* Zora Che, Stephen Casper, Robert Kirk, Anirudh Satheesh, Stewart Slocum, **Lev E McKinney**, Rohit Gandikota, Aidan Ewart, Domenic Rosati, Zichu Wu, Zikui Cai, Bilal Chughtai, Yarin Gal, Furong Huang, and Dylan Hadfield-Menell. Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities. _Transactions on Machine Learning Research_, 2025. URL https://arxiv.org/abs/2502.05209 >* Nora Belrose, Zach Furman, **Lev E McKinney**, Logan Smith, Danny Halawi, Igor Ostrovsky, Stella Biderman, and Jacob Steinhardt. Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv preprint arXiv:2303.08112, 2023. URL https://arxiv.org/abs/2303.08112 >* **Lev E McKinney**\*, Yawen Duan\*, David Krueger, and Adam Gleave. On the Fragility of Learned Reward Functions. In _Deep Reinforcement Learning Workshop NeurIPS 2022_, 2022. URL https://arxiv.org/abs/2301.03652 >* Keiran Paster\*, **Lev E McKinney**\*, Sheila A McIlraith, and Jimmy Ba. BLAST: Latent Dynamics Models from Bootstrapping. In _Deep RL Workshop NeurIPS 2021_, 2021. URL https://openreview.net/forum?id=VwA_hKnX_kR > > \* denotes equal contribution. > [!info]- Contact Info > levmckinney [at] cs [dot] toronto [dot] edu