Samuel Nellessen

© Samuel Nellessen
  • Master’s Student
Samuel Nellessen is an AI security researcher working on automated red-teaming and adversarial evaluation of LLM agents. He is currently an independent Foresight AI Safety Grantee, a student researcher in Tal Kachman’s lab at Radboud University, and an incoming LASR Labs Fellow. His research focuses on agentic AI security, tool-use vulnerabilities, adversarial training, and mechanistic interpretability of refusal behavior. He has led work on Tag-Along Attacks and Slingshot, a reinforcement learning framework for discovering verifiable agent-to-agent jailbreaks. His recent paper, David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning with Tal Kachman (2026), formalizes this threat model and shows strong transfer to frontier models. He has also contributed to UK AISI’s inspect_ai and served as a technical advisor for the Foresight AI Safety Grant Program.