Diogo Cruz
AI Safety Researcher
I'm a technical researcher focused on AI safety, with a physics PhD in quantum computing. My research centers on agent evaluations—understanding safety issues in autonomous AI systems during long-horizon tasks.
I recently wrapped up mentoring 10 researchers across 3 projects in SPAR Fall 2025, with multiple papers accepted at ICLR and COLM workshops. I'm now mentoring 2 new projects in SPAR Spring 2026, on agentic situational awareness and efficient safety benchmarking, alongside independent research on alloy agents for AI control.
My recent publications include a TMLR paper on chess neural network interpretability, two ICLR 2026 workshop papers on goal drift in LLM agents, an AAAI 2026 workshop paper on long-context agent safety, and two COLM SoLaR papers on multi-turn jailbreaks and LLM unlearning. I'm currently preparing a COLM main conference submission on reward hacking under tool failure.