Latest work
Behavioral Self-Rewarding (BSR)
Learning from debugging behavior to improve software engineering agents beyond binary test rewards.
Preprint, 2026
I am a PhD candidate at Penn State’s College of Information Sciences and Technology, where I have been advised by Prof. Qingyun Wu since 2022. Before joining Penn State, I received my master’s degree in AI & ML from Imperial College London and my bachelor’s degree in Computer Science from the University of California, Davis.
My research focuses on large language models, especially multi-agent systems and reinforcement learning for language agents.
I was a research intern at Microsoft Research, Redmond in 2024 and 2025.
I’m looking for full-time research scientist positions, preferably in North America. Résumé · Contact me
Latest work
Learning from debugging behavior to improve software engineering agents beyond binary test rewards.
Preprint, 2026
Unified environments for digital agents, connecting graphical and text interaction.
ICML 2026 · Position Paper Track
A benchmark for evaluating LLM agents on cyber threat investigation tasks.
Creator · ICML 2026
Reinforced self-play for reasoning, with no human-curated training data.
Co-creator & maintainer · NeurIPS 2025
An open-source framework for building LLM applications through multi-agent conversation.
Co-creator & maintainer
No selected papers in this topic yet. Switch to Full to explore all publications.
Preprint 2026
Preprint 2026
* Equal contribution.
Back to top