This is Ye, welcome to my page. I am a Lecturer (aka. Assistant Professor) in the Department of Computer Science at University College London. My research builds the next generation of coding agents that work like software engineers inside a real command-line environment, reasoning across an entire codebase and its tools to tackle tasks such as issue resolution, agent failure repair, and more. My group focuses on CLI coding agents: we design the agent harnesses they operate in, train them in realistic terminal-world environments, and study why they fail and how they can repair themselves. What I care about most is making these agents secure and cost-efficient, so that they are actually affordable in practice. We also maintain Awesome Code Agents, a curated list of cutting-edge coding-agent projects and research.

Prior to joining UCL, I worked as a Postdoctoral Researcher at Carnegie Mellon University with Prof. Claire Le Goues. I obtained my PhD from KTH Royal Institute of Technology, where I was fortunate to be supervised by Prof. Martin Monperrus and Prof. Benoit Baudry. I received my bachelor’s degree from Sichuan University.

Let’s connect: I’m always happy to talk research, swap ideas, or just say hi — whether we share interests or come from very different fields. Reach me at [email protected].

Funding: We greatly appreciate that our research is supported by Google, AWS, Mistral AI, and Delysium (Jerry Zhang) through computing credits and gift funding.

News

  • 2026.07: 🤗 TerminalWorld dataset exceeded 10,000 downloads on HuggingFace!
  • 2026.06: Organizing the 7th APR workshop at ASE — welcome to submit to APR@ASE 2026!
  • 2026.06: Happy to receive the FSE 2026 Distinguished Reviewer Award.
  • 2026.06: Our paper “Agent-based Automated Remediation for Vulnerabilities in Maven Projects” is accepted to OOPSLA 2026. Congrats to Lyuye!
  • 2026.06: 🎙️ TerminalWorld was featured on Last Week in AI (ep. #246), a newsletter and podcast with 181k+ listeners.
  • 2026.03: Our paper ExecVerify is accepted to ACL 2026 (main). Congrats to Lingxiao!
  • 2025.11: Our scaffold Prometheus achieved TOP 1 🏆 by resolving 33.77% of issues on the AWS SWE-PolyBench leaderboard. Check out our experiment results.
  • 2025.10: Our scaffold Prometheus achieved TOP 5 overall on the SWE-bench Verified leaderboard 🎉! Check out our project repository.

Projects

TerminalWorld overview
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks.
Zhaoyang Chu, Jiarui Hu*, Xingyu Jiang*, Pengyu Zou*, Han Li, Chao Peng, Peter O’Hearn, Earl T. Barr, Mark Harman, Federica Sarro, He Ye†.
Preprint.
TerminalWorld is a benchmark that evaluates AI agents on the real-world terminal workflows developers run every day; its novelty is an automated pipeline that mines real terminal recordings into reproducible, test-verified task environments that stay live as engineering practices evolve.
ContextBench overview ContextBench leaderboard
ContextBench: A Benchmark for Context Retrieval in Coding Agents.
Han Li, Letian Zhu, Bohan Zhang, Rili Feng, Jiaming Wang, Yue Pan, Earl T Barr, Federica Sarro, Zhaoyang Chu, He Ye.
Preprint.
ContextBench is a benchmark that evaluates how coding agents perform multi-file context retrieval across a repository; its novelty is measuring the dynamics of retrieval — not just whether the right files are found, but the accuracy (context F1), efficiency, and cost of how agents gather that context.
Prometheus overview Prometheus results
Prometheus: Towards Long-Horizon Codebase Navigation for Repository-Level Problem Solving.
Yue Pan, Zimin Chen, Siyu Lu, Zhaoyang Chu, Xiang Li, Han Li, Yang Feng, Claire Le Goues, Federica Sarro, Martin Monperrus, He Ye.
Preprint.
Prometheus is a coding agent for repository-level problem solving that navigates large codebases over long horizons; its novelty is unifying embedding-based retrieval, structure-aware knowledge-graph navigation, and a working memory over the agent's trajectory, so it gathers the right context across many files and reasoning steps — reaching top results on SWE-bench Verified and SWE-PolyBench.