Publications
For a complete list, please visit my Google Scholar profile Citations2495.
Benchmarking and Understanding Coding Agents
- TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks . Zhaoyang Chu, Jiarui Hu, Xingyu Jiang, Pengyu Zou, Han Li, Chao Peng, Peter O'Hearn, Earl T Barr, Mark Harman, Federica Sarro, He Ye. Preprint Arxiv 2026.
- Failure as a Process: An Anatomy of CLI Coding Agent Trajectories . Xiangxin Zhao, Han Li, Shuaiting Li, Tianyi Zhao, Earl T Barr, Federica Sarro, He Ye. Preprint Arxiv 2026.
- Prometheus: Unified Knowledge Graphs for Issue Resolution in Multilingual Codebases . Zimin Chen, Yue Pan, Siyu Lu, Jiayi Xu, Claire Le Goues, Martin Monperrus, He Ye. Arxiv 2025.
- HerAgent: Rethinking the Automated Environment Deployment via Hierarchical Test Pyramid . Xiang Li, Siyu Lu, Federica Sarro, Claire Le Goues, He Ye. Preprint Arxiv 2026.
- ContextBench: A Benchmark for Context Retrieval in Coding Agents . Han Li, Letian Zhu, Bohan Zhang, Rili Feng, Jiaming Wang, Yue Pan, Earl T Barr, Federica Sarro, Zhaoyang Chu, He Ye. Preprint Arxiv 2026.
- Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories . Oorja Majgaonkar, Zhiwei Fei, Xiang Li, Federica Sarro, He Ye. Arxiv 2025.
- AdverIntent-Agent: Adversarial Reasoning for Repair Based on Inferred Program Intent . He Ye, Aidan ZH Yang, Chang Hu, Yanlin Wang, Tao Zhang, Claire Le Goues. Proceedings of the ACM on Software Engineering 2 (ISSTA), 1398-1420, 2025.
- MemoryBank: Enhancing Large Language Models with Long-Term Memory . Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, Yanlin Wang. Proceedings of the 38th AAAI Conference on Artificial Intelligence, 38(17), 19724-19731.
Model Post-Training and Reasoning
- ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning . Lingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye, Zhongxin Liu, Xiaoxue Ren, Lingfeng Bao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13850-13875.
- Security vulnerability detection with multitask self-instructed fine-tuning of large language models . Aidan ZH Yang, Haoye Tian, He Ye, Ruben Martins, Claire Le Goues. Arxiv 2024.
Learning-based Automated Program Repair
- Iter: Iterative neural repair for multi-location patches . He Ye, Martin Monperrus. Proceedings of the 46th IEEE/ACM international conference on software engineering,2024.
- Neural program repair with execution-based backpropagation . He Ye, Matias Martinez, Martin Monperrus. Proceedings of the 44th international conference on software engineering,2022.
- Selfapr: Self-supervised program repair with test execution diagnostics . He Ye, Matias Martinez, Xiapu Luo, Tao Zhang, Martin Monperrus. Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering,2022.



