Binjie Zhang
Ph.D. Candidate, Show Lab, National University of Singapore.
Email: binjie97 [at] u.nus.edu
Show Lab, NUS
Singapore
I am a fourth-year Ph.D. candidate in Computer Science at the National University of Singapore (NUS), advised by Prof. Mike Zheng Shou in Show Lab.
My primary focus is LLM agents: agent harnesses, tool use, memory, context management, and self-evolution. My research background spans multimodal understanding and continual / compatible representation learning. I study how AI systems can improve while preserving useful knowledge and reliable behavior.
I currently intern with TikTok Search, working on Agentic Search and Tako. My contributions focus on scenario-aware agent policies and an evidence-driven iteration loop for diagnosing failures, evaluating changes, and improving the search experience. My ongoing model-training work connects faithful trajectory data, SFT / reinforcement learning / on-policy distillation, and separate evaluation of agent behavior and the delivered product.
Previously, at TikTok Shop in Singapore (January–June 2026), I led three agent systems from design to production within five months: FIRE, a reusable skill-based agent harness; EvoA, an autonomous model-iteration system; and APA, an evidence-driven risk-discovery pipeline. EvoA reduced the human audit rate by approximately 4.6% relative in a controlled online experiment. Across these projects, I owned architecture, implementation, experimentation, and cross-team integration, with business and infrastructure partners supporting rollout.
I have four first-author or co-first-author conference papers at ICLR, IJCAI (long oral), AAAI, and ECCV, an additional TaCA preprint, and three ICLR 2027 submissions plus one AAAI submission currently under review. My latest published work, ReGRPO, learns grounded reflection and recovery for tool-using agents.
Before NUS, I received my M.Eng. in Computer Science and Technology from Tsinghua University, supervised by Prof. Chun Yuan, and my B.Eng. in Information Engineering from East China University of Science and Technology (ECUST), where I ranked 1 / 92. At Tencent PCG / ARC Lab, I worked on compatible model upgrades for image and video retrieval, helping accelerate production hot updates and reduce feature-refresh cost. This work received the SZCCF Science and Technology Award and the Tencent Technology Breakthrough Award.
news
| Oct 08, 2026 | Currently working with TikTok Search on Agentic Search and Tako, with contributions to scenario-aware policies, evidence-driven iteration, and ongoing data / training / evaluation design. |
|---|---|
| Oct 08, 2026 | Current submissions: SGRC, StateTrackBench, and PathComp are under review at ICLR 2027; POU-OPD is under review at AAAI. |
| Jun 20, 2026 | ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents accepted to ECCV 2026. |
| Mar 07, 2026 | Submitted ReGRPO (reflection-augmented RL for tool-using agents) to ECCV 2026 as first author. |
| Jan 30, 2026 | Exploring continual vision–language–action adaptation and egocentric future prediction. |
selected publications
- Preprint
TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible AdapterarXiv preprint arXiv:2306.12642, 2023Preprint - AAAI’23Darwinian Model Upgrades: Model Evolving with Selective CompatibilityIn AAAI Conference on Artificial Intelligence (AAAI), 2023Co-first author
- IJCAI’22
Towards Universal Backward-Compatible Representation LearningIn International Joint Conference on Artificial Intelligence (IJCAI), 2022Long oral - ICLR’22
Hot-Refresh Model Upgrades with Regression-Alleviating Compatible Training in Image RetrievalIn International Conference on Learning Representations (ICLR), 2022