Bio. I am a Research Assistant at the Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Chengwei Qin. I am also a Research Intern at Zhejiang University (IFRC Lab), working on trustworthy language models with Prof. Meng Han and Dr. Wenpeng Xing. Additionally, I am honored to receive guidance on AI safety research from Prof. Simin Chen (George Mason University), Prof. Yi Liu (Griffith University), and Prof. Ying Zhang (Wake Forest University).

Beyond research, I gained valuable industry experience through an AI agent development internship (July–October 2025). I am grateful to my mentor, Wenjie Zhang, a P9 engineer at Alibaba Group, for his guidance and support.

Research. From latent knowledge to reliable action, memory, and verifiable trust. I study how models internally represent knowledge, evidence provenance, conflict, and risk; why those signals fail to guide generation, tool use, and memory; and how monitoring, system safeguards, and verification can make AI systems more reliable.

Research Agenda

Trustworthy agents must not only represent provenance, conflict, and risk internally. These signals should reliably inform what agents say, do, remember, reuse, and learn, while the resulting models and artifacts remain externally verifiable.

01

Internal representation

Understand how models encode knowledge, retrieved evidence, source provenance, conflict, and hallucination risk through white-box monitoring and hidden-state analysis.

02

Action and memory

Study why decodable safety signals fail to guide generation, tool use, and memory updates, and turn monitoring into reliable system safeguards.

03

Verifiable trust

Build accountable deployment mechanisms through information-flow control, model fingerprinting, blockchain, and privacy-preserving verification.

Current themes

  • Multi-agent memory consistency: write validity, contamination, persistence, propagation, and repair.
  • Poisoning defense and credit: separating the source, amplifier, executor, and best remediation target.
  • Runtime monitoring of agent skills: identifying risky skills before unsafe tool actions.

My longer-term goal is to understand environment–verifier–memory co-evolution in lifelong agents: how tasks, evaluation, memory, and oversight should adapt together as agents act, learn, and update over time.

Academic Services

Reviewer

  • EMNLP 2026
  • NeurIPS 2026
  • AAAI 2027
  • ICLR 2027

Publications

* Equal contribution.

Experience & Education

George Mason University

Fairfax, VA, USA

Research Intern, supervised by Prof. Simin Chen.

Hong Kong University of Science and Technology (Guangzhou)

Guangzhou, China

Research Assistant, advised by Prof. Chengwei Qin.

Binjiang Institute of Zhejiang University · IFRC Lab

Hangzhou, China

Research Intern, supervised by Prof. Meng Han and Dr. Wenpeng Xing.

University of Malaya

Kuala Lumpur, Malaysia

Visiting Student.

Westlake University

Hangzhou, China

Research Assistant, supervised by Dr. Ziyang Zhang.

Communication University of Zhejiang

Hangzhou, China

B.Eng. in Artificial Intelligence, supervised by Dr. Hao Zeng.

Patents

  1. Zhe Yu, Wenpeng Xing, Meng Han. A hallucination detection method based on dual-path internal state forcing logic for retrieval-augmented generation in large language models. Application No. 202610260408X. Under review.
  2. Meng Han, Zhe Yu, Jiayan Hu, Rongchang Li, Wenpeng Xing, Jingyi Yu, Zhen Hong, et al. A post-processing method, system, device, and medium for hallucination detection in large language models based on adaptive order statistics aggregation. Application No. 2026107898102, filed June 3, 2026. Pending.
  3. Zhe Yu, Jiayan Hu, Jingyi Yu, Weihang Yu, Wenpeng Xing, Jing Xiong, Yourong Chen, Zhen Hong, et al. A hallucination detection method, system, and device for large language models based on multi-dimensional heterogeneous feature fusion. Application No. 2026107899270, filed June 3, 2026. Pending.