Internal representation
Understand how models encode knowledge, retrieved evidence, source provenance, conflict, and hallucination risk through white-box monitoring and hidden-state analysis.
Bio. I am a Research Assistant at the Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Chengwei Qin. I am also a Research Intern at Zhejiang University (IFRC Lab), working on trustworthy language models with Prof. Meng Han and Dr. Wenpeng Xing. Additionally, I am honored to receive guidance on AI safety research from Prof. Simin Chen (George Mason University), Prof. Yi Liu (Griffith University), and Prof. Ying Zhang (Wake Forest University).
Beyond research, I gained valuable industry experience through an AI agent development internship (July–October 2025). I am grateful to my mentor, Wenjie Zhang, a P9 engineer at Alibaba Group, for his guidance and support.
Research. From latent knowledge to reliable action, memory, and verifiable trust. I study how models internally represent knowledge, evidence provenance, conflict, and risk; why those signals fail to guide generation, tool use, and memory; and how monitoring, system safeguards, and verification can make AI systems more reliable.
Trustworthy agents must not only represent provenance, conflict, and risk internally. These signals should reliably inform what agents say, do, remember, reuse, and learn, while the resulting models and artifacts remain externally verifiable.
Understand how models encode knowledge, retrieved evidence, source provenance, conflict, and hallucination risk through white-box monitoring and hidden-state analysis.
Study why decodable safety signals fail to guide generation, tool use, and memory updates, and turn monitoring into reliable system safeguards.
Build accountable deployment mechanisms through information-flow control, model fingerprinting, blockchain, and privacy-preserving verification.
My longer-term goal is to understand environment–verifier–memory co-evolution in lifelong agents: how tasks, evaluation, memory, and oversight should adapt together as agents act, learn, and update over time.
* Equal contribution.
Accepted at ACL 2026
Detects RAG hallucinations through dual-path internal-state forcing and white-box activation signals.
Accepted at EMNLP 2026
Uses three complementary internal signals for token-level, training-free intervention under retrieval-memory conflict.
Accepted at EMNLP 2026
Introduces a double-gate protocol that separates atomic knowledge stability from compositional reasoning.
Under review at AAAI 2027
PLAIT converts hallucination scores into parent-preserving response–claim–span audit plans, using whole-plan learning and exact budgeted optimization to capture more unsupported content per review minute.
Under review at AAAI 2027
Separates identical surface outputs that follow different internal parametric-memory and retrieved-evidence routes.
Under review at AAAI 2027
GAVEL shows item calibration cannot validate experimental contrasts: paired audits reveal that Qwen judges greatly overstate a reminder intervention.
Under review at NeurIPS 2026
Shows that source role is decodable early while tool decisions become controllable only in a late commitment band.
Under review at NeurIPS 2026
Exposes a source-override vulnerability in which assistant-side reasoning for another question can redirect the final answer.
To be submitted to ICLR 2027
A compartmentalized multi-agent defense that restricts how untrusted evidence reaches final synthesis.
Under review at ARR (October 2026)
Enables scalable model fingerprint transfer through vector addition for ownership verification.
Accepted at NeurIPS 2026 Workshop
Monitors faithfulness using distances between residual-stream activations and evidence representations.
Accepted at NeurIPS 2026 Workshop
A 12,522-sample evidence-grounded benchmark and a two-stage white-box hallucination-risk triage framework.
Accepted at ACM TURC 2026
Combines zero-knowledge proofs and blockchain for privacy-preserving model ownership attribution.
Electronics (MDPI)
Trusted metadata coordination and tiered off-chain storage for recovery-safe, low-latency IoT data management.
Optics Communications, 2025
Parallel dual all-fiber interferometers for orthogonal salinity and temperature detection.
ACM ICETM, 2025
A bibliometric analysis of research trends, keyword co-occurrence, and institutional collaboration.
Guangzhou, China
Research Assistant, advised by Prof. Chengwei Qin.
Hangzhou, China
Research Intern, supervised by Prof. Meng Han and Dr. Wenpeng Xing.
Kuala Lumpur, Malaysia
Visiting Student.
Hangzhou, China
Research Assistant, supervised by Dr. Ziyang Zhang.
Hangzhou, China
B.Eng. in Artificial Intelligence, supervised by Dr. Hao Zeng.