I am a final-year Ph.D. candidate (class of 2027) in the School of Automation Science and Electrical Engineering (SASEE), Beihang University, supervised by Prof. Zheng Zheng and collaborating with Prof. Tsong Yueh Chen (IEEE Fellow, Swinburne University). I received my B.Eng. from the Shenyuan Honors College of Beihang University.
My research studies how to evaluate intelligent systems when no ground truth exists. Using metamorphic testing as the principal tool, I design verifiable and attributable metrics for oracle-less scenarios — from a completeness theory for testing deep-learning operators, to meta-evaluation of LLM reasoning reliability (LGMT), to validity assessment of neuron coverage metrics.
Recently I work on Eval for Training: turning evaluation signals into training signals. At Kling AI (Kuaishou), I build MLLM-as-Judge pipelines whose quality signals feed reward models and reinforcement learning for generative models. At ByteDance, I designed agent evaluation frameworks and built SenWorld, a digital-twin simulation that generates oracle-labeled evaluation data at scale.
News
- 2026.09 Our empirical study on code comments for quantum SDKs was published in ACM TOSEM (CCF-A).First-author work on comment practices in quantum software development kits (Qiskit).
- 2026.08 Joined Kling AI (Kuaishou) as a Multimodal Evaluation Algorithm Intern.Working on MLLM-as-Judge evaluation pipelines and reward modeling.
- 2026.07 SenWorld was accepted by ISSREW 2026 (co-located with IEEE ISSRE).A digital-twin simulation for generating context-rich, oracle-labeled evaluation data.
- 2025.12 Honored as the Merit Student of Beijing (City-level Honor).
- 2025.11 Open-sourced a Response Letter LaTeX Template on GitHub .
Experience
- 2026.08 – Present Multimodal Evaluation Algorithm Intern, Kling AI, Kuaishou (可灵 AI · 快手).Building MLLM-as-Judge evaluation pipelines for image stylization (human-agreement lifted from 56% to 71%); designed a VLLM-OCR coupled detector for text-rendering defects in generated content; distilling VQA/Rubric quality signals into reward models for RL-stage training alignment.
- 2026.04 – 2026.08 LLM Evaluation Algorithm Intern, Doubao Mobile Assistant, ByteDance (豆包手机助手 · 字节跳动).Designed the evaluation framework for proactive agent suggestions; built a PRD-to-test-case generation workflow (AI-generated cases up to 60%); led SenWorld, a persona-agent and digital-twin simulation environment that yields oracle-labeled, annotation-free evaluation data (accepted at ISSREW 2026).
Publications
Selected Publications
Manuscripts Under Review
Honors and Awards
Academic Honors
Scholarships & Competitions
Leadership & Service