Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
Reinforcement learning for agentic VQA that balances clarification and answering under underspecified context.
I am a final-year Ph.D. candidate at the University of Washington, advised by Prof. Bill Howe and Prof. Lucy Lu Wang. I am a member of the UW RAISE Center and collaborate with Prof. Yulia Tsvetkov.
My research builds closed loops for data-model co-evolution: models expose capability gaps, those gaps drive new training data, learning produces stronger models, and rigorous evaluation verifies the gains.
PhD in Information Science (Natural Language Processing)
University of Washington
MS in Computational Science & Engineering (Artificial Intelligence)
University of Hong Kong
BS in Control Science & Engineering (Robotics)
Zhejiang University
Reinforcement learning for agentic VQA that balances clarification and answering under underspecified context.
A benchmark and analysis of when tool-using LLM agents should stop and abstain rather than continue acting.
Uncertainty-aware data mixture optimization for multimodal LLM midtraining via interpretable domain decomposition.
A modular abstention framework for reliable expert LLMs that enables selective abstention from uncertain questions.
Automatic prediction of compute-optimal data composition for efficient LLM training.
8/2026 Our paper Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation won the ACL 2026 SAC Award!
8/2026 Our papers Clarify or Answer (CoA) and MMMG have been accepted to COLM 2026!
5/2026 We submitted Agentic Abstention to NeurIPS! Stay tuned for the full paper.
3/2026 Our paper on uncertainty-aware data mixture optimization for MLLM midtraining has been accepted by ICLR 2026 Workshop DATA-FM!
1/2026 Our paper on reinforcement learning for agentic VQA has been released on arXiv.
1/2026 Our benchmark on dark pattern susceptibility of computer-use agents has been accepted by IUI 2026.
9/2025 Our paper about MLLM spurious correlation has been accepted by NeurIPS 2025!
7/2025 I presented our abstention survey in LLMs (oral presentation) and confidence calibration (poster) at ACL 2025!
7/2025 Our paper about modular abstention has been accepted by ICML 2025!
6/2025 I will start my summer internship at Apple as a research intern!
5/2025 Our paper about optimal data mixing in pretraining has been accepted by [COLM 2025]!
5/2025 Our paper about confidence calibration has been accepted by ACL 2025!
2/2025 Our paper about abstention survey in LLMs has been accepted by TACL 2025!