Bingbing Wen 🎓

👋 About Me

I am a final-year Ph.D. candidate at the University of Washington, advised by Prof. Bill Howe and Prof. Lucy Lu Wang. I am a member of the UW RAISE Center and collaborate with Prof. Yulia Tsvetkov.

My research asks how models and their training data can evolve together safely and efficiently. I develop methods that use model behavior to determine what the model should learn from next: optimizing data mixtures and curating tasks and reasoning trajectories (MixAtlas, AutoScale, STOP); training agents to clarify, abstain, and coordinate specialized components (Agentic Abstention, Clarify or Answer, MARVEL); and testing whether apparent gains reflect genuine improvements in calibration, robustness, and safety (Know Your Limits, Confidence Calibration, ScienceQA Abstention, SusBench). More broadly, I am interested in self-improving AI systems that can recognize their limitations.

Education

PhD in Information Science (Natural Language Processing)

University of Washington

MS in Computational Science & Engineering (Artificial Intelligence)

University of Hong Kong

BS in Control Science & Engineering (Robotics)

Zhejiang University

Featured Publications
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification featured image

Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification

Reinforcement learning for agentic VQA that balances clarification and answering under underspecified context.

Read more
Agentic Abstention: Do Agents Know When to Stop Instead of Act? featured image

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

A benchmark and analysis of when tool-using LLM agents should stop and abstain rather than continue acting.

Read more
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining featured image

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

Uncertainty-aware data mixture optimization for multimodal LLM midtraining via interpretable domain decomposition.

Read more
MARVEL: Modular Abstention for Reliable and Versatile Expert LLMs featured image

MARVEL: Modular Abstention for Reliable and Versatile Expert LLMs

A modular abstention framework for reliable expert LLMs that enables selective abstention from uncertain questions.

Read more
AutoScale-Automatic Prediction of Compute-optimal Data Composition for Training LLMs featured image

AutoScale-Automatic Prediction of Compute-optimal Data Composition for Training LLMs

Automatic prediction of compute-optimal data composition for efficient LLM training.

Read more
Recent Publications
(2026). Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification. COLM 2026.
(2026). MMMG: A Comprehensive and Reliable Benchmark for Multitask Multimodal Generation. COLM 2026.
(2026). Agentic Abstention: Do Agents Know When to Stop Instead of Act?. NeurIPS 2026.
(2026). MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining. ICLR 2026 DATA-FM.
(2026). SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents. IUI 2026.
📰 News

8/2026 Our paper Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation won the ACL 2026 SAC Award!

8/2026 Our papers Clarify or Answer (CoA) and MMMG have been accepted to COLM 2026!

5/2026 We submitted Agentic Abstention to NeurIPS! Stay tuned for the full paper.

3/2026 Our paper on uncertainty-aware data mixture optimization for MLLM midtraining has been accepted by ICLR 2026 Workshop DATA-FM!

1/2026 Our paper on reinforcement learning for agentic VQA has been released on arXiv.

1/2026 Our benchmark on dark pattern susceptibility of computer-use agents has been accepted by IUI 2026.

9/2025 Our paper about MLLM spurious correlation has been accepted by NeurIPS 2025!

7/2025 I presented our abstention survey in LLMs (oral presentation) and confidence calibration (poster) at ACL 2025!

7/2025 Our paper about modular abstention has been accepted by ICML 2025!

6/2025 I will start my summer internship at Apple as a research intern!

5/2025 Our paper about optimal data mixing in pretraining has been accepted by [COLM 2025]!

5/2025 Our paper about confidence calibration has been accepted by ACL 2025!

2/2025 Our paper about abstention survey in LLMs has been accepted by TACL 2025!