Publications

(2026). MMMG: A Comprehensive and Reliable Benchmark for Multitask Multimodal Generation. COLM 2026.
(2026). Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification. COLM 2026.
(2026). Agentic Abstention: Do Agents Know When to Stop Instead of Act?. NeurIPS 2026.
(2026). MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining. ICLR 2026 DATA-FM.
(2026). SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents. IUI 2026.
(2025). Asking the Missing Piece: Context-Driven Clarification for Ambiguous VQA. NeurIPS 2025 FoRLM.
(2025). Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?. NeurIPS 2025 D&B.
(2025). Tensorized Clustered LoRA Merging for Multi-Task Interference. arXiv.
(2025). MARVEL: Modular Abstention for Reliable and Versatile Expert LLMs. ICML 2025.
(2025). Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs. ACL 2025.