Agentic Systems

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

Understanding and mitigating shortcut tool selection in RL-trained agents through rewards for tool necessity.

Read more
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification featured image

Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification

Reinforcement learning for agentic VQA that balances clarification and answering under underspecified context.

Read more
Agentic Abstention: Do Agents Know When to Stop Instead of Act? featured image

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

A benchmark and analysis of when tool-using LLM agents should stop and abstain rather than continue acting.

Read more

SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents

Benchmarking dark pattern susceptibility of computer-use agents in realistic UI environments.

Read more