Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act
Understanding and mitigating shortcut tool selection in RL-trained agents through rewards for tool necessity.
•
1 min read
Read more
Understanding and mitigating shortcut tool selection in RL-trained agents through rewards for tool necessity.
A benchmark and analysis of when tool-using LLM agents should stop and abstain rather than continue acting.
Tensorized clustered LoRA merging to reduce multi-task interference in adapter-based LLM fine-tuning.
An informative visual dialogue dataset created by bridging large multimodal and language models.
Improving robustness, fairness, and emotion-awareness of explanations in recommender systems.
Generating explanations in multi-turn conversational recommendation with EGCR.