Back to news
arXiv cs.LG · 2026-08-12 00:00 UTC
research

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real intera

Why it matters

A benchmark hiding user profiles tests true personalization, helping teams avoid overfitting to leaked cues and better evaluate privacy-preserving agent behavior.

Read the original story

Published to Cognify News · Week 33, 2026