tau-bench - A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Summary
Ï„-bench is a test benchmark for AI agents that need to use tools (APIs) in real-world domains.
What it has:
- Two domains: retail (book orders) and airline (flight bookings)
- Conversations between simulated users, AI agents, and tools
- Success and failure examples (labeled 1 and 0)
How they made it:
- Users simulated by LLMs following templates (act like real humans)
- APIs created manually and with LLMs
- Database examples created manually then expanded
New metric (pass^k):
- pass@k = did agent succeed in ANY of k tries
- pass^k = did agent succeed in ALL k tries (consistency matters)
- This shows reliability, not just luck
Annotations
« Our evaluation scheme compares the database state at the end of each episode with the ground truth expected state. »(2)
« We also introduce the metric of pass^k, which measures the consistency and robustness of the agent across k i.i.d. trials. »(2)
« For instance, even state-of-the-art LMs like gpt-4o achieve low task success rates (pass^1) using function calling (∼61% on τ -retail and ∼35% on τ -airline). With increasing k, the chance of consistently solving a task drops rapidly, to as low as ∼25% for pass^8 on τ -retail for the same model »(2)
« For example, if the preferred payment method is not specified, the user might answer differently and cause the final database to be different across trials »(5)
« we run each τ -retail task with > 40 gpt-4-turbo trials and check all tasks with zero or low success rates) »(5)
« so the cost is mainly due to long system prompt (domain policy + function definitions). »(7)
Date : 06-17-2024
Authors : Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan
Paper Link : http://arxiv.org/abs/2406.12045
Zotero Link: Full Text PDF
Tags : #Computer-Science---Artificial-Intelligence, #Computer-Science---Computation-and-Language
Citation :