ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Xiaozhe, Lyu, TianYi, Yang, Siyi, Gong, Yuxi, Yang, Yizhao, Huang, Jinxuan, Zhang, Ligao, Huang, Zhuoyi, Liu, Qingwen |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding
par: Li, Xiaozhe, et autres
Publié: (2026)
par: Li, Xiaozhe, et autres
Publié: (2026)
Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning
par: Li, Xiaozhe, et autres
Publié: (2026)
par: Li, Xiaozhe, et autres
Publié: (2026)
MDDFNet: Mamba-based Dynamic Dual Fusion Network for Traffic Sign Detection
par: Yu, TianYi
Publié: (2025)
par: Yu, TianYi
Publié: (2025)
Design description of Wisdom Computing Persperctive
par: Yu, TianYi
Publié: (2025)
par: Yu, TianYi
Publié: (2025)
From Logic to Toolchains: An Empirical Study of Bugs in the TypeScript Ecosystem
par: Tang, TianYi, et autres
Publié: (2026)
par: Tang, TianYi, et autres
Publié: (2026)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
par: Hong, Wenyi, et autres
Publié: (2025)
par: Hong, Wenyi, et autres
Publié: (2025)
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
par: Hu, Xiaomeng, et autres
Publié: (2026)
par: Hu, Xiaomeng, et autres
Publié: (2026)
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
par: Wang, Haoxin, et autres
Publié: (2025)
par: Wang, Haoxin, et autres
Publié: (2025)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
par: Mateega, Spencer, et autres
Publié: (2026)
par: Mateega, Spencer, et autres
Publié: (2026)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
par: Hou, Yutao, et autres
Publié: (2026)
par: Hou, Yutao, et autres
Publié: (2026)
DeepThink: Aligning Language Models with Domain-Specific User Intents
par: Li, Yang, et autres
Publié: (2025)
par: Li, Yang, et autres
Publié: (2025)
Connecting the Dots: Inferring Patent Phrase Similarity with Retrieved Phrase Graphs
par: Peng, Zhuoyi, et autres
Publié: (2024)
par: Peng, Zhuoyi, et autres
Publié: (2024)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
par: Yu, Haorui, et autres
Publié: (2026)
par: Yu, Haorui, et autres
Publié: (2026)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
par: Yang, Shuo, et autres
Publié: (2025)
par: Yang, Shuo, et autres
Publié: (2025)
ECLM: Entity Level Language Model for Spoken Language Understanding with Chain of Intent
par: Yin, Shangjian, et autres
Publié: (2024)
par: Yin, Shangjian, et autres
Publié: (2024)
Automated Functional Testing for Malleable Mobile Application Driven from User Intent
par: Wang, Yuying, et autres
Publié: (2026)
par: Wang, Yuying, et autres
Publié: (2026)
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
par: Wang, Jiangyuan, et autres
Publié: (2025)
par: Wang, Jiangyuan, et autres
Publié: (2025)
Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
par: Tan, Hanzhuo, et autres
Publié: (2025)
par: Tan, Hanzhuo, et autres
Publié: (2025)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
par: Zong, Xuanjun, et autres
Publié: (2025)
par: Zong, Xuanjun, et autres
Publié: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
par: Ding, Shuangrui, et autres
Publié: (2026)
par: Ding, Shuangrui, et autres
Publié: (2026)
Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges
par: Wu, Di, et autres
Publié: (2025)
par: Wu, Di, et autres
Publié: (2025)
Chapter 1 Sustainable cities and landscapes
par: Yang, Yizhao, et autres
Publié: (2022)
par: Yang, Yizhao, et autres
Publié: (2022)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
par: Lu, Jiaxuan, et autres
Publié: (2026)
par: Lu, Jiaxuan, et autres
Publié: (2026)
Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use
par: Fu, Yicheng, et autres
Publié: (2025)
par: Fu, Yicheng, et autres
Publié: (2025)
QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
par: Wu, Yao, et autres
Publié: (2026)
par: Wu, Yao, et autres
Publié: (2026)
RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs
par: Jin, Pengwei, et autres
Publié: (2025)
par: Jin, Pengwei, et autres
Publié: (2025)
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
par: Chen, Xueyao, et autres
Publié: (2026)
par: Chen, Xueyao, et autres
Publié: (2026)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
par: Chen, Jiacheng, et autres
Publié: (2024)
par: Chen, Jiacheng, et autres
Publié: (2024)
SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems
par: Cheng, Qian, et autres
Publié: (2025)
par: Cheng, Qian, et autres
Publié: (2025)
Adversarial Mixup Unlearning
par: Peng, Zhuoyi, et autres
Publié: (2025)
par: Peng, Zhuoyi, et autres
Publié: (2025)
The interplay between lymphatic vessels and macrophages in inflammation response
par: Cheng Zhou, et autres
Publié: (2024)
par: Cheng Zhou, et autres
Publié: (2024)
TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles
par: Yu, Qingchen, et autres
Publié: (2024)
par: Yu, Qingchen, et autres
Publié: (2024)
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
par: Wu, Xiaobao, et autres
Publié: (2024)
par: Wu, Xiaobao, et autres
Publié: (2024)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
par: Shi, Yuzhen, et autres
Publié: (2026)
par: Shi, Yuzhen, et autres
Publié: (2026)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
par: Yang, Chenghao, et autres
Publié: (2025)
par: Yang, Chenghao, et autres
Publié: (2025)
A Novel Koopman-Inspired Method for the Secondary Control of Microgrids with Grid-Forming and Grid-Following Sources
par: Gong, Xun, et autres
Publié: (2023)
par: Gong, Xun, et autres
Publié: (2023)
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
par: Tang, Yuanbo, et autres
Publié: (2026)
par: Tang, Yuanbo, et autres
Publié: (2026)
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
par: Li, Jingyao, et autres
Publié: (2025)
par: Li, Jingyao, et autres
Publié: (2025)
Unlocking the Power of Time-Since-Infection Models: Data Augmentation for Improved Instantaneous Reproduction Number Estimation
par: Shi, Jiasheng, et autres
Publié: (2024)
par: Shi, Jiasheng, et autres
Publié: (2024)
Can Large Language Models Understand Real-World Complex Instructions?
par: He, Qianyu, et autres
Publié: (2023)
par: He, Qianyu, et autres
Publié: (2023)
Documents similaires
-
COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding
par: Li, Xiaozhe, et autres
Publié: (2026) -
Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning
par: Li, Xiaozhe, et autres
Publié: (2026) -
MDDFNet: Mamba-based Dynamic Dual Fusion Network for Traffic Sign Detection
par: Yu, TianYi
Publié: (2025) -
Design description of Wisdom Computing Persperctive
par: Yu, TianYi
Publié: (2025) -
From Logic to Toolchains: An Empirical Study of Bugs in the TypeScript Ecosystem
par: Tang, TianYi, et autres
Publié: (2026)