From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Haoyu, Zhuang, Jinyu, Chu, Haoran, Yu, Shuhang, J, Group, T AI, Wang, Hao, Han, Kunpeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking and Learning Real-World Customer Service Dialogue
por: Gao, Tianhong, et al.
Publicado: (2025)
por: Gao, Tianhong, et al.
Publicado: (2025)
Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
por: Hong, Mengze, et al.
Publicado: (2024)
por: Hong, Mengze, et al.
Publicado: (2024)
REIC: RAG-Enhanced Intent Classification at Scale
por: Zhang, Ziji, et al.
Publicado: (2025)
por: Zhang, Ziji, et al.
Publicado: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
por: Fabbri, Alexander R., et al.
Publicado: (2025)
por: Fabbri, Alexander R., et al.
Publicado: (2025)
From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
por: Liu, Junhua, et al.
Publicado: (2024)
por: Liu, Junhua, et al.
Publicado: (2024)
Conformal Intent Classification and Clarification for Fast and Accurate Intent Recognition
por: Hengst, Floris den, et al.
Publicado: (2024)
por: Hengst, Floris den, et al.
Publicado: (2024)
SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
por: Yang, Yuqi, et al.
Publicado: (2025)
por: Yang, Yuqi, et al.
Publicado: (2025)
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
por: Zhang, Yidan, et al.
Publicado: (2024)
por: Zhang, Yidan, et al.
Publicado: (2024)
Uddessho: An Extensive Benchmark Dataset for Multimodal Author Intent Classification in Low-Resource Bangla Language
por: Faria, Fatema Tuj Johora, et al.
Publicado: (2024)
por: Faria, Fatema Tuj Johora, et al.
Publicado: (2024)
Deceptive Humor: A Synthetic Multilingual Benchmark Dataset for Bridging Fabricated Claims with Humorous Content
por: Kasu, Sai Kartheek Reddy, et al.
Publicado: (2025)
por: Kasu, Sai Kartheek Reddy, et al.
Publicado: (2025)
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
por: Pu, Yuan, et al.
Publicado: (2024)
por: Pu, Yuan, et al.
Publicado: (2024)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
por: Borisov, Vadim
Publicado: (2026)
por: Borisov, Vadim
Publicado: (2026)
AutoIntent: AutoML for Text Classification
por: Alekseev, Ilya, et al.
Publicado: (2025)
por: Alekseev, Ilya, et al.
Publicado: (2025)
Exploring Description-Augmented Dataless Intent Classification
por: Hu, Ruoyu, et al.
Publicado: (2024)
por: Hu, Ruoyu, et al.
Publicado: (2024)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
por: He, Yun, et al.
Publicado: (2024)
por: He, Yun, et al.
Publicado: (2024)
OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language Models
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
por: Liang, Hao, et al.
Publicado: (2025)
por: Liang, Hao, et al.
Publicado: (2025)
CharacterBench: Benchmarking Character Customization of Large Language Models
por: Zhou, Jinfeng, et al.
Publicado: (2024)
por: Zhou, Jinfeng, et al.
Publicado: (2024)
Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts
por: Liu, Weisi, et al.
Publicado: (2025)
por: Liu, Weisi, et al.
Publicado: (2025)
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
por: Yin, Yuwei, et al.
Publicado: (2026)
por: Yin, Yuwei, et al.
Publicado: (2026)
Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
por: Zhang, Xue, et al.
Publicado: (2025)
por: Zhang, Xue, et al.
Publicado: (2025)
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
por: Sharma, Shivam, et al.
Publicado: (2025)
por: Sharma, Shivam, et al.
Publicado: (2025)
Effectiveness of Pre-training for Few-shot Intent Classification
por: Zhang, Haode, et al.
Publicado: (2021)
por: Zhang, Haode, et al.
Publicado: (2021)
Main Predicate and Their Arguments as Explanation Signals For Intent Classification
por: Pimparkhede, Sameer, et al.
Publicado: (2025)
por: Pimparkhede, Sameer, et al.
Publicado: (2025)
Domain Adaptation in Intent Classification Systems: A Review
por: Atuhurra, Jesse, et al.
Publicado: (2024)
por: Atuhurra, Jesse, et al.
Publicado: (2024)
FAMMA: A Benchmark for Financial Domain Multilingual Multimodal Question Answering
por: Xue, Siqiao, et al.
Publicado: (2024)
por: Xue, Siqiao, et al.
Publicado: (2024)
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
por: Xu, Shuhang, et al.
Publicado: (2025)
por: Xu, Shuhang, et al.
Publicado: (2025)
ITALIC: An Italian Intent Classification Dataset
por: Koudounas, Alkis, et al.
Publicado: (2023)
por: Koudounas, Alkis, et al.
Publicado: (2023)
Contamination Report for Multilingual Benchmarks
por: Ahuja, Sanchit, et al.
Publicado: (2024)
por: Ahuja, Sanchit, et al.
Publicado: (2024)
Towards Data-efficient Customer Intent Recognition with Prompt-based Learning Paradigm
por: Luo, Hengyu, et al.
Publicado: (2023)
por: Luo, Hengyu, et al.
Publicado: (2023)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
por: Pecher, Branislav, et al.
Publicado: (2026)
por: Pecher, Branislav, et al.
Publicado: (2026)
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
por: Mora, David, et al.
Publicado: (2025)
por: Mora, David, et al.
Publicado: (2025)
Capsule Network-Based Semantic Intent Modeling for Human-Computer Interaction
por: Wang, Shixiao, et al.
Publicado: (2025)
por: Wang, Shixiao, et al.
Publicado: (2025)
The Open-World Lottery Ticket Hypothesis for OOD Intent Classification
por: Zhou, Yunhua, et al.
Publicado: (2022)
por: Zhou, Yunhua, et al.
Publicado: (2022)
Estonian Native Large Language Model Benchmark
por: Lillepalu, Helena Grete, et al.
Publicado: (2025)
por: Lillepalu, Helena Grete, et al.
Publicado: (2025)
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
por: He, Linyang, et al.
Publicado: (2025)
por: He, Linyang, et al.
Publicado: (2025)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
por: Hao, Yijie, et al.
Publicado: (2025)
por: Hao, Yijie, et al.
Publicado: (2025)
LLMs for Customized Marketing Content Generation and Evaluation at Scale
por: Liu, Haoran, et al.
Publicado: (2025)
por: Liu, Haoran, et al.
Publicado: (2025)
Ellipsoid-Based Decision Boundaries for Open Intent Classification
por: Zou, Yuetian, et al.
Publicado: (2025)
por: Zou, Yuetian, et al.
Publicado: (2025)
Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
por: Rojas-Galeano, Sergio
Publicado: (2025)
por: Rojas-Galeano, Sergio
Publicado: (2025)
Ejemplares similares
-
Benchmarking and Learning Real-World Customer Service Dialogue
por: Gao, Tianhong, et al.
Publicado: (2025) -
Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
por: Hong, Mengze, et al.
Publicado: (2024) -
REIC: RAG-Enhanced Intent Classification at Scale
por: Zhang, Ziji, et al.
Publicado: (2025) -
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
por: Fabbri, Alexander R., et al.
Publicado: (2025) -
From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
por: Liu, Junhua, et al.
Publicado: (2024)