LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeong, Seogyeong, Park, Kiwoong, Song, Seyoung, Kim, Eunsu, Friedl, Ken E., Kim, Jaeho, Oh, Alice |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
von: Park, Junyeong, et al.
Veröffentlicht: (2025)
von: Park, Junyeong, et al.
Veröffentlicht: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
von: Oh, Juhyun, et al.
Veröffentlicht: (2025)
von: Oh, Juhyun, et al.
Veröffentlicht: (2025)
Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
Defect modulation and in‐situ exsolution in Y 2 Ru 2 O 7 @NiFeP/Ru heterostructure for enhanced oxygen evolution reaction
von: Eunsu Jang, et al.
Veröffentlicht: (2024)
von: Eunsu Jang, et al.
Veröffentlicht: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
Knowledge-Aware Iterative Retrieval for Multi-Agent Systems
von: Song, Seyoung
Veröffentlicht: (2025)
von: Song, Seyoung
Veröffentlicht: (2025)
Systemic Credit Risk Premium: Insights From Credit Derivatives Markets
von: Kiwoong Byun, et al.
Veröffentlicht: (2025)
von: Kiwoong Byun, et al.
Veröffentlicht: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Transfer Learning for Benign Overfitting in High-Dimensional Linear Regression
von: Kim, Yeichan, et al.
Veröffentlicht: (2025)
von: Kim, Yeichan, et al.
Veröffentlicht: (2025)
Benchmarking Contextual Understanding for In-Car Conversational Systems
von: Habicht, Philipp, et al.
Veröffentlicht: (2025)
von: Habicht, Philipp, et al.
Veröffentlicht: (2025)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models
von: Giebisch, Rafael, et al.
Veröffentlicht: (2025)
von: Giebisch, Rafael, et al.
Veröffentlicht: (2025)
Sociolinguistic dimensions of immigration to the United States
von: Kim Potowski
Veröffentlicht: (2013)
von: Kim Potowski
Veröffentlicht: (2013)
LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
Does FOMC Tone Really Matter? Statistical Evidence from Spectral Graph Network Analysis
von: Choi, Jaeho, et al.
Veröffentlicht: (2025)
von: Choi, Jaeho, et al.
Veröffentlicht: (2025)
When Tom Eats Kimchi: Evaluating Cultural Bias of Multimodal Large Language Models in Cultural Mixture Contexts
von: Kim, Jun Seong, et al.
Veröffentlicht: (2025)
von: Kim, Jun Seong, et al.
Veröffentlicht: (2025)
Measuring Interest Group Positions on Legislation: An AI-Driven Analysis of Lobbying Reports
von: Kim, Jiseon, et al.
Veröffentlicht: (2025)
von: Kim, Jiseon, et al.
Veröffentlicht: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA
von: Lee, Sangyoon, et al.
Veröffentlicht: (2026)
von: Lee, Sangyoon, et al.
Veröffentlicht: (2026)
Unexplored Faces of Robustness and Out-of-Distribution: Covariate Shifts in Environment and Sensor Domains
von: Baek, Eunsu, et al.
Veröffentlicht: (2024)
von: Baek, Eunsu, et al.
Veröffentlicht: (2024)
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
von: Choi, Yoonseo, et al.
Veröffentlicht: (2025)
von: Choi, Yoonseo, et al.
Veröffentlicht: (2025)
Riemannian Optimization for LoRA on the Stiefel Manifold
von: Park, Juneyoung, et al.
Veröffentlicht: (2025)
von: Park, Juneyoung, et al.
Veröffentlicht: (2025)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
von: Yu, Haeun, et al.
Veröffentlicht: (2025)
von: Yu, Haeun, et al.
Veröffentlicht: (2025)
Innovation in Financial Inclusion Policies with Digital Transformation: Evidence from South Korea*
von: Kiwoong Byun, et al.
Veröffentlicht: (2024)
von: Kiwoong Byun, et al.
Veröffentlicht: (2024)
Shared Heritage, Distinct Writing: Rethinking Resource Selection for East Asian Historical Documents
von: Song, Seyoung, et al.
Veröffentlicht: (2024)
von: Song, Seyoung, et al.
Veröffentlicht: (2024)
HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
von: Bayramli, Zahra, et al.
Veröffentlicht: (2025)
von: Bayramli, Zahra, et al.
Veröffentlicht: (2025)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
von: Park, Seonghwan, et al.
Veröffentlicht: (2025)
von: Park, Seonghwan, et al.
Veröffentlicht: (2025)
Quantum prime factorization algorithms using binary carry propagation
von: Ryou, Arim, et al.
Veröffentlicht: (2025)
von: Ryou, Arim, et al.
Veröffentlicht: (2025)
Quantum compressed sensing tomographic reconstruction algorithm
von: Ryou, Arim, et al.
Veröffentlicht: (2025)
von: Ryou, Arim, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
von: Song, Seyoung, et al.
Veröffentlicht: (2025) -
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
von: Kim, Eunsu, et al.
Veröffentlicht: (2025) -
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
von: Park, Junyeong, et al.
Veröffentlicht: (2025) -
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
von: Oh, Juhyun, et al.
Veröffentlicht: (2024) -
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
von: Oh, Juhyun, et al.
Veröffentlicht: (2025)