Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Bryan, Luo, Fiona, Haider, Samar, Agashe, Adwait, Li, Tammy, Liu, Runqi, Miao, Muqing, Ramakrishnan, Shriya, Yuan, Yuan, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
von: Li, Bryan, et al.
Veröffentlicht: (2023)
von: Li, Bryan, et al.
Veröffentlicht: (2023)
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
von: Li, Bryan, et al.
Veröffentlicht: (2024)
von: Li, Bryan, et al.
Veröffentlicht: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
mStyleDistance: Multilingual Style Embeddings and their Evaluation
von: Qiu, Justin, et al.
Veröffentlicht: (2025)
von: Qiu, Justin, et al.
Veröffentlicht: (2025)
The Media Bias Detector: A Framework for Annotating and Analyzing the News at Scale
von: Haider, Samar, et al.
Veröffentlicht: (2025)
von: Haider, Samar, et al.
Veröffentlicht: (2025)
Choice-75: A Dataset on Decision Branching in Script Learning
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2023)
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2023)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News Coverage
von: Wang, Jenny S, et al.
Veröffentlicht: (2025)
von: Wang, Jenny S, et al.
Veröffentlicht: (2025)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
von: Song, Jaewoo, et al.
Veröffentlicht: (2024)
von: Song, Jaewoo, et al.
Veröffentlicht: (2024)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
Towards Faithful Model Explanation in NLP: A Survey
von: Lyu, Qing, et al.
Veröffentlicht: (2022)
von: Lyu, Qing, et al.
Veröffentlicht: (2022)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
von: Jin, Meiqing, et al.
Veröffentlicht: (2025)
von: Jin, Meiqing, et al.
Veröffentlicht: (2025)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Evaluating Vision-Language Models on Bistable Images
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
CtrlRAG: Black-box Document Poisoning Attacks for Retrieval-Augmented Generation of Large Language Models
von: Sui, Runqi
Veröffentlicht: (2025)
von: Sui, Runqi
Veröffentlicht: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models
von: Huang, Runsheng "Anson", et al.
Veröffentlicht: (2024)
von: Huang, Runsheng "Anson", et al.
Veröffentlicht: (2024)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
von: Huang, Runsheng, et al.
Veröffentlicht: (2024)
von: Huang, Runsheng, et al.
Veröffentlicht: (2024)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
von: Rao, Delip, et al.
Veröffentlicht: (2024)
von: Rao, Delip, et al.
Veröffentlicht: (2024)
HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
von: Bian, Zixuan, et al.
Veröffentlicht: (2025)
von: Bian, Zixuan, et al.
Veröffentlicht: (2025)
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
von: Rao, Delip, et al.
Veröffentlicht: (2025)
von: Rao, Delip, et al.
Veröffentlicht: (2025)
OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
von: Zhang, Li, et al.
Veröffentlicht: (2023)
von: Zhang, Li, et al.
Veröffentlicht: (2023)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
Oreo: A Plug-in Context Reconstructor to Enhance Retrieval-Augmented Generation
von: Li, Sha, et al.
Veröffentlicht: (2025)
von: Li, Sha, et al.
Veröffentlicht: (2025)
Multilingual Generative Retrieval via Cross-lingual Semantic Compression
von: Huang, Yuxin, et al.
Veröffentlicht: (2025)
von: Huang, Yuxin, et al.
Veröffentlicht: (2025)
XRAG: Cross-lingual Retrieval-Augmented Generation
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Large Language Models Can Self-Improve At Web Agent Tasks
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
von: Ranaldi, Leonardo, et al.
Veröffentlicht: (2025)
von: Ranaldi, Leonardo, et al.
Veröffentlicht: (2025)
CALYPSO: LLMs as Dungeon Masters' Assistants
von: Zhu, Andrew, et al.
Veröffentlicht: (2023)
von: Zhu, Andrew, et al.
Veröffentlicht: (2023)
Distributed Retrieval-Augmented Generation
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
von: Li, Bryan, et al.
Veröffentlicht: (2023) -
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
von: Li, Bryan, et al.
Veröffentlicht: (2024) -
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025) -
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026) -
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)