Evaluating Cultural and Social Awareness of LLM Web Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Qiu, Haoyi, Fabbri, Alexander R., Agarwal, Divyansh, Huang, Kung-Hsiang, Tan, Sarah, Peng, Nanyun, Wu, Chien-Sheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
di: Qiu, Haoyi, et al.
Pubblicazione: (2025)
di: Qiu, Haoyi, et al.
Pubblicazione: (2025)
AMRFact: Enhancing Summarization Factuality Evaluation with AMR-Driven Negative Samples Generation
di: Qiu, Haoyi, et al.
Pubblicazione: (2023)
di: Qiu, Haoyi, et al.
Pubblicazione: (2023)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
di: Qiu, Haoyi, et al.
Pubblicazione: (2025)
di: Qiu, Haoyi, et al.
Pubblicazione: (2025)
SafeWorld: Geo-Diverse Safety Alignment
di: Yin, Da, et al.
Pubblicazione: (2024)
di: Yin, Da, et al.
Pubblicazione: (2024)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
di: Agarwal, Divyansh, et al.
Pubblicazione: (2024)
di: Agarwal, Divyansh, et al.
Pubblicazione: (2024)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
di: Huang, Tenghao, et al.
Pubblicazione: (2026)
di: Huang, Tenghao, et al.
Pubblicazione: (2026)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2023)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2023)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
Decoupling Task-Solving and Output Formatting in LLM Generation
di: Deng, Haikang, et al.
Pubblicazione: (2025)
di: Deng, Haikang, et al.
Pubblicazione: (2025)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2024)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2024)
Nudging the Boundaries of LLM Reasoning
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2025)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2025)
Art or Artifice? Large Language Models and the False Promise of Creativity
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
Benchmarking Deep Search over Heterogeneous Enterprise Data
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2025)
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2025)
BingoGuard: LLM Content Moderation Tools with Risk Levels
di: Yin, Fan, et al.
Pubblicazione: (2025)
di: Yin, Fan, et al.
Pubblicazione: (2025)
Agentic Uncertainty Quantification
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
di: Laban, Philippe, et al.
Pubblicazione: (2024)
di: Laban, Philippe, et al.
Pubblicazione: (2024)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2026)
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2026)
MMGR: Multi-Modal Generative Reasoning
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
di: Venkit, Pranav Narayanan, et al.
Pubblicazione: (2025)
di: Venkit, Pranav Narayanan, et al.
Pubblicazione: (2025)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2024)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2024)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
LLM-REVal: Can We Trust LLM Reviewers Yet?
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2023)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2023)
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
di: Yerukola, Akhila, et al.
Pubblicazione: (2025)
di: Yerukola, Akhila, et al.
Pubblicazione: (2025)
NewsEdits 2.0: Learning the Intentions Behind Updating News
di: Spangher, Alexander, et al.
Pubblicazione: (2024)
di: Spangher, Alexander, et al.
Pubblicazione: (2024)
ReIFE: Re-evaluating Instruction-Following Evaluation
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
di: Liou, Yow-Fu, et al.
Pubblicazione: (2026)
di: Liou, Yow-Fu, et al.
Pubblicazione: (2026)
Adaptable Logical Control for Large Language Models
di: Zhang, Honghua, et al.
Pubblicazione: (2024)
di: Zhang, Honghua, et al.
Pubblicazione: (2024)
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2024)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2024)
Context-Aware SQL Error Correction Using Few-Shot Learning -- A Novel Approach Based on NLQ, Error, and SQL Similarity
di: Jain, Divyansh, et al.
Pubblicazione: (2024)
di: Jain, Divyansh, et al.
Pubblicazione: (2024)
Unanswerability Evaluation for Retrieval Augmented Generation
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
di: Guan, Shengyue, et al.
Pubblicazione: (2025)
di: Guan, Shengyue, et al.
Pubblicazione: (2025)
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2025)
di: Choubey, Prafulla Kumar, et al.
Pubblicazione: (2025)
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
di: Wu, You, et al.
Pubblicazione: (2024)
di: Wu, You, et al.
Pubblicazione: (2024)
SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation
di: Tian, Yufei, et al.
Pubblicazione: (2025)
di: Tian, Yufei, et al.
Pubblicazione: (2025)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
di: Wu, Di, et al.
Pubblicazione: (2026)
di: Wu, Di, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
di: Qiu, Haoyi, et al.
Pubblicazione: (2025) -
AMRFact: Enhancing Summarization Factuality Evaluation with AMR-Driven Negative Samples Generation
di: Qiu, Haoyi, et al.
Pubblicazione: (2023) -
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025) -
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
di: Qiu, Haoyi, et al.
Pubblicazione: (2025) -
SafeWorld: Geo-Diverse Safety Alignment
di: Yin, Da, et al.
Pubblicazione: (2024)