EgoNormia: Benchmarking Physical Social Norm Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rezaei, MohammadHossein, Fu, Yicheng, Cuvin, Phil, Ziems, Caleb, Zhang, Yanzhe, Zhu, Hao, Yang, Diyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Paraphrasing in Affirmative Terms Improves Negation Understanding
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2024)
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2024)
Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future
von: Li, Minzhi, et al.
Veröffentlicht: (2024)
von: Li, Minzhi, et al.
Veröffentlicht: (2024)
Making Language Models Robust Against Negation
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025)
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025)
Commonsense Knowledge with Negation: A Resource to Enhance Negation Understanding
von: Wang, Zijie, et al.
Veröffentlicht: (2026)
von: Wang, Zijie, et al.
Veröffentlicht: (2026)
Social Skill Training with Large Language Models
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
von: Zhu, Hao, et al.
Veröffentlicht: (2025)
von: Zhu, Hao, et al.
Veröffentlicht: (2025)
Measuring and Addressing Indexical Bias in Information Retrieval
von: Ziems, Caleb, et al.
Veröffentlicht: (2024)
von: Ziems, Caleb, et al.
Veröffentlicht: (2024)
Can Large Language Models Transform Computational Social Science?
von: Ziems, Caleb, et al.
Veröffentlicht: (2023)
von: Ziems, Caleb, et al.
Veröffentlicht: (2023)
Culture Cartography: Mapping the Landscape of Cultural Knowledge
von: Ziems, Caleb, et al.
Veröffentlicht: (2025)
von: Ziems, Caleb, et al.
Veröffentlicht: (2025)
Searching for Privacy Risks in LLM Agents via Simulation
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2025)
Attacking Vision-Language Computer Agents via Pop-ups
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
DECEPTICON: How Dark Patterns Manipulate Web Agents
von: Cuvin, Phil, et al.
Veröffentlicht: (2025)
von: Cuvin, Phil, et al.
Veröffentlicht: (2025)
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping
von: Li, Ryan, et al.
Veröffentlicht: (2024)
von: Li, Ryan, et al.
Veröffentlicht: (2024)
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
von: Li, Minzhi, et al.
Veröffentlicht: (2023)
von: Li, Minzhi, et al.
Veröffentlicht: (2023)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
Auditing Gender Presentation Differences in Text-to-Image Models
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles
von: Kruk, Julia, et al.
Veröffentlicht: (2024)
von: Kruk, Julia, et al.
Veröffentlicht: (2024)
CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies
von: Shi, Weiyan, et al.
Veröffentlicht: (2024)
von: Shi, Weiyan, et al.
Veröffentlicht: (2024)
Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use
von: Fu, Yicheng, et al.
Veröffentlicht: (2025)
von: Fu, Yicheng, et al.
Veröffentlicht: (2025)
Generative Interfaces for Language Models
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
von: Liu, Zijun, et al.
Veröffentlicht: (2023)
von: Liu, Zijun, et al.
Veröffentlicht: (2023)
Online Rubrics Elicitation from Pairwise Comparisons
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025)
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025)
AI in Mental Health: Emotional and Sentiment Analysis of Large Language Models' Responses to Depression, Anxiety, and Stress Queries
von: VarastehNezhad, Arya, et al.
Veröffentlicht: (2025)
von: VarastehNezhad, Arya, et al.
Veröffentlicht: (2025)
Anchor Points: Benchmarking Models with Much Fewer Examples
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
Distilling an End-to-End Voice Assistant Without Instruction Training Data
von: Held, William, et al.
Veröffentlicht: (2024)
von: Held, William, et al.
Veröffentlicht: (2024)
The Call for Socially Aware Language Technologies
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
Benchmarking Machine Translation with Cultural Awareness
von: Yao, Binwei, et al.
Veröffentlicht: (2023)
von: Yao, Binwei, et al.
Veröffentlicht: (2023)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
SocialGaze: Improving the Integration of Human Social Norms in Large Language Models
von: Vijjini, Anvesh Rao, et al.
Veröffentlicht: (2024)
von: Vijjini, Anvesh Rao, et al.
Veröffentlicht: (2024)
Dynamic Skill Adaptation for Large Language Models
von: Chen, Jiaao, et al.
Veröffentlicht: (2024)
von: Chen, Jiaao, et al.
Veröffentlicht: (2024)
Contextualized Privacy Defense for LLM Agents
von: Wen, Yule, et al.
Veröffentlicht: (2026)
von: Wen, Yule, et al.
Veröffentlicht: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Algorithmic Thinking Theory
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025)
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
NormXLogit: The Head-on-Top Never Lies
von: Abbasi, Sina, et al.
Veröffentlicht: (2024)
von: Abbasi, Sina, et al.
Veröffentlicht: (2024)
EgoSocialArena: Benchmarking the Social Intelligence of Large Language Models from a First-person Perspective
von: Hou, Guiyang, et al.
Veröffentlicht: (2024)
von: Hou, Guiyang, et al.
Veröffentlicht: (2024)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Paraphrasing in Affirmative Terms Improves Negation Understanding
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2024) -
Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future
von: Li, Minzhi, et al.
Veröffentlicht: (2024) -
Making Language Models Robust Against Negation
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025) -
Commonsense Knowledge with Negation: A Resource to Enhance Negation Understanding
von: Wang, Zijie, et al.
Veröffentlicht: (2026) -
Social Skill Training with Large Language Models
von: Yang, Diyi, et al.
Veröffentlicht: (2024)