ExpressivityBench: Can LLMs Communicate Implicitly?
Fuente:
arXiv
Saved in:
| Main Authors: | Tint, Joshua, Sagar, Som, Taparia, Aditya, Raines, Kelly, Pathiraja, Bimsara, Liu, Caleb, Senanayake, Ransalu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language Models
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
Fairness in Autonomous Driving: Towards Understanding Confounding Factors in Object Detection under Challenging Weather
by: Pathiraja, Bimsara, et al.
Published: (2024)
by: Pathiraja, Bimsara, et al.
Published: (2024)
Explainable Concept Generation through Vision-Language Preference Learning for Understanding Neural Networks' Internal Representations
by: Taparia, Aditya, et al.
Published: (2024)
by: Taparia, Aditya, et al.
Published: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Personality, Role, and Expressive Style in Large Language Models: An Interactionist Analysis
by: Nagao, Moe, et al.
Published: (2026)
by: Nagao, Moe, et al.
Published: (2026)
Can LLMs Compute with Reasons?
by: Sandilya, Harshit, et al.
Published: (2024)
by: Sandilya, Harshit, et al.
Published: (2024)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
by: Feucht, Sheridan, et al.
Published: (2024)
by: Feucht, Sheridan, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
AI Can Learn Scientific Taste
by: Tong, Jingqi, et al.
Published: (2026)
by: Tong, Jingqi, et al.
Published: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
by: Tu, Songjun, et al.
Published: (2026)
by: Tu, Songjun, et al.
Published: (2026)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation
by: Yang, Pei, et al.
Published: (2026)
by: Yang, Pei, et al.
Published: (2026)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
by: Yu, Jinzheng, et al.
Published: (2025)
by: Yu, Jinzheng, et al.
Published: (2025)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
by: Zhang, Yizhuo, et al.
Published: (2024)
by: Zhang, Yizhuo, et al.
Published: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
BlasBench: An Open Benchmark for Irish Speech Recognition
by: Raj, Jyoutir, et al.
Published: (2026)
by: Raj, Jyoutir, et al.
Published: (2026)
XferBench: a Data-Driven Benchmark for Emergent Language
by: Boldt, Brendon, et al.
Published: (2024)
by: Boldt, Brendon, et al.
Published: (2024)
Vocabulary Transfer for Biomedical Texts: Add Tokens if You Can Not Add Data
by: Singh, Priyanka, et al.
Published: (2022)
by: Singh, Priyanka, et al.
Published: (2022)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
by: Hosseini, Peyman, et al.
Published: (2024)
by: Hosseini, Peyman, et al.
Published: (2024)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
by: Fagnou, Erwan, et al.
Published: (2026)
by: Fagnou, Erwan, et al.
Published: (2026)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
by: Qi, Jinhu, et al.
Published: (2024)
by: Qi, Jinhu, et al.
Published: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
by: Zheng, Qinyue, et al.
Published: (2025)
by: Zheng, Qinyue, et al.
Published: (2025)
Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task
by: Evelo, Bart, et al.
Published: (2026)
by: Evelo, Bart, et al.
Published: (2026)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis
by: Gandarela, João Pedro, et al.
Published: (2024)
by: Gandarela, João Pedro, et al.
Published: (2024)
Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation
by: Nasim, Mehwish, et al.
Published: (2026)
by: Nasim, Mehwish, et al.
Published: (2026)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
by: Šindelář, Pavel, et al.
Published: (2025)
by: Šindelář, Pavel, et al.
Published: (2025)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
by: Ewais, Ahmed, et al.
Published: (2026)
by: Ewais, Ahmed, et al.
Published: (2026)
MALT: Mechanistic Ablation of Lossy Translation in LLMs for a Low-Resource Language: Urdu
by: Bajwa, Taaha Saleem
Published: (2025)
by: Bajwa, Taaha Saleem
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
by: Bajpai, Ashutosh, et al.
Published: (2025)
by: Bajpai, Ashutosh, et al.
Published: (2025)
Similar Items
-
LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"
by: Sagar, Som, et al.
Published: (2024) -
Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language Models
by: Sagar, Som, et al.
Published: (2024) -
Fairness in Autonomous Driving: Towards Understanding Confounding Factors in Object Detection under Challenging Weather
by: Pathiraja, Bimsara, et al.
Published: (2024) -
Explainable Concept Generation through Vision-Language Preference Learning for Understanding Neural Networks' Internal Representations
by: Taparia, Aditya, et al.
Published: (2024) -
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)