REBUS: A Robust Evaluation Benchmark of Understanding Symbols
Fuente:
arXiv
Saved in:
| Main Authors: | Gritsevskiy, Andrew, Panickssery, Arjun, Kirtland, Aaron, Kauffman, Derik, Gundlach, Hans, Gritsevskaya, Irina, Cavanagh, Joe, Chiang, Jonathan, La Roux, Lydia, Hung, Michelle |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
LLM Evaluators Recognize and Favor Their Own Generations
by: Panickssery, Arjun, et al.
Published: (2024)
by: Panickssery, Arjun, et al.
Published: (2024)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)
by: Ball, Sarah, et al.
Published: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
by: Ackerman, Christopher, et al.
Published: (2024)
by: Ackerman, Christopher, et al.
Published: (2024)
LOGIC-LM++: Multi-Step Refinement for Symbolic Formulations
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
by: Faure, Gueter Josmy, et al.
Published: (2026)
by: Faure, Gueter Josmy, et al.
Published: (2026)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
by: Shalyt, Michael, et al.
Published: (2025)
by: Shalyt, Michael, et al.
Published: (2025)
OntoURL: A Benchmark for Evaluating Large Language Models on Symbolic Ontological Understanding, Reasoning and Learning
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
To Tailor Schedules, Students Log in to Online Classes
by: Cavanagh, Sean
Published: (2006)
by: Cavanagh, Sean
Published: (2006)
Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations
by: Chiang, Cheng-Han, et al.
Published: (2024)
by: Chiang, Cheng-Han, et al.
Published: (2024)
Over-Reasoning and Redundant Calculation of Large Language Models
by: Chiang, Cheng-Han, et al.
Published: (2024)
by: Chiang, Cheng-Han, et al.
Published: (2024)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
Rethinking Symbolic Regression Datasets and Benchmarks for Scientific Discovery
by: Matsubara, Yoshitomo, et al.
Published: (2022)
by: Matsubara, Yoshitomo, et al.
Published: (2022)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
by: Chiang, Cheng-Han, et al.
Published: (2025)
by: Chiang, Cheng-Han, et al.
Published: (2025)
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
by: Ayoobi, Navid, et al.
Published: (2025)
by: Ayoobi, Navid, et al.
Published: (2025)
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
by: Dan, Nifu, et al.
Published: (2025)
by: Dan, Nifu, et al.
Published: (2025)
LLM-Symbolic Integration for Robust Temporal Tabular Reasoning
by: Kulkarni, Atharv, et al.
Published: (2025)
by: Kulkarni, Atharv, et al.
Published: (2025)
Making Sentence Embeddings Robust to User-Generated Content
by: Nishimwe, Lydia, et al.
Published: (2024)
by: Nishimwe, Lydia, et al.
Published: (2024)
Once Again, with Style: Understanding and Supporting Partial Reuse in Dashboard Authoring
by: Sultanum, Nicole, et al.
Published: (2026)
by: Sultanum, Nicole, et al.
Published: (2026)
Can Large Language Models Understand Symbolic Graphics Programs?
by: Qiu, Zeju, et al.
Published: (2024)
by: Qiu, Zeju, et al.
Published: (2024)
Dynamic Locality Sensitive Orderings in Doubling Metrics
by: La, An, et al.
Published: (2024)
by: La, An, et al.
Published: (2024)
Safer in Translation? Presupposition Robustness in Indic Languages
by: Palnitkar, Aadi, et al.
Published: (2025)
by: Palnitkar, Aadi, et al.
Published: (2025)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
Prove Symbolic Regression is NP-hard by Symbol Graph
by: Song, Jinglu, et al.
Published: (2024)
by: Song, Jinglu, et al.
Published: (2024)
Symbolic Graph Inference for Compound Scene Understanding
by: Aryan, FNU, et al.
Published: (2024)
by: Aryan, FNU, et al.
Published: (2024)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Interpretable by AI Mother Tongue: Native Symbolic Reasoning in Neural Models
by: Liu, Hung Ming
Published: (2025)
by: Liu, Hung Ming
Published: (2025)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Building Robust and Scalable Multilingual ASR for Indian Languages
by: Gangwar, Arjun, et al.
Published: (2025)
by: Gangwar, Arjun, et al.
Published: (2025)
Exhaustive Symbolic Integration: Integration by Differentiation and the Landscape of Symbolic Integrability
by: Desmond, Harry
Published: (2026)
by: Desmond, Harry
Published: (2026)
The Complexity of Data-Free Nfer
by: Kauffman, Sean, et al.
Published: (2024)
by: Kauffman, Sean, et al.
Published: (2024)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
Similar Items
-
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023) -
LLM Evaluators Recognize and Favor Their Own Generations
by: Panickssery, Arjun, et al.
Published: (2024) -
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024) -
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024) -
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
by: Ackerman, Christopher, et al.
Published: (2024)