AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Xiao, Wang, Andrew, Choi, Jacob, Lu, Yining, Sharma, Shreya, Shen, Lingfeng, Tiyyala, Vijay, Andrews, Nicholas, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution
by: Lu, Yining, et al.
Published: (2023)
by: Lu, Yining, et al.
Published: (2023)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
by: Jiang, Dongwei, et al.
Published: (2025)
by: Jiang, Dongwei, et al.
Published: (2025)
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025)
by: Huang, Heyuan, et al.
Published: (2025)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
by: Fang, Shicheng, et al.
Published: (2026)
by: Fang, Shicheng, et al.
Published: (2026)
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
by: Tan, Weiting, et al.
Published: (2024)
by: Tan, Weiting, et al.
Published: (2024)
Benchmarking Language Model Creativity: A Case Study on Code Generation
by: Lu, Yining, et al.
Published: (2024)
by: Lu, Yining, et al.
Published: (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
AnalogNAS-Bench: A NAS Benchmark for Analog In-Memory Computing
by: Bessalah, Aniss, et al.
Published: (2025)
by: Bessalah, Aniss, et al.
Published: (2025)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025)
by: Cai, Muzhen, et al.
Published: (2025)
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
by: Byerly, Adam, et al.
Published: (2024)
by: Byerly, Adam, et al.
Published: (2024)
RiskBench: A Scenario-based Benchmark for Risk Identification
by: Kung, Chi-Hsi, et al.
Published: (2023)
by: Kung, Chi-Hsi, et al.
Published: (2023)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
by: Uzunoglu, Arda, et al.
Published: (2025)
by: Uzunoglu, Arda, et al.
Published: (2025)
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
TeleEmbedBench: A Multi-Corpus Embedding Benchmark for RAG in Telecommunications
by: Gajjar, Pranshav, et al.
Published: (2026)
by: Gajjar, Pranshav, et al.
Published: (2026)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
by: Bai, Yushi, et al.
Published: (2023)
by: Bai, Yushi, et al.
Published: (2023)
RORA: Robust Free-Text Rationale Evaluation
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Actions on the Picard group of smooth Fano threefolds
by: Sharma, Shreya
Published: (2025)
by: Sharma, Shreya
Published: (2025)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Characterizing public comments via Regulations.gov in response to proposed cannabis rescheduling in the United States
by: Vijay M. Tiyyala, et al.
Published: (2026)
by: Vijay M. Tiyyala, et al.
Published: (2026)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
MileBench: Benchmarking MLLMs in Long Context
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
ORAN-Bench-13K: An Open Source Benchmark for Assessing LLMs in Open Radio Access Networks
by: Gajjar, Pranshav, et al.
Published: (2024)
by: Gajjar, Pranshav, et al.
Published: (2024)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
by: Hu, Caiyu, et al.
Published: (2025)
by: Hu, Caiyu, et al.
Published: (2025)
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
by: Ruan, Jie, et al.
Published: (2025)
by: Ruan, Jie, et al.
Published: (2025)
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
by: Ling, Zhan, et al.
Published: (2025)
by: Ling, Zhan, et al.
Published: (2025)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Landsat-Bench: Datasets and Benchmarks for Landsat Foundation Models
by: Corley, Isaac, et al.
Published: (2025)
by: Corley, Isaac, et al.
Published: (2025)
OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology
by: Zhou, Chengfeng, et al.
Published: (2025)
by: Zhou, Chengfeng, et al.
Published: (2025)
Interdiscursive collaboration in public relations contexts
by: Vijay K. Bhatia
Published: (2013)
by: Vijay K. Bhatia
Published: (2013)
AMS-IO-Bench and AMS-IO-Agent: Benchmarking and Structured Reasoning for Analog and Mixed-Signal Integrated Circuit Input/Output Design
by: Zhang, Zhishuai, et al.
Published: (2025)
by: Zhang, Zhishuai, et al.
Published: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
by: Zhou, Wenqi, et al.
Published: (2025)
by: Zhou, Wenqi, et al.
Published: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
Effect of Accretion on the evolution of Primordial Black Holes in the context of Modified Gravity Theories
by: Banerjee, Shreya, et al.
Published: (2024)
by: Banerjee, Shreya, et al.
Published: (2024)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
Similar Items
-
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023) -
GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution
by: Lu, Yining, et al.
Published: (2023) -
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024) -
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
by: Jiang, Dongwei, et al.
Published: (2025) -
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)