AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yifan, Zhao, Wenkuan, Zhong, Shanshan, Qin, Jinghui, Liang, Mingfu, Huang, Zhongzhan, Wen, Wushao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ASR: Attention-alike Structural Re-parameterization
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima
by: Zhong, Shanshan, et al.
Published: (2024)
by: Zhong, Shanshan, et al.
Published: (2024)
Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
MoExtend: Tuning New Experts for Modality and Task Extension
by: Zhong, Shanshan, et al.
Published: (2024)
by: Zhong, Shanshan, et al.
Published: (2024)
Boundary-Driven Table-Filling with Cross-Granularity Contrastive Learning for Aspect Sentiment Triplet Extraction
by: Li, Qingling, et al.
Published: (2025)
by: Li, Qingling, et al.
Published: (2025)
AttNS: Attention-Inspired Numerical Solving For Limited Data Scenarios
by: Huang, Zhongzhan, et al.
Published: (2023)
by: Huang, Zhongzhan, et al.
Published: (2023)
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
by: Huang, Zhongzhan, et al.
Published: (2025)
by: Huang, Zhongzhan, et al.
Published: (2025)
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
by: Huang, Zhongzhan, et al.
Published: (2025)
by: Huang, Zhongzhan, et al.
Published: (2025)
AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models
by: Ling, Guoming, et al.
Published: (2026)
by: Ling, Guoming, et al.
Published: (2026)
A Generic Shared Attention Mechanism for Various Backbone Neural Networks
by: Huang, Zhongzhan, et al.
Published: (2022)
by: Huang, Zhongzhan, et al.
Published: (2022)
I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
by: Kharchenko, Julia, et al.
Published: (2025)
by: Kharchenko, Julia, et al.
Published: (2025)
CCiV: A Benchmark for Structure, Rhythm and Quality in LLM-Generated Chinese \textit{Ci} Poetry
by: Zhao, Shangqing, et al.
Published: (2026)
by: Zhao, Shangqing, et al.
Published: (2026)
Can Speech LLMs Think while Listening?
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
ThinkSwitcher: When to Think Hard, When to Think Fast
by: Liang, Guosheng, et al.
Published: (2025)
by: Liang, Guosheng, et al.
Published: (2025)
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
by: Huang, Zhongzhan, et al.
Published: (2025)
by: Huang, Zhongzhan, et al.
Published: (2025)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
by: Huang, Shulin, et al.
Published: (2023)
by: Huang, Shulin, et al.
Published: (2023)
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation
by: Du, Xuan, et al.
Published: (2026)
by: Du, Xuan, et al.
Published: (2026)
CLEAR-KGQA: Clarification-Enhanced Ambiguity Resolution for Knowledge Graph Question Answering
by: Wen, Liqiang, et al.
Published: (2025)
by: Wen, Liqiang, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
by: Lu, Qiuyang, et al.
Published: (2025)
by: Lu, Qiuyang, et al.
Published: (2025)
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
Evaluating Chinese Ambiguity Understanding in Large Language Models
by: Mo, Junwen, et al.
Published: (2026)
by: Mo, Junwen, et al.
Published: (2026)
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
by: Du, Yanrui, et al.
Published: (2023)
by: Du, Yanrui, et al.
Published: (2023)
Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
by: Luo, Wenzhen, et al.
Published: (2025)
by: Luo, Wenzhen, et al.
Published: (2025)
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
CiMaTe: Citation Count Prediction Effectively Leveraging the Main Text
by: Hirako, Jun, et al.
Published: (2024)
by: Hirako, Jun, et al.
Published: (2024)
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
by: Wen, Hao, et al.
Published: (2025)
by: Wen, Hao, et al.
Published: (2025)
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
by: Zeng, Zhongshen, et al.
Published: (2024)
by: Zeng, Zhongshen, et al.
Published: (2024)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
by: Cheng, Xiaoxue, et al.
Published: (2025)
by: Cheng, Xiaoxue, et al.
Published: (2025)
CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization
by: Li, Junyi, et al.
Published: (2026)
by: Li, Junyi, et al.
Published: (2026)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
by: Jia, Qi, et al.
Published: (2024)
by: Jia, Qi, et al.
Published: (2024)
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
by: Zhang, Zhibo, et al.
Published: (2025)
by: Zhang, Zhibo, et al.
Published: (2025)
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
by: Park, Jeonghyun, et al.
Published: (2025)
by: Park, Jeonghyun, et al.
Published: (2025)
Achromoporus andujari Perez-Asso 2009
by: Bouzan, Rodrigo S., et al.
Published: (2026)
by: Bouzan, Rodrigo S., et al.
Published: (2026)
Similar Items
-
ASR: Attention-alike Structural Re-parameterization
by: Zhong, Shanshan, et al.
Published: (2023) -
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
by: Zhong, Shanshan, et al.
Published: (2023) -
Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima
by: Zhong, Shanshan, et al.
Published: (2024) -
Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation
by: Wang, Hui, et al.
Published: (2025) -
MoExtend: Tuning New Experts for Modality and Task Extension
by: Zhong, Shanshan, et al.
Published: (2024)