MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Kai, Ling, Zhan, Liu, Kang, Yang, Yifan, Fan, Ting-Han, Shen, Lingfeng, Du, Zhengyin, Chen, Jiecao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
von: Ling, Zhan, et al.
Veröffentlicht: (2025)
von: Ling, Zhan, et al.
Veröffentlicht: (2025)
Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?
von: Yan, Kai, et al.
Veröffentlicht: (2025)
von: Yan, Kai, et al.
Veröffentlicht: (2025)
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
von: Lu, Miao, et al.
Veröffentlicht: (2025)
von: Lu, Miao, et al.
Veröffentlicht: (2025)
Scaling Long-Horizon LLM Agent via Context-Folding
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
von: Yuan, Siyu, et al.
Veröffentlicht: (2025)
von: Yuan, Siyu, et al.
Veröffentlicht: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
OckBench: Measuring the Efficiency of LLM Reasoning
von: Du, Zheng, et al.
Veröffentlicht: (2025)
von: Du, Zheng, et al.
Veröffentlicht: (2025)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
von: Hong, Joey, et al.
Veröffentlicht: (2025)
von: Hong, Joey, et al.
Veröffentlicht: (2025)
ConvexBench: Can LLMs Recognize Convex Functions?
von: Liu, Yepeng, et al.
Veröffentlicht: (2026)
von: Liu, Yepeng, et al.
Veröffentlicht: (2026)
Can Large Language Models Really Recognize Your Name?
von: Pham, Dzung, et al.
Veröffentlicht: (2025)
von: Pham, Dzung, et al.
Veröffentlicht: (2025)
Many-Shot In-Context Learning
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2024)
MAPLE: Many-Shot Adaptive Pseudo-Labeling for In-Context Learning
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
Compressing Many-Shots in In-Context Learning
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2024)
Towards Compute-Optimal Many-Shot In-Context Learning
von: Golchin, Shahriar, et al.
Veröffentlicht: (2025)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2025)
On Many-Shot In-Context Learning for Long-Context Evaluation
von: Zou, Kaijian, et al.
Veröffentlicht: (2024)
von: Zou, Kaijian, et al.
Veröffentlicht: (2024)
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
von: Song, Mingyang, et al.
Veröffentlicht: (2024)
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation
von: Wan, Xingchen, et al.
Veröffentlicht: (2025)
von: Wan, Xingchen, et al.
Veröffentlicht: (2025)
Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems
von: Zhu, Junze, et al.
Veröffentlicht: (2026)
von: Zhu, Junze, et al.
Veröffentlicht: (2026)
RedacBench: Can AI Erase Your Secrets?
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
BenchScope: How Many Independent Signals Does Your Benchmark Provide?
von: Sha, Tommy, et al.
Veröffentlicht: (2026)
von: Sha, Tommy, et al.
Veröffentlicht: (2026)
Still's Lung Disease: Time to Recognize This Complication in Adults
von: Lauren A. Henderson
Veröffentlicht: (2025)
von: Lauren A. Henderson
Veröffentlicht: (2025)
Evaluating Zero-Shot Long-Context LLM Compression
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
Many-Shot In-Context Learning for Molecular Inverse Design
von: Moayedpour, Saeed, et al.
Veröffentlicht: (2024)
von: Moayedpour, Saeed, et al.
Veröffentlicht: (2024)
Can this Model Also Recognize Dogs? Zero-Shot Model Search from Weights
von: Kahana, Jonathan, et al.
Veröffentlicht: (2025)
von: Kahana, Jonathan, et al.
Veröffentlicht: (2025)
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
Fewer is More: Boosting LLM Reasoning with Reinforced Context Pruning
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
Originality: Who Can Recognize It?
von: West, Susie
Veröffentlicht: (1976)
von: West, Susie
Veröffentlicht: (1976)
Distilling Many-Shot In-Context Learning into a Cheat Sheet
von: Honda, Ukyo, et al.
Veröffentlicht: (2025)
von: Honda, Ukyo, et al.
Veröffentlicht: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
Emotions are Recognized Patterns of Cognitive Activities
von: Jin, Yue
Veröffentlicht: (2025)
von: Jin, Yue
Veröffentlicht: (2025)
FREYR: A Framework for Recognizing and Executing Your Requests
von: Gallotta, Roberto, et al.
Veröffentlicht: (2025)
von: Gallotta, Roberto, et al.
Veröffentlicht: (2025)
AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache Reuse
von: Ou, Jie, et al.
Veröffentlicht: (2026)
von: Ou, Jie, et al.
Veröffentlicht: (2026)
From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
von: Du, Hang, et al.
Veröffentlicht: (2025)
von: Du, Hang, et al.
Veröffentlicht: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
von: Zhang, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2025)
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
von: Xiao, Emily, et al.
Veröffentlicht: (2025)
von: Xiao, Emily, et al.
Veröffentlicht: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
von: Ling, Zhan, et al.
Veröffentlicht: (2025) -
Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?
von: Yan, Kai, et al.
Veröffentlicht: (2025) -
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
von: Lu, Miao, et al.
Veröffentlicht: (2025) -
Scaling Long-Horizon LLM Agent via Context-Folding
von: Sun, Weiwei, et al.
Veröffentlicht: (2025) -
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
von: Yuan, Siyu, et al.
Veröffentlicht: (2025)