Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahmadpour, Mohammadjavad, Meighani, Amirmahdi, Taebi, Payam, Ghahroodi, Omid, Izadi, Amirmohammad, Baghshah, Mahdieh Soleymani |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
por: Saghafian, Armin, et al.
Publicado: (2024)
por: Saghafian, Armin, et al.
Publicado: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
CER: Confidence Enhanced Reasoning in LLMs
por: Razghandi, Ali, et al.
Publicado: (2025)
por: Razghandi, Ali, et al.
Publicado: (2025)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
por: Hasani, Hosein, et al.
Publicado: (2025)
por: Hasani, Hosein, et al.
Publicado: (2025)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
por: Mehri, Faridoun, et al.
Publicado: (2024)
por: Mehri, Faridoun, et al.
Publicado: (2024)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2026)
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2026)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
por: Hosseini, Seyed Mohsen, et al.
Publicado: (2024)
por: Hosseini, Seyed Mohsen, et al.
Publicado: (2024)
PreND: Enhancing Intrinsic Motivation in Reinforcement Learning through Pre-trained Network Distillation
por: Davoodabadi, Mohammadamin, et al.
Publicado: (2024)
por: Davoodabadi, Mohammadamin, et al.
Publicado: (2024)
Classification of Breast Cancer Histopathology Images using a Modified Supervised Contrastive Learning Method
por: Sani, Matina Mahdizadeh, et al.
Publicado: (2024)
por: Sani, Matina Mahdizadeh, et al.
Publicado: (2024)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
por: Zohrabi, Reihaneh, et al.
Publicado: (2026)
por: Zohrabi, Reihaneh, et al.
Publicado: (2026)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
por: Marioriyad, Arash, et al.
Publicado: (2026)
por: Marioriyad, Arash, et al.
Publicado: (2026)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2026)
por: Hosseini, Seyed Mohammad Hadi, et al.
Publicado: (2026)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
por: Abbasi, Reza, et al.
Publicado: (2024)
por: Abbasi, Reza, et al.
Publicado: (2024)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
por: Mozafari, Mohammad, et al.
Publicado: (2024)
por: Mozafari, Mohammad, et al.
Publicado: (2024)
ComAlign: Compositional Alignment in Vision-Language Models
por: Abdollah, Ali, et al.
Publicado: (2024)
por: Abdollah, Ali, et al.
Publicado: (2024)
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
por: Ghahroodi, Omid, et al.
Publicado: (2025)
por: Ghahroodi, Omid, et al.
Publicado: (2025)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
por: Izadi, Amir Mohammad, et al.
Publicado: (2025)
por: Izadi, Amir Mohammad, et al.
Publicado: (2025)
Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
por: Hasani, Hosein, et al.
Publicado: (2025)
por: Hasani, Hosein, et al.
Publicado: (2025)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
por: Dijujin, Negin Hashemi, et al.
Publicado: (2025)
por: Dijujin, Negin Hashemi, et al.
Publicado: (2025)
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
ELAB: Extensive LLM Alignment Benchmark in Persian Language
por: Pourbahman, Zahra, et al.
Publicado: (2025)
por: Pourbahman, Zahra, et al.
Publicado: (2025)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
por: Hasani, Hosein, et al.
Publicado: (2026)
por: Hasani, Hosein, et al.
Publicado: (2026)
Causal Attribution via Activation Patching
por: Izadi, Amirmohammad, et al.
Publicado: (2026)
por: Izadi, Amirmohammad, et al.
Publicado: (2026)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
por: Zohrabi, Reihaneh, et al.
Publicado: (2025)
por: Zohrabi, Reihaneh, et al.
Publicado: (2025)
Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation
por: Noohdani, Fahimeh Hosseini, et al.
Publicado: (2024)
por: Noohdani, Fahimeh Hosseini, et al.
Publicado: (2024)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
por: Abdollahi, Ali, et al.
Publicado: (2024)
por: Abdollahi, Ali, et al.
Publicado: (2024)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
por: Ghahroodi, Omid, et al.
Publicado: (2024)
por: Ghahroodi, Omid, et al.
Publicado: (2024)
The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
por: Samiei, Mahdi, et al.
Publicado: (2025)
por: Samiei, Mahdi, et al.
Publicado: (2025)
SOInter: A Novel Deep Energy Based Interpretation Method for Explaining Structured Output Models
por: Seyyedsalehi, S. Fatemeh, et al.
Publicado: (2022)
por: Seyyedsalehi, S. Fatemeh, et al.
Publicado: (2022)
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
por: Cekinmez, Jasin, et al.
Publicado: (2025)
por: Cekinmez, Jasin, et al.
Publicado: (2025)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2025)
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2025)
Statistically Valid Information Bottleneck via Multiple Hypothesis Testing
por: Farzaneh, Amirmohammad, et al.
Publicado: (2024)
por: Farzaneh, Amirmohammad, et al.
Publicado: (2024)
Multi-Objective Hyperparameter Selection via Hypothesis Testing on Reliability Graphs
por: Farzaneh, Amirmohammad, et al.
Publicado: (2025)
por: Farzaneh, Amirmohammad, et al.
Publicado: (2025)
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
por: Lin, Junhong, et al.
Publicado: (2025)
por: Lin, Junhong, et al.
Publicado: (2025)
Efficient Test-Time Scaling for Small Vision-Language Models
por: Kaya, Mehmet Onurcan, et al.
Publicado: (2025)
por: Kaya, Mehmet Onurcan, et al.
Publicado: (2025)
LLM-ODE: Data-driven Discovery of Dynamical Systems with Large Language Models
por: Bideh, Amirmohammad Ziaei, et al.
Publicado: (2026)
por: Bideh, Amirmohammad Ziaei, et al.
Publicado: (2026)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
por: Chen, Feng, et al.
Publicado: (2025)
por: Chen, Feng, et al.
Publicado: (2025)
Efficient Reasoning at Fixed Test-Time Cost via Length-Aware Attention Priors and Gain-Aware Training
por: Atri, Rian
Publicado: (2026)
por: Atri, Rian
Publicado: (2026)
Quantile Learn-Then-Test: Quantile-Based Risk Control for Hyperparameter Optimization
por: Farzaneh, Amirmohammad, et al.
Publicado: (2024)
por: Farzaneh, Amirmohammad, et al.
Publicado: (2024)
Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing
por: Soleymani, Mahdi, et al.
Publicado: (2025)
por: Soleymani, Mahdi, et al.
Publicado: (2025)
Ejemplares similares
-
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
por: Saghafian, Armin, et al.
Publicado: (2024) -
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
por: Izadi, Amirmohammad, et al.
Publicado: (2025) -
CER: Confidence Enhanced Reasoning in LLMs
por: Razghandi, Ali, et al.
Publicado: (2025) -
Understanding Counting Mechanisms in Large Language and Vision-Language Models
por: Hasani, Hosein, et al.
Publicado: (2025) -
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
por: Mehri, Faridoun, et al.
Publicado: (2024)