Saved in:
| Main Authors: | Chang, Hao, Wang, Zhihui, Wu, Lingxiang, An, Wei, Li, Boyang, Lin, Zaiping, Sheng, Weidong, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.19640 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Chain of Thought Prompting in Large Language Models via Reasoning Patterns
by: Zhang, Yufeng, et al.
Published: (2024)
by: Zhang, Yufeng, et al.
Published: (2024)
Australia's Wellbeing Framework: Is It Really ‘Measuring What Matters’?
by: Kate Sollis, et al.
Published: (2025)
by: Kate Sollis, et al.
Published: (2025)
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Enhancing Text-to-SQL Capabilities of Large Language Models via Domain Database Knowledge Injection
by: Ma, Xingyu, et al.
Published: (2024)
by: Ma, Xingyu, et al.
Published: (2024)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Rethinking Household Food Waste: What Really Matters in Everyday Food Management
by: Lucie Veselá, et al.
Published: (2026)
by: Lucie Veselá, et al.
Published: (2026)
Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
What Really Matters in Matrix-Whitening Optimizers?
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
by: He, Chenwei, et al.
Published: (2026)
by: He, Chenwei, et al.
Published: (2026)
PFDM: Parser-Free Virtual Try-on via Diffusion Model
by: Niu, Yunfang, et al.
Published: (2024)
by: Niu, Yunfang, et al.
Published: (2024)
What do Blind and Low-Vision People Really Want from Assistive Smart Devices? Comparison of the Literature with a Focus Study
by: Gamage, Bhanuka, et al.
Published: (2025)
by: Gamage, Bhanuka, et al.
Published: (2025)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
by: Yu, Yijiong
Published: (2024)
by: Yu, Yijiong
Published: (2024)
Testing Autonomous Driving Systems -- What Really Matters and What Doesn't
by: Li, Changwen, et al.
Published: (2025)
by: Li, Changwen, et al.
Published: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
by: Guo, Longteng, et al.
Published: (2026)
by: Guo, Longteng, et al.
Published: (2026)
Visible-Thermal Tiny Object Detection: A Benchmark Dataset and Baselines
by: Ying, Xinyi, et al.
Published: (2024)
by: Ying, Xinyi, et al.
Published: (2024)
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
Coordinated Beamforming for RIS-Empowered ISAC Systems over Secure Low-Altitude Networks
by: Wang, Chunjie, et al.
Published: (2025)
by: Wang, Chunjie, et al.
Published: (2025)
RRCANet: Recurrent Reusable-Convolution Attention Network for Infrared Small Target Detection
by: Liu, Yongxian, et al.
Published: (2025)
by: Liu, Yongxian, et al.
Published: (2025)
AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
by: Niu, Yunfang, et al.
Published: (2024)
by: Niu, Yunfang, et al.
Published: (2024)
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2024)
by: Wang, Langyu, et al.
Published: (2024)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025)
by: Hao, Haihong, et al.
Published: (2025)
Shared Sky, Shared Spectrum: Coordinated Satellite-5G Networks for Low-Altitude Economy
by: Wang, Yanmin, et al.
Published: (2026)
by: Wang, Yanmin, et al.
Published: (2026)
What Really is Commonsense Knowledge?
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
by: Wang, Yuandong, et al.
Published: (2025)
by: Wang, Yuandong, et al.
Published: (2025)
Efficient Coordination Among Chinese Provinces in Managing Supply and Demand for Staple Crops
by: Yifei Wang, et al.
Published: (2024)
by: Yifei Wang, et al.
Published: (2024)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
by: Wang, Yunnan, et al.
Published: (2025)
by: Wang, Yunnan, et al.
Published: (2025)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point Supervision
by: Ying, Xinyi, et al.
Published: (2023)
by: Ying, Xinyi, et al.
Published: (2023)
What Really Matters for Robust Multi-Sensor HD Map Construction?
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
What Really Matters for Learning-based LiDAR-Camera Calibration
by: Huang, Shujuan, et al.
Published: (2025)
by: Huang, Shujuan, et al.
Published: (2025)
Earnings Quality and ESG Performance in Energy and Utilities: What Really Matters?
by: Antonios Persakis, et al.
Published: (2025)
by: Antonios Persakis, et al.
Published: (2025)
Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Data-Centric AI Governance: Addressing the Limitations of Model-Focused Policies
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
The Things That Really Matter
Published: (2022)
Published: (2022)
Topology-Aware Coordination for Multi-Functional Low-Altitude Wireless Networks
by: He, Jiajun, et al.
Published: (2026)
by: He, Jiajun, et al.
Published: (2026)
Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation
by: Zhang, Ruiqi, et al.
Published: (2026)
by: Zhang, Ruiqi, et al.
Published: (2026)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Agentic AI for Low-Altitude Semantic Wireless Networks: An Energy Efficient Design
by: Zhao, Zhouxiang, et al.
Published: (2025)
by: Zhao, Zhouxiang, et al.
Published: (2025)
Similar Items
-
Enhancing Chain of Thought Prompting in Large Language Models via Reasoning Patterns
by: Zhang, Yufeng, et al.
Published: (2024) -
Australia's Wellbeing Framework: Is It Really ‘Measuring What Matters’?
by: Kate Sollis, et al.
Published: (2025) -
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
by: Yu, Jiachen, et al.
Published: (2025) -
Enhancing Text-to-SQL Capabilities of Large Language Models via Domain Database Knowledge Injection
by: Ma, Xingyu, et al.
Published: (2024) -
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)