SwimBird: Eliciting Switchable Reasoning Mode in Hybrid Autoregressive MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tong, Jintao, Yan, Shilin, Xue, Hongwei, Tang, Xiaojun, Shi, Kunyu, Zhang, Guannan, Li, Ruixuan, Zou, Yixiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
von: Yan, Shilin, et al.
Veröffentlicht: (2026)
von: Yan, Shilin, et al.
Veröffentlicht: (2026)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Lightweight Frequency Masker for Cross-Domain Few-Shot Semantic Segmentation
von: Tong, Jintao, et al.
Veröffentlicht: (2024)
von: Tong, Jintao, et al.
Veröffentlicht: (2024)
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
von: Shen, Haozhan, et al.
Veröffentlicht: (2026)
von: Shen, Haozhan, et al.
Veröffentlicht: (2026)
Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot Segmentation
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Beyond Scattered Acceptance: Fast and Coherent Inference for DLMs via Longest Stable Prefixes
von: Li, Pengxiang, et al.
Veröffentlicht: (2026)
von: Li, Pengxiang, et al.
Veröffentlicht: (2026)
Adapter Naturally Serves as Decoupler for Cross-Domain Few-Shot Semantic Segmentation
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
von: Yi, Shuai, et al.
Veröffentlicht: (2026)
von: Yi, Shuai, et al.
Veröffentlicht: (2026)
Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning
von: Yi, Shuai, et al.
Veröffentlicht: (2025)
von: Yi, Shuai, et al.
Veröffentlicht: (2025)
Revisiting Pool-based Prompt Learning for Few-shot Class-incremental Learning
von: Jiang, Yongwei, et al.
Veröffentlicht: (2025)
von: Jiang, Yongwei, et al.
Veröffentlicht: (2025)
Reconstruction Target Matters in Masked Image Modeling for Cross-Domain Few-Shot Learning
von: Ma, Ran, et al.
Veröffentlicht: (2024)
von: Ma, Ran, et al.
Veröffentlicht: (2024)
Random Registers for Cross-Domain Few-Shot Learning
von: Yi, Shuai, et al.
Veröffentlicht: (2025)
von: Yi, Shuai, et al.
Veröffentlicht: (2025)
Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection
von: Jiang, Yongwei, et al.
Veröffentlicht: (2026)
von: Jiang, Yongwei, et al.
Veröffentlicht: (2026)
Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning
von: Yi, Shuai, et al.
Veröffentlicht: (2026)
von: Yi, Shuai, et al.
Veröffentlicht: (2026)
FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
The Devil is in Low-Level Features for Cross-Domain Few-Shot Segmentation
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment
von: Zhao, Yaze, et al.
Veröffentlicht: (2026)
von: Zhao, Yaze, et al.
Veröffentlicht: (2026)
Learning Unknowns from Unknowns: Diversified Negative Prototypes Generator for Few-Shot Open-Set Recognition
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2026)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2026)
Flatten Long-Range Loss Landscapes for Cross-Domain Few-Shot Learning
von: Zou, Yixiong, et al.
Veröffentlicht: (2024)
von: Zou, Yixiong, et al.
Veröffentlicht: (2024)
Delve into Base-Novel Confusion: Redundancy Exploration for Few-Shot Class-Incremental Learning
von: Zhou, Haichen, et al.
Veröffentlicht: (2024)
von: Zhou, Haichen, et al.
Veröffentlicht: (2024)
Compositional Few-Shot Class-Incremental Learning
von: Zou, Yixiong, et al.
Veröffentlicht: (2024)
von: Zou, Yixiong, et al.
Veröffentlicht: (2024)
Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning
von: Zhao, Yaze, et al.
Veröffentlicht: (2026)
von: Zhao, Yaze, et al.
Veröffentlicht: (2026)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
von: Liu, Xianjie, et al.
Veröffentlicht: (2026)
von: Liu, Xianjie, et al.
Veröffentlicht: (2026)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
von: Liu, Xianjie, et al.
Veröffentlicht: (2025)
von: Liu, Xianjie, et al.
Veröffentlicht: (2025)
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
von: Zheng, Guanghao, et al.
Veröffentlicht: (2025)
von: Zheng, Guanghao, et al.
Veröffentlicht: (2025)
MICM: Rethinking Unsupervised Pretraining for Enhanced Few-shot Learning
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding
von: Yan, Qingyang, et al.
Veröffentlicht: (2025)
von: Yan, Qingyang, et al.
Veröffentlicht: (2025)
RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
von: Lawrence, Logan, et al.
Veröffentlicht: (2026)
von: Lawrence, Logan, et al.
Veröffentlicht: (2026)
Speculative Decoding for Autoregressive Video Generation
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
von: Zhang, Xintong, et al.
Veröffentlicht: (2026)
von: Zhang, Xintong, et al.
Veröffentlicht: (2026)
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
von: Chen, Jintao, et al.
Veröffentlicht: (2026)
von: Chen, Jintao, et al.
Veröffentlicht: (2026)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
von: Chen, Ruoyu, et al.
Veröffentlicht: (2025)
von: Chen, Ruoyu, et al.
Veröffentlicht: (2025)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
von: Yan, Shilin, et al.
Veröffentlicht: (2026) -
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025) -
Lightweight Frequency Masker for Cross-Domain Few-Shot Semantic Segmentation
von: Tong, Jintao, et al.
Veröffentlicht: (2024) -
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
von: Shen, Haozhan, et al.
Veröffentlicht: (2026) -
Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot Segmentation
von: Tong, Jintao, et al.
Veröffentlicht: (2025)