LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Linquan, Jiang, Tianxiang, Dong, Yifei, Yang, Haoyu, Zhang, Fengji, Meng, Shichaang, Xuan, Ai, Song, Linqi, Keung, Jacky |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
by: Zhang, Fengji, et al.
Published: (2024)
by: Zhang, Fengji, et al.
Published: (2024)
FaceSleuth-R: Adaptive Orientation-Aware Attention for Robust Micro-Expression Recognition
by: Wu, Linquan, et al.
Published: (2025)
by: Wu, Linquan, et al.
Published: (2025)
Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
by: Mao, Zhenyu, et al.
Published: (2025)
by: Mao, Zhenyu, et al.
Published: (2025)
SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics
by: Cao, Yuchen, et al.
Published: (2026)
by: Cao, Yuchen, et al.
Published: (2026)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
Multi-Strategy Enhanced COA for Path Planning in Autonomous Navigation
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
Data Preparation for Deep Learning based Code Smell Detection: A Systematic Literature Review
by: Zhang, Fengji, et al.
Published: (2024)
by: Zhang, Fengji, et al.
Published: (2024)
Latent Chain-of-Thought for Visual Reasoning
by: Sun, Guohao, et al.
Published: (2025)
by: Sun, Guohao, et al.
Published: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
by: Zhang, Huanyu, et al.
Published: (2025)
by: Zhang, Huanyu, et al.
Published: (2025)
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Tree Learning: A Multi-Skill Continual Learning Framework for Humanoid Robots
by: Yan, Yifei, et al.
Published: (2026)
by: Yan, Yifei, et al.
Published: (2026)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
by: Jiang, Tianxiang, et al.
Published: (2025)
by: Jiang, Tianxiang, et al.
Published: (2025)
Hybrid Privacy Policy-Code Consistency Check using Knowledge Graphs and LLMs
by: Mao, Zhenyu, et al.
Published: (2025)
by: Mao, Zhenyu, et al.
Published: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
by: Dong, Sibo, et al.
Published: (2025)
by: Dong, Sibo, et al.
Published: (2025)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
by: Jiang, Gangwei, et al.
Published: (2025)
by: Jiang, Gangwei, et al.
Published: (2025)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
by: Wu, Fengyi, et al.
Published: (2025)
by: Wu, Fengyi, et al.
Published: (2025)
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
by: Shao, Yifei, et al.
Published: (2026)
by: Shao, Yifei, et al.
Published: (2026)
R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
by: Lu, Wenting, et al.
Published: (2026)
by: Lu, Wenting, et al.
Published: (2026)
ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Latent Visual Reasoning
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
R2Code: A Self-Reflective LLM Framework for Requirements-to-Code Traceability
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Towards Requirements Engineering for GenAI-Enabled Software: Bridging Responsibility Gaps through Human Oversight Requirements
by: Mao, Zhenyu, et al.
Published: (2025)
by: Mao, Zhenyu, et al.
Published: (2025)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
by: Tong, Haoyu, et al.
Published: (2026)
by: Tong, Haoyu, et al.
Published: (2026)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
A$^2$Search: Ambiguity-Aware Question Answering with Reinforcement Learning
by: Zhang, Fengji, et al.
Published: (2025)
by: Zhang, Fengji, et al.
Published: (2025)
LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning
by: Xu, Zifan, et al.
Published: (2023)
by: Xu, Zifan, et al.
Published: (2023)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
by: Li, Zhaoyi, et al.
Published: (2026)
by: Li, Zhaoyi, et al.
Published: (2026)
Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
by: Chen, Xinghao, et al.
Published: (2025)
by: Chen, Xinghao, et al.
Published: (2025)
Efficient Long-distance Latent Relation-aware Graph Neural Network for Multi-modal Emotion Recognition in Conversations
by: Shou, Yuntao, et al.
Published: (2024)
by: Shou, Yuntao, et al.
Published: (2024)
Infinitely many associated primes of local cohomology modules of ramified regular local rings
by: Ma, Linquan
Published: (2026)
by: Ma, Linquan
Published: (2026)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
LLM Reasoning Is Latent, Not the Chain of Thought
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
Reasoning to Learn from Latent Thoughts
by: Ruan, Yangjun, et al.
Published: (2025)
by: Ruan, Yangjun, et al.
Published: (2025)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
by: He, Hulingxiao, et al.
Published: (2026)
by: He, Hulingxiao, et al.
Published: (2026)
Similar Items
-
HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
by: Zhang, Fengji, et al.
Published: (2024) -
FaceSleuth-R: Adaptive Orientation-Aware Attention for Robust Micro-Expression Recognition
by: Wu, Linquan, et al.
Published: (2025) -
Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
by: Mao, Zhenyu, et al.
Published: (2025) -
SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics
by: Cao, Yuchen, et al.
Published: (2026) -
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)