Gespeichert in:
| Hauptverfasser: | Xu, Haiying, Wang, Zihan, Dai, Song, Zhang, Zhengxuan, Dou, Kairan, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.12166 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MGAug: Multimodal Geometric Augmentation in Latent Spaces of Image Deformations
von: Hossain, Tonmoy, et al.
Veröffentlicht: (2023)
von: Hossain, Tonmoy, et al.
Veröffentlicht: (2023)
GMapLatent: Geometric Mapping in Latent Space
von: Zeng, Wei, et al.
Veröffentlicht: (2025)
von: Zeng, Wei, et al.
Veröffentlicht: (2025)
LaRe: Latent Refocusing for Multimodal Reasoning
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
Efficient Implicit Neural Compression of Point Clouds via Learnable Activation in Latent Space
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
The Learnability Gap in Medical Latent Diffusion
von: Dombrowski, Mischa, et al.
Veröffentlicht: (2026)
von: Dombrowski, Mischa, et al.
Veröffentlicht: (2026)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
von: He, Chenwei, et al.
Veröffentlicht: (2026)
von: He, Chenwei, et al.
Veröffentlicht: (2026)
Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
von: Wang, Chenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Chenfeng, et al.
Veröffentlicht: (2026)
LatentUMM: Dual Latent Alignment for Unified Multimodal Models
von: Luo, Yinyi, et al.
Veröffentlicht: (2026)
von: Luo, Yinyi, et al.
Veröffentlicht: (2026)
Semantic-Enriched Latent Visual Reasoning
von: Xu, Tianrun, et al.
Veröffentlicht: (2026)
von: Xu, Tianrun, et al.
Veröffentlicht: (2026)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
von: Wang, Qixun, et al.
Veröffentlicht: (2025)
von: Wang, Qixun, et al.
Veröffentlicht: (2025)
PERL: Parameter Efficient Reasoning in CLIP Latent Space
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
Nodule-Aligned Latent Space Learning with LLM-Driven Multimodal Diffusion for Lung Nodule Progression Prediction
von: Song, James, et al.
Veröffentlicht: (2026)
von: Song, James, et al.
Veröffentlicht: (2026)
ShaLa: Multimodal Shared Latent Space Modelling
von: Cui, Jiali, et al.
Veröffentlicht: (2025)
von: Cui, Jiali, et al.
Veröffentlicht: (2025)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
Rectifying Latent Space for Generative Single-Image Reflection Removal
von: Li, Mingjia, et al.
Veröffentlicht: (2025)
von: Li, Mingjia, et al.
Veröffentlicht: (2025)
Generative Human Motion Stylization in Latent Space
von: Guo, Chuan, et al.
Veröffentlicht: (2024)
von: Guo, Chuan, et al.
Veröffentlicht: (2024)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
von: Ma, Yanbiao, et al.
Veröffentlicht: (2025)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026)
von: Han, Yudong, et al.
Veröffentlicht: (2026)
Latent Watermark: Inject and Detect Watermarks in Latent Diffusion Space
von: Meng, Zheling, et al.
Veröffentlicht: (2024)
von: Meng, Zheling, et al.
Veröffentlicht: (2024)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models
von: Cai, Zhejia, et al.
Veröffentlicht: (2025)
von: Cai, Zhejia, et al.
Veröffentlicht: (2025)
BYOCL: Build Your Own Consistent Latent with Hierarchical Representative Latent Clustering
von: Dai, Jiayue, et al.
Veröffentlicht: (2024)
von: Dai, Jiayue, et al.
Veröffentlicht: (2024)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Improving Deep Representation Learning via Auxiliary Learnable Target Coding
von: Liu, Kangjun, et al.
Veröffentlicht: (2023)
von: Liu, Kangjun, et al.
Veröffentlicht: (2023)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
von: Shaar, Eitan, et al.
Veröffentlicht: (2026)
von: Shaar, Eitan, et al.
Veröffentlicht: (2026)
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026)
von: Liu, Ruiheng, et al.
Veröffentlicht: (2026)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
Constructing Fair Latent Space for Intersection of Fairness and Explainability
von: Joo, Hyungjun, et al.
Veröffentlicht: (2024)
von: Joo, Hyungjun, et al.
Veröffentlicht: (2024)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
von: Mi, Yapeng, et al.
Veröffentlicht: (2025)
von: Mi, Yapeng, et al.
Veröffentlicht: (2025)
Learning Transformation-Isomorphic Latent Space for Accurate Hand Pose Estimation
von: Ren, Kaiwen, et al.
Veröffentlicht: (2025)
von: Ren, Kaiwen, et al.
Veröffentlicht: (2025)
Leveraging Latent Visual Reasoning in Silence
von: Zhu, Dongyao, et al.
Veröffentlicht: (2026)
von: Zhu, Dongyao, et al.
Veröffentlicht: (2026)
Conditional Latent Coding with Learnable Synthesized Reference for Deep Image Compression
von: Wu, Siqi, et al.
Veröffentlicht: (2025)
von: Wu, Siqi, et al.
Veröffentlicht: (2025)
Learning Multimodal Latent Space with EBM Prior and MCMC Inference
von: Yuan, Shiyu, et al.
Veröffentlicht: (2024)
von: Yuan, Shiyu, et al.
Veröffentlicht: (2024)
Latent Diffusion Inversion Requires Understanding the Latent Space
von: Rao, Mingxing, et al.
Veröffentlicht: (2025)
von: Rao, Mingxing, et al.
Veröffentlicht: (2025)
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MGAug: Multimodal Geometric Augmentation in Latent Spaces of Image Deformations
von: Hossain, Tonmoy, et al.
Veröffentlicht: (2023) -
GMapLatent: Geometric Mapping in Latent Space
von: Zeng, Wei, et al.
Veröffentlicht: (2025) -
LaRe: Latent Refocusing for Multimodal Reasoning
von: Ma, Jizheng, et al.
Veröffentlicht: (2025) -
Efficient Implicit Neural Compression of Point Clouds via Learnable Activation in Latent Space
von: Zhang, Yichi, et al.
Veröffentlicht: (2025) -
The Learnability Gap in Medical Latent Diffusion
von: Dombrowski, Mischa, et al.
Veröffentlicht: (2026)