The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhaoyang, Shao, Run, Wu, Dongyue, Teng, Jiajie, Tao, Chao, Chen, Jingdong, Li, Haifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
von: Shao, Run, et al.
Veröffentlicht: (2024)
von: Shao, Run, et al.
Veröffentlicht: (2024)
Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
von: Yi, Lingjie, et al.
Veröffentlicht: (2025)
von: Yi, Lingjie, et al.
Veröffentlicht: (2025)
The Lattice Representation Hypothesis of Large Language Models
von: Xiong, Bo
Veröffentlicht: (2026)
von: Xiong, Bo
Veröffentlicht: (2026)
AllSpark: A Multimodal Spatio-Temporal General Intelligence Model with Ten Modalities via Language as a Reference Framework
von: Shao, Run, et al.
Veröffentlicht: (2023)
von: Shao, Run, et al.
Veröffentlicht: (2023)
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI
von: Xue, Dong, et al.
Veröffentlicht: (2025)
von: Xue, Dong, et al.
Veröffentlicht: (2025)
Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
von: Li, Haifeng, et al.
Veröffentlicht: (2025)
von: Li, Haifeng, et al.
Veröffentlicht: (2025)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
von: Li, Hengzhuang, et al.
Veröffentlicht: (2025)
von: Li, Hengzhuang, et al.
Veröffentlicht: (2025)
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
von: Hu, Tianyi, et al.
Veröffentlicht: (2026)
von: Hu, Tianyi, et al.
Veröffentlicht: (2026)
Probing Multimodal Large Language Models for Global and Local Semantic Representations
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
Asking like Socrates: Socrates helps VLMs understand remote sensing images
von: Shao, Run, et al.
Veröffentlicht: (2025)
von: Shao, Run, et al.
Veröffentlicht: (2025)
The Linear Representation Hypothesis and the Geometry of Large Language Models
von: Park, Kiho, et al.
Veröffentlicht: (2023)
von: Park, Kiho, et al.
Veröffentlicht: (2023)
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
von: Deng, Zehao, et al.
Veröffentlicht: (2026)
von: Deng, Zehao, et al.
Veröffentlicht: (2026)
BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
A Non-autoregressive Multi-Horizon Flight Trajectory Prediction Framework with Gray Code Representation
von: Guo, Dongyue, et al.
Veröffentlicht: (2023)
von: Guo, Dongyue, et al.
Veröffentlicht: (2023)
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
von: Jin, Ruihan, et al.
Veröffentlicht: (2025)
von: Jin, Ruihan, et al.
Veröffentlicht: (2025)
Dreaming User Multimodal Representation Guided by The Platonic Representation Hypothesis for Micro-Video Recommendation
von: Lin, Chengzhi, et al.
Veröffentlicht: (2024)
von: Lin, Chengzhi, et al.
Veröffentlicht: (2024)
Hierarchical Structure Enhances the Convergence and Generalizability of Linear Molecular Representation
von: Wu, Juan-Ni, et al.
Veröffentlicht: (2024)
von: Wu, Juan-Ni, et al.
Veröffentlicht: (2024)
Hypothesis Generation via LLM-Automated Language Bias for ILP
von: Yang, Yang, et al.
Veröffentlicht: (2025)
von: Yang, Yang, et al.
Veröffentlicht: (2025)
Scaling Law Hypothesis for Multimodal Model
von: Sun, Qingyun, et al.
Veröffentlicht: (2024)
von: Sun, Qingyun, et al.
Veröffentlicht: (2024)
An item is worth one token in Multimodal Large Language Models-based Sequential Recommendation
von: Zhong, Qiyong, et al.
Veröffentlicht: (2025)
von: Zhong, Qiyong, et al.
Veröffentlicht: (2025)
Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment
von: Wang, Haoyuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoyuan, et al.
Veröffentlicht: (2026)
Adapting and Evaluating Multimodal Large Language Models for Adolescent Idiopathic Scoliosis Self-Management: A Divide and Conquer Framework
von: Wu, Zhaolong, et al.
Veröffentlicht: (2025)
von: Wu, Zhaolong, et al.
Veröffentlicht: (2025)
No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
von: Yu, Run, et al.
Veröffentlicht: (2025)
von: Yu, Run, et al.
Veröffentlicht: (2025)
Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis
von: Lin, Ronghao, et al.
Veröffentlicht: (2022)
von: Lin, Ronghao, et al.
Veröffentlicht: (2022)
Generalizable Multimodal Large Language Model Editing via Invariant Trajectory Learning
von: Su, Jiajie, et al.
Veröffentlicht: (2026)
von: Su, Jiajie, et al.
Veröffentlicht: (2026)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
von: Chen, Tao, et al.
Veröffentlicht: (2026)
von: Chen, Tao, et al.
Veröffentlicht: (2026)
AlignMamba-2: Enhancing Multimodal Fusion and Sentiment Analysis with Modality-Aware Mamba
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Predictive Learning in Energy-based Models with Attractor Structures
von: Dong, Xingsi, et al.
Veröffentlicht: (2025)
von: Dong, Xingsi, et al.
Veröffentlicht: (2025)
Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
Mitigating Knowledge Conflicts in Language Model-Driven Question Answering
von: Cao, Han, et al.
Veröffentlicht: (2024)
von: Cao, Han, et al.
Veröffentlicht: (2024)
From Texts to Shields: Convergence of Large Language Models and Cybersecurity
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
Distilling Transitional Pattern to Large Language Models for Multimodal Session-based Recommendation
von: Su, Jiajie, et al.
Veröffentlicht: (2025)
von: Su, Jiajie, et al.
Veröffentlicht: (2025)
Material Property Prediction with Element Attribute Knowledge Graphs and Multimodal Representation Learning
von: Huang, Chao, et al.
Veröffentlicht: (2024)
von: Huang, Chao, et al.
Veröffentlicht: (2024)
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
von: Tao, Xijia, et al.
Veröffentlicht: (2025)
von: Tao, Xijia, et al.
Veröffentlicht: (2025)
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
von: Shao, Zibo, et al.
Veröffentlicht: (2026)
von: Shao, Zibo, et al.
Veröffentlicht: (2026)
MIDG: Mixture of Invariant Experts with knowledge injection for Domain Generalization in Multimodal Sentiment Analysis
von: Li, Yangle, et al.
Veröffentlicht: (2025)
von: Li, Yangle, et al.
Veröffentlicht: (2025)
Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
von: Hu, Xinmiao, et al.
Veröffentlicht: (2025)
von: Hu, Xinmiao, et al.
Veröffentlicht: (2025)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
von: Shao, Run, et al.
Veröffentlicht: (2024) -
Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
von: Yi, Lingjie, et al.
Veröffentlicht: (2025) -
The Lattice Representation Hypothesis of Large Language Models
von: Xiong, Bo
Veröffentlicht: (2026) -
AllSpark: A Multimodal Spatio-Temporal General Intelligence Model with Ten Modalities via Language as a Reference Framework
von: Shao, Run, et al.
Veröffentlicht: (2023) -
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
von: Guo, Mingning, et al.
Veröffentlicht: (2025)