Exploring the Potential of Encoder-free Architectures in 3D LMMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Yiwen, Guo, Zoey, Wang, Zhuhao, Zhang, Ray, Chen, Qizhi, Liu, Junli, Qu, Delin, Wang, Zhigang, Wang, Dong, Zhao, Bin, Li, Xuelong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
di: Tang, Yiwen, et al.
Pubblicazione: (2025)
di: Tang, Yiwen, et al.
Pubblicazione: (2025)
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
di: Tang, Yiwen, et al.
Pubblicazione: (2023)
di: Tang, Yiwen, et al.
Pubblicazione: (2023)
FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives
di: Chen, Qizhi, et al.
Pubblicazione: (2024)
di: Chen, Qizhi, et al.
Pubblicazione: (2024)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
di: Tang, Yiwen, et al.
Pubblicazione: (2024)
di: Tang, Yiwen, et al.
Pubblicazione: (2024)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
di: Liu, Junli, et al.
Pubblicazione: (2025)
di: Liu, Junli, et al.
Pubblicazione: (2025)
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
di: Gao, Xianqiang, et al.
Pubblicazione: (2026)
di: Gao, Xianqiang, et al.
Pubblicazione: (2026)
GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
di: Yan, Chi, et al.
Pubblicazione: (2023)
di: Yan, Chi, et al.
Pubblicazione: (2023)
A Novel Method to Metigate Demographic and Expert Bias in ICD Coding with Causal Inference
di: Zhang, Bin, et al.
Pubblicazione: (2024)
di: Zhang, Bin, et al.
Pubblicazione: (2024)
A Novel ICD Coding Method Based on Associated and Hierarchical Code Description Distillation
di: Zhang, Bin, et al.
Pubblicazione: (2024)
di: Zhang, Bin, et al.
Pubblicazione: (2024)
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
di: Yao, Yuanqi, et al.
Pubblicazione: (2025)
di: Yao, Yuanqi, et al.
Pubblicazione: (2025)
LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control
di: Qu, Delin, et al.
Pubblicazione: (2024)
di: Qu, Delin, et al.
Pubblicazione: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
di: Gao, Xianqiang, et al.
Pubblicazione: (2024)
di: Gao, Xianqiang, et al.
Pubblicazione: (2024)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
di: Jing, Linglin, et al.
Pubblicazione: (2025)
di: Jing, Linglin, et al.
Pubblicazione: (2025)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
di: Zhang, Pingrui, et al.
Pubblicazione: (2025)
di: Zhang, Pingrui, et al.
Pubblicazione: (2025)
Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?
di: Wang, Xiaochen, et al.
Pubblicazione: (2025)
di: Wang, Xiaochen, et al.
Pubblicazione: (2025)
NAG: A Unified Native Architecture for Encoder-free Text-Graph Modeling in Language Models
di: Gong, Haisong, et al.
Pubblicazione: (2026)
di: Gong, Haisong, et al.
Pubblicazione: (2026)
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
di: Ye, Zhoutong, et al.
Pubblicazione: (2025)
di: Ye, Zhoutong, et al.
Pubblicazione: (2025)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
di: Wang, Shun, et al.
Pubblicazione: (2025)
di: Wang, Shun, et al.
Pubblicazione: (2025)
Can MLLMs Generalize to Multi-Party dialog? Exploring Multilingual Response Generation in Complex Scenarios
di: Hu, Zhongtian, et al.
Pubblicazione: (2025)
di: Hu, Zhongtian, et al.
Pubblicazione: (2025)
Implicit Event-RGBD Neural SLAM
di: Qu, Delin, et al.
Pubblicazione: (2023)
di: Qu, Delin, et al.
Pubblicazione: (2023)
LLM-RG4: Flexible and Factual Radiology Report Generation across Diverse Input Contexts
di: Wang, Zhuhao, et al.
Pubblicazione: (2024)
di: Wang, Zhuhao, et al.
Pubblicazione: (2024)
MMSearch-R1: Incentivizing LMMs to Search
di: Wu, Jinming, et al.
Pubblicazione: (2025)
di: Wu, Jinming, et al.
Pubblicazione: (2025)
Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation
di: Guan, Yiwen, et al.
Pubblicazione: (2025)
di: Guan, Yiwen, et al.
Pubblicazione: (2025)
Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR
di: Wang, Minghan, et al.
Pubblicazione: (2024)
di: Wang, Minghan, et al.
Pubblicazione: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs
di: Liang, Jiafeng, et al.
Pubblicazione: (2026)
di: Liang, Jiafeng, et al.
Pubblicazione: (2026)
Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
di: Tian, Xiaoyu, et al.
Pubblicazione: (2025)
di: Tian, Xiaoyu, et al.
Pubblicazione: (2025)
How Well Do LLMs Identify Cultural Unity in Diversity?
di: Li, Jialin, et al.
Pubblicazione: (2024)
di: Li, Jialin, et al.
Pubblicazione: (2024)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
di: Wang, Zixin, et al.
Pubblicazione: (2024)
di: Wang, Zixin, et al.
Pubblicazione: (2024)
All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
di: Vayani, Ashmal, et al.
Pubblicazione: (2024)
di: Vayani, Ashmal, et al.
Pubblicazione: (2024)
Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
di: Zhang, Biao, et al.
Pubblicazione: (2025)
di: Zhang, Biao, et al.
Pubblicazione: (2025)
Open-Vocabulary Octree-Graph for 3D Scene Understanding
di: Wang, Zhigang, et al.
Pubblicazione: (2024)
di: Wang, Zhigang, et al.
Pubblicazione: (2024)
C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
di: Ma, Chengqian, et al.
Pubblicazione: (2025)
di: Ma, Chengqian, et al.
Pubblicazione: (2025)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
di: Huang, Hailang, et al.
Pubblicazione: (2024)
di: Huang, Hailang, et al.
Pubblicazione: (2024)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
di: Lu, Xin, et al.
Pubblicazione: (2025)
di: Lu, Xin, et al.
Pubblicazione: (2025)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
di: Xiong, Shengwu., et al.
Pubblicazione: (2025)
di: Xiong, Shengwu., et al.
Pubblicazione: (2025)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
di: Murthy, Rithesh, et al.
Pubblicazione: (2024)
di: Murthy, Rithesh, et al.
Pubblicazione: (2024)
Languages Transferred Within the Encoder: On Representation Transfer in Zero-Shot Multilingual Translation
di: Qu, Zhi, et al.
Pubblicazione: (2024)
di: Qu, Zhi, et al.
Pubblicazione: (2024)
Night-to-Day Translation via Illumination Degradation Disentanglement
di: Lan, Guanzhou, et al.
Pubblicazione: (2024)
di: Lan, Guanzhou, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
di: Tang, Yiwen, et al.
Pubblicazione: (2025) -
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
di: Tang, Yiwen, et al.
Pubblicazione: (2023) -
FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives
di: Chen, Qizhi, et al.
Pubblicazione: (2024) -
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
di: Tang, Yiwen, et al.
Pubblicazione: (2024) -
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
di: Liu, Junli, et al.
Pubblicazione: (2025)