LangBridge: Interpreting Image as a Combination of Language Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Jiaqi, Niu, Yuwei, Meng, Fanqing, Li, Hao, Tian, Changyao, Du, Yinuo, Xiong, Yuwen, Li, Dianqi, Zhu, Xizhou, Yuan, Li, Dai, Jifeng, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024)
by: Yoon, Dongkeun, et al.
Published: (2024)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
by: Tian, Changyao, et al.
Published: (2023)
by: Tian, Changyao, et al.
Published: (2023)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024)
by: Tian, Changyao, et al.
Published: (2024)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
by: Luo, Gen, et al.
Published: (2025)
by: Luo, Gen, et al.
Published: (2025)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Learning 1D Causal Visual Representation with De-focus Attention Networks
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
by: Tian, Changyao, et al.
Published: (2025)
by: Tian, Changyao, et al.
Published: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
by: Liao, Jiaqi, et al.
Published: (2025)
by: Liao, Jiaqi, et al.
Published: (2025)
big.LITTLE Vision Transformer for Efficient Visual Recognition
by: Guo, He, et al.
Published: (2024)
by: Guo, He, et al.
Published: (2024)
LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
CoMemo: LVLMs Need Image Context with Image Memory
by: Liu, Shi, et al.
Published: (2025)
by: Liu, Shi, et al.
Published: (2025)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
by: Xiong, Yuwen, et al.
Published: (2024)
by: Xiong, Yuwen, et al.
Published: (2024)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
GenExam: A Multidisciplinary Text-to-Image Exam
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Parameter-Inverted Image Pyramid Networks
by: Zhu, Xizhou, et al.
Published: (2024)
by: Zhu, Xizhou, et al.
Published: (2024)
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
Chemical Oxygen Demand: A Key Determinant in Shaping Biological Community Structure.
by: Li, Yao, et al.
Published: (2026)
by: Li, Yao, et al.
Published: (2026)
Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces
by: Galoaa, Bishoy, et al.
Published: (2025)
by: Galoaa, Bishoy, et al.
Published: (2025)
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
by: Jin, Weiyang, et al.
Published: (2025)
by: Jin, Weiyang, et al.
Published: (2025)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
by: Yang, Chenyu, et al.
Published: (2024)
by: Yang, Chenyu, et al.
Published: (2024)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
Structural Analysis of Different Combined‐System Bridges with Cable‐Stayed Bridge without back cables
by: Xianxin Li, et al.
Published: (2025)
by: Xianxin Li, et al.
Published: (2025)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
by: Ge, Junqi, et al.
Published: (2024)
by: Ge, Junqi, et al.
Published: (2024)
MIF‐Mediated NLRP3 Inflammasome‐Dependent Pyroptosis in Spinal Neurons and Microglial Polarization Facilitate Neuropathic Pain Progression
by: Feng Zhou, et al.
Published: (2025)
by: Feng Zhou, et al.
Published: (2025)
DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
Clinical and genetic characteristics of Cornelia de Lange syndrome in pediatric patients
by: Xiaoqiao Li, et al.
Published: (2025)
by: Xiaoqiao Li, et al.
Published: (2025)
Shui Su Ravine Bridge Design and its Steel Joint Fatigue Performance Evaluation
by: Zhihua Xiong, et al.
Published: (2024)
by: Zhihua Xiong, et al.
Published: (2024)
Nodal auxiliary space preconditioning for the surface de Rham complex
by: Li, Yuwen
Published: (2021)
by: Li, Yuwen
Published: (2021)
A new analysis of empirical interpolation methods and Chebyshev greedy algorithms
by: Li, Yuwen
Published: (2024)
by: Li, Yuwen
Published: (2024)
Some p-robust a posteriori error estimates based on auxiliary spaces
by: Li, Yuwen
Published: (2025)
by: Li, Yuwen
Published: (2025)
Sustainable investing in emerging markets: Evidence from the Sustainable Stock Exchanges initiative
by: Yuwen Dai
Published: (2024)
by: Yuwen Dai
Published: (2024)
Useful Memories Become Faulty When Continuously Updated by LLMs
by: Zhang, Dylan, et al.
Published: (2026)
by: Zhang, Dylan, et al.
Published: (2026)
Microbial Biosynthesis of Monoterpenoic Acid from Glycerol
by: Dianqi Yang, et al.
Published: (2026)
by: Dianqi Yang, et al.
Published: (2026)
Microbial Biosynthesis of Natural Esters via Enzyme and Metabolic Engineering
by: Dianqi Yang, et al.
Published: (2025)
by: Dianqi Yang, et al.
Published: (2025)
Similar Items
-
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024) -
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024) -
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
by: Tian, Changyao, et al.
Published: (2023) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024) -
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
by: Li, Hao, et al.
Published: (2024)