MLLMs-Augmented Visual-Language Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yanqing, Wang, Kai, Shao, Wenqi, Luo, Ping, Qiao, Yu, Shou, Mike Zheng, Zhang, Kaipeng, You, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
by: Zhao, Henry Hengyuan, et al.
Published: (2024)
by: Zhao, Henry Hengyuan, et al.
Published: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
by: Chen, Zeren, et al.
Published: (2023)
by: Chen, Zeren, et al.
Published: (2023)
Factorized Learning for Temporally Grounded Video-Language Models
by: Zeng, Wenzheng, et al.
Published: (2025)
by: Zeng, Wenzheng, et al.
Published: (2025)
Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
by: Nath, Sujoy, et al.
Published: (2025)
by: Nath, Sujoy, et al.
Published: (2025)
Enhance-A-Video: Better Generated Video for Free
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
by: Wu, Daiqing, et al.
Published: (2025)
by: Wu, Daiqing, et al.
Published: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
by: Zhao, Henry Hengyuan, et al.
Published: (2025)
by: Zhao, Henry Hengyuan, et al.
Published: (2025)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
by: Luo, Gen, et al.
Published: (2024)
by: Luo, Gen, et al.
Published: (2024)
P-Flow: Prompting Visual Effects Generation
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
by: Zhang, Huanyu, et al.
Published: (2025)
by: Zhang, Huanyu, et al.
Published: (2025)
AdaCodec: A Predictive Visual Code for Video MLLMs
by: Hou, Haowen, et al.
Published: (2026)
by: Hou, Haowen, et al.
Published: (2026)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
by: Lin, Yuqi, et al.
Published: (2025)
by: Lin, Yuqi, et al.
Published: (2025)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
by: Zhao, Lirui, et al.
Published: (2024)
by: Zhao, Lirui, et al.
Published: (2024)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
by: Kim, Jungeun, et al.
Published: (2024)
by: Kim, Jungeun, et al.
Published: (2024)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
by: Weng, Fenghua, et al.
Published: (2025)
by: Weng, Fenghua, et al.
Published: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
by: Du, Yuetian, et al.
Published: (2026)
by: Du, Yuetian, et al.
Published: (2026)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
by: Hu, Siyuan, et al.
Published: (2024)
by: Hu, Siyuan, et al.
Published: (2024)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
Visual Program Distillation with Template-Based Augmentation
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Similar Items
-
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
by: Zhang, Hao, et al.
Published: (2024) -
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024) -
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
by: Meng, Fanqing, et al.
Published: (2024) -
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024) -
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024)