Phantom of Latent for Large Language and Vision Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Byung-Kwan, Chung, Sangyun, Kim, Chae Won, Park, Beomchan, Ro, Yong Man |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TroL: Traversal of Layers for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
MoAI: Mixture of All Intelligence for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
CoLLaVO: Crayon Large Language and Vision mOdel
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
von: Chung, Sangyun, et al.
Veröffentlicht: (2024)
von: Chung, Sangyun, et al.
Veröffentlicht: (2024)
Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
von: Park, Beomchan, et al.
Veröffentlicht: (2026)
von: Park, Beomchan, et al.
Veröffentlicht: (2026)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
Unified Reinforcement and Imitation Learning for Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
Causal Unsupervised Semantic Segmentation
von: Kim, Junho, et al.
Veröffentlicht: (2023)
von: Kim, Junho, et al.
Veröffentlicht: (2023)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
Integrating Language-Derived Appearance Elements with Visual Cues in Pedestrian Detection
von: Park, Sungjune, et al.
Veröffentlicht: (2023)
von: Park, Sungjune, et al.
Veröffentlicht: (2023)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Robust Pedestrian Detection via Constructing Versatile Pedestrian Knowledge Bank
von: Park, Sungjune, et al.
Veröffentlicht: (2024)
von: Park, Sungjune, et al.
Veröffentlicht: (2024)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
von: Park, Sungjune, et al.
Veröffentlicht: (2026)
von: Park, Sungjune, et al.
Veröffentlicht: (2026)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2024)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2024)
Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
von: Kim, Jiwan, et al.
Veröffentlicht: (2026)
von: Kim, Jiwan, et al.
Veröffentlicht: (2026)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language Models
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
von: Lee, Hosu, et al.
Veröffentlicht: (2025)
von: Lee, Hosu, et al.
Veröffentlicht: (2025)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
von: Chung, Dahyun, et al.
Veröffentlicht: (2025)
von: Chung, Dahyun, et al.
Veröffentlicht: (2025)
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
von: Ro, Juneyoung, et al.
Veröffentlicht: (2025)
von: Ro, Juneyoung, et al.
Veröffentlicht: (2025)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning
von: Yun, Seokju, et al.
Veröffentlicht: (2024)
von: Yun, Seokju, et al.
Veröffentlicht: (2024)
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
von: Ko, Donggeun, et al.
Veröffentlicht: (2024)
von: Ko, Donggeun, et al.
Veröffentlicht: (2024)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
On the Adversarial Robustness of 3D Large Vision-Language Models
von: Liu, Chao, et al.
Veröffentlicht: (2026)
von: Liu, Chao, et al.
Veröffentlicht: (2026)
ReConPatch : Contrastive Patch Representation Learning for Industrial Anomaly Detection
von: Hyun, Jeeho, et al.
Veröffentlicht: (2023)
von: Hyun, Jeeho, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
TroL: Traversal of Layers for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024) -
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024) -
MoAI: Mixture of All Intelligence for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024) -
CoLLaVO: Crayon Large Language and Vision mOdel
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024) -
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)