Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Insu, Park, Wooje, Jang, Jaeyun, Noh, Minyoung, Shim, Kyuhong, Shim, Byonghyo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Learning Primitive Relations for Compositional Zero-Shot Learning
by: Lee, Insu, et al.
Published: (2025)
by: Lee, Insu, et al.
Published: (2025)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024)
by: Kim, Donghoon, et al.
Published: (2024)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
by: Kim, Donghoon, et al.
Published: (2025)
by: Kim, Donghoon, et al.
Published: (2025)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024)
by: Kim, Byeonggeun, et al.
Published: (2024)
Mitigating Cross-Image Information Leakage in LVLMs for Multi-Image Tasks
by: Park, Yeji, et al.
Published: (2025)
by: Park, Yeji, et al.
Published: (2025)
Role of Sensing and Computer Vision in 6G Wireless Communications
by: Kim, Seungnyun, et al.
Published: (2024)
by: Kim, Seungnyun, et al.
Published: (2024)
SceneAware: Scene-Constrained Pedestrian Trajectory Prediction with LLM-Guided Walkability
by: Bai, Juho, et al.
Published: (2025)
by: Bai, Juho, et al.
Published: (2025)
VOMTC: Vision Objects for Millimeter and Terahertz Communications
by: Kim, Sunwoo, et al.
Published: (2024)
by: Kim, Sunwoo, et al.
Published: (2024)
Video-Oasis: Rethinking Evaluation of Video Understanding
by: Lim, Geuntaek, et al.
Published: (2026)
by: Lim, Geuntaek, et al.
Published: (2026)
Precision matters: Precision-aware ensemble for weakly supervised semantic segmentation
by: Park, Junsung, et al.
Published: (2024)
by: Park, Junsung, et al.
Published: (2024)
Sampling Bag of Views for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2024)
by: Choi, Hojun, et al.
Published: (2024)
Grounding Driving VLA via Inverse Kinematics
by: Park, Junsung, et al.
Published: (2026)
by: Park, Junsung, et al.
Published: (2026)
No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
by: Park, Junsung, et al.
Published: (2024)
by: Park, Junsung, et al.
Published: (2024)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
by: Kim, Dongseob, et al.
Published: (2025)
by: Kim, Dongseob, et al.
Published: (2025)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
by: Lim, Youngsun, et al.
Published: (2024)
by: Lim, Youngsun, et al.
Published: (2024)
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
by: Ko, Donggeun, et al.
Published: (2024)
by: Ko, Donggeun, et al.
Published: (2024)
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
by: Grauman, Kristen, et al.
Published: (2023)
by: Grauman, Kristen, et al.
Published: (2023)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
by: Park, NaHyeon, et al.
Published: (2024)
by: Park, NaHyeon, et al.
Published: (2024)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
by: Park, Seojeong, et al.
Published: (2024)
by: Park, Seojeong, et al.
Published: (2024)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
by: Kim, Myunsoo, et al.
Published: (2025)
by: Kim, Myunsoo, et al.
Published: (2025)
Weakly Supervised Semantic Segmentation for Driving Scenes
by: Kim, Dongseob, et al.
Published: (2023)
by: Kim, Dongseob, et al.
Published: (2023)
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
Towards Visual Text Design Transfer Across Languages
by: Choi, Yejin, et al.
Published: (2024)
by: Choi, Yejin, et al.
Published: (2024)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
by: Zhao, Ziwei, et al.
Published: (2025)
by: Zhao, Ziwei, et al.
Published: (2025)
Preoperative Rotator Cuff Tear Prediction from Shoulder Radiographs using a Convolutional Block Attention Module-Integrated Neural Network
by: Jo, Chris Hyunchul, et al.
Published: (2024)
by: Jo, Chris Hyunchul, et al.
Published: (2024)
DITTO: Dual and Integrated Latent Topologies for Implicit 3D Reconstruction
by: Shim, Jaehyeok, et al.
Published: (2024)
by: Shim, Jaehyeok, et al.
Published: (2024)
Label-Augmented Dataset Distillation
by: Kang, Seoungyoon, et al.
Published: (2024)
by: Kang, Seoungyoon, et al.
Published: (2024)
Memory-Efficient Fine-Tuning for Quantized Diffusion Model
by: Ryu, Hyogon, et al.
Published: (2024)
by: Ryu, Hyogon, et al.
Published: (2024)
Knowledge-based learning in Text-RAG and Image-RAG
by: Shim, Alexander, et al.
Published: (2026)
by: Shim, Alexander, et al.
Published: (2026)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
by: Lim, Youngsun, et al.
Published: (2024)
by: Lim, Youngsun, et al.
Published: (2024)
SNP: Structured Neuron-level Pruning to Preserve Attention Scores
by: Shim, Kyunghwan, et al.
Published: (2024)
by: Shim, Kyunghwan, et al.
Published: (2024)
CrimEdit: Controllable Editing for Counterfactual Object Removal, Insertion, and Movement
by: Jeon, Boseong, et al.
Published: (2025)
by: Jeon, Boseong, et al.
Published: (2025)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Similar Items
-
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026) -
Learning Primitive Relations for Compositional Zero-Shot Learning
by: Lee, Insu, et al.
Published: (2025) -
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024) -
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
by: Kim, Donghoon, et al.
Published: (2025) -
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024)