Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jung, Ji Hyeok, Kim, Eun Tae, Kim, Seoyeon, Lee, Joo Ho, Kim, Bumsoo, Chang, Buru |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2024)
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2024)
Retrieval Enhanced Feedback via In-context Neural Error-book
von: Hyun, Jongyeop, et al.
Veröffentlicht: (2025)
von: Hyun, Jongyeop, et al.
Veröffentlicht: (2025)
CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model
von: Go, Dongyoung, et al.
Veröffentlicht: (2024)
von: Go, Dongyoung, et al.
Veröffentlicht: (2024)
mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
von: Kim, Kyeong Seon, et al.
Veröffentlicht: (2026)
von: Kim, Kyeong Seon, et al.
Veröffentlicht: (2026)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
von: Park, Yeji, et al.
Veröffentlicht: (2024)
von: Park, Yeji, et al.
Veröffentlicht: (2024)
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
von: Lee, Seohyeong, et al.
Veröffentlicht: (2025)
von: Lee, Seohyeong, et al.
Veröffentlicht: (2025)
Object Aware Egocentric Online Action Detection
von: An, Joungbin, et al.
Veröffentlicht: (2024)
von: An, Joungbin, et al.
Veröffentlicht: (2024)
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing
von: Kim, Euntae, et al.
Veröffentlicht: (2026)
von: Kim, Euntae, et al.
Veröffentlicht: (2026)
SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script
von: Kim, Eunwon, et al.
Veröffentlicht: (2024)
von: Kim, Eunwon, et al.
Veröffentlicht: (2024)
Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
von: Kim, Jieyong, et al.
Veröffentlicht: (2024)
von: Kim, Jieyong, et al.
Veröffentlicht: (2024)
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
von: Kim, Yejin, et al.
Veröffentlicht: (2025)
von: Kim, Yejin, et al.
Veröffentlicht: (2025)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
von: Kim, Taehyeon, et al.
Veröffentlicht: (2023)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2023)
Impact of Leadless Pacemaker Implantation Position on Subclinical Right Ventricular Perforation
von: Young Shin Lee, et al.
Veröffentlicht: (2026)
von: Young Shin Lee, et al.
Veröffentlicht: (2026)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
von: Kim, Jungwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2025)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
von: Kim, Seoyeon, et al.
Veröffentlicht: (2026)
von: Kim, Seoyeon, et al.
Veröffentlicht: (2026)
Evalet: Evaluating Large Language Models through Functional Fragmentation
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
Generative Modeling of Class Probability for Multi-Modal Representation Learning
von: Shin, Jungkyoo, et al.
Veröffentlicht: (2025)
von: Shin, Jungkyoo, et al.
Veröffentlicht: (2025)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
Carbon Dot/Polypyrrole Nanoparticle Complexes as Multifunctional Theranostic Agents
von: TaeEun Kim
Veröffentlicht: (2019)
von: TaeEun Kim
Veröffentlicht: (2019)
Understanding and Tackling Over-Dilution in Graph Neural Networks
von: Lee, Junhyun, et al.
Veröffentlicht: (2025)
von: Lee, Junhyun, et al.
Veröffentlicht: (2025)
Enhancing Psychotherapy Counseling: A Data Augmentation Pipeline Leveraging Large Language Models for Counseling Conversations
von: Kim, Jun-Woo, et al.
Veröffentlicht: (2024)
von: Kim, Jun-Woo, et al.
Veröffentlicht: (2024)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
von: Choi, Tae-Min, et al.
Veröffentlicht: (2025)
von: Choi, Tae-Min, et al.
Veröffentlicht: (2025)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
von: Lee, YuEun, et al.
Veröffentlicht: (2025)
von: Lee, YuEun, et al.
Veröffentlicht: (2025)
The spatial correlation between CN line and dust continuum emitting regions in high-mass star-forming cloud
von: Hwang, Jihye, et al.
Veröffentlicht: (2024)
von: Hwang, Jihye, et al.
Veröffentlicht: (2024)
Enhanced Durability and Catalytic Performance of Pt–SnO2/Multi‐Walled Carbon Nanotube with Shifted d‐Band Center for Proton‐Exchange Membrane Fuel Cells
von: Hyeongwoo Min, et al.
Veröffentlicht: (2024)
von: Hyeongwoo Min, et al.
Veröffentlicht: (2024)
Second Examination of the Right Colon Using Narrow‐Band Imaging Increases Adenoma Detection Rates in the Right Colon: A Multicenter, Randomized Controlled Trial
von: Shin Hee Kim, et al.
Veröffentlicht: (2025)
von: Shin Hee Kim, et al.
Veröffentlicht: (2025)
LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education
von: Lee, Unggi, et al.
Veröffentlicht: (2024)
von: Lee, Unggi, et al.
Veröffentlicht: (2024)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
Bi-directional Contextual Attention for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
Bi-MCQ: Reformulating Vision-Language Alignment for Negation Understanding
von: Kim, Tae Hun, et al.
Veröffentlicht: (2026)
von: Kim, Tae Hun, et al.
Veröffentlicht: (2026)
Comparison of Dexmedetomidine Administration Strategy for Propofol‐Based Pediatric Sedation for Magnetic Resonance Imaging: A Retrospective Study
von: Tae‐Won Kim, et al.
Veröffentlicht: (2026)
von: Tae‐Won Kim, et al.
Veröffentlicht: (2026)
Effect of bowel preparation completion time on bowel cleansing efficacy: Prospective randomized controlled trial of different bowel preparation completion times precolonoscopy
von: Hye Min Kim, et al.
Veröffentlicht: (2024)
von: Hye Min Kim, et al.
Veröffentlicht: (2024)
Dynamics of Oxygen Reserve Index and Arterial Oxygen Partial Pressure in Children: A Prospective Observational Study
von: Jin‐Tae Kim, et al.
Veröffentlicht: (2025)
von: Jin‐Tae Kim, et al.
Veröffentlicht: (2025)
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
von: Hyun, Lee, et al.
Veröffentlicht: (2023)
von: Hyun, Lee, et al.
Veröffentlicht: (2023)
ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
von: Son, Kihoon, et al.
Veröffentlicht: (2025)
von: Son, Kihoon, et al.
Veröffentlicht: (2025)
Immunohistochemical differentiation of keratins and involucrin between palmar psoriasis, chronic hand eczema and hyperkeratotic hand eczema
von: Eun Joo Baek, et al.
Veröffentlicht: (2024)
von: Eun Joo Baek, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025) -
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2024) -
Retrieval Enhanced Feedback via In-context Neural Error-book
von: Hyun, Jongyeop, et al.
Veröffentlicht: (2025) -
CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model
von: Go, Dongyoung, et al.
Veröffentlicht: (2024) -
mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
von: Kim, Kyeong Seon, et al.
Veröffentlicht: (2026)