Bi-directional Contextual Attention for 3D Dense Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minjung, Lim, Hyung Suk, Lee, Soonyoung, Kim, Bumsoo, Kim, Gunhee |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling
by: Byeon, Junhyeong, et al.
Published: (2026)
by: Byeon, Junhyeong, et al.
Published: (2026)
ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition
by: Yoa, Seungdong, et al.
Published: (2024)
by: Yoa, Seungdong, et al.
Published: (2024)
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
by: Kim, Sangwook, et al.
Published: (2025)
by: Kim, Sangwook, et al.
Published: (2025)
Gaussian Blending: Rethinking Alpha Blending in 3D Gaussian Splatting
by: Koo, Junseo, et al.
Published: (2025)
by: Koo, Junseo, et al.
Published: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
by: Lee, Ji Soo, et al.
Published: (2025)
by: Lee, Ji Soo, et al.
Published: (2025)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
by: Yun, Heeseung, et al.
Published: (2025)
by: Yun, Heeseung, et al.
Published: (2025)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
by: Kim, Ye-Chan, et al.
Published: (2026)
by: Kim, Ye-Chan, et al.
Published: (2026)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
BiTT: Bi-directional Texture Reconstruction of Interacting Two Hands from a Single Image
by: Kim, Minje, et al.
Published: (2024)
by: Kim, Minje, et al.
Published: (2024)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
by: Choi, Seung hee, et al.
Published: (2026)
by: Choi, Seung hee, et al.
Published: (2026)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
Time-Scaling State-Space Models for Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2025)
by: Piergiovanni, AJ, et al.
Published: (2025)
Diffusion Model for Dense Matching
by: Nam, Jisu, et al.
Published: (2023)
by: Nam, Jisu, et al.
Published: (2023)
URECA: Unique Region Caption Anything
by: Lim, Sangbeom, et al.
Published: (2025)
by: Lim, Sangbeom, et al.
Published: (2025)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
OmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable Capabilities
by: Lee, Suyoung, et al.
Published: (2024)
by: Lee, Suyoung, et al.
Published: (2024)
FIMP: Future Interaction Modeling for Multi-Agent Motion Prediction
by: Woo, Sungmin, et al.
Published: (2024)
by: Woo, Sungmin, et al.
Published: (2024)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
by: Song, Seokwon, et al.
Published: (2025)
by: Song, Seokwon, et al.
Published: (2025)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
by: Baik, Sangwon, et al.
Published: (2026)
by: Baik, Sangwon, et al.
Published: (2026)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
by: Kang, Minseok, et al.
Published: (2025)
by: Kang, Minseok, et al.
Published: (2025)
Learning 3D Scene Analogies with Neural Contextual Scene Maps
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
Contextrast: Contextual Contrastive Learning for Semantic Segmentation
by: Sung, Changki, et al.
Published: (2024)
by: Sung, Changki, et al.
Published: (2024)
ConcreTizer: Model Inversion Attack via Occupancy Classification and Dispersion Control for 3D Point Cloud Restoration
by: Kim, Youngseok, et al.
Published: (2025)
by: Kim, Youngseok, et al.
Published: (2025)
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
by: Jung, Ji Hyeok, et al.
Published: (2024)
by: Jung, Ji Hyeok, et al.
Published: (2024)
Integrating Meshes and 3D Gaussians for Indoor Scene Reconstruction with SAM Mask Guidance
by: Kim, Jiyeop, et al.
Published: (2024)
by: Kim, Jiyeop, et al.
Published: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
by: Jin, Bu, et al.
Published: (2024)
by: Jin, Bu, et al.
Published: (2024)
ESR-NeRF: Emissive Source Reconstruction Using LDR Multi-view Images
by: Jeong, Jinseo, et al.
Published: (2024)
by: Jeong, Jinseo, et al.
Published: (2024)
SenseShift6D: Multimodal RGB-D Benchmarking for Robust 6D Pose Estimation across Environment and Sensor Variations
by: Han, Yegyu, et al.
Published: (2025)
by: Han, Yegyu, et al.
Published: (2025)
Similar Items
-
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024) -
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025) -
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025) -
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025) -
Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling
by: Byeon, Junhyeong, et al.
Published: (2026)