See It All: Contextualized Late Aggregation for 3D Dense Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Minjung, Lim, Hyung Suk, Kim, Seung Hwan, Lee, Soonyoung, Kim, Bumsoo, Kim, Gunhee |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bi-directional Contextual Attention for 3D Dense Captioning
di: Kim, Minjung, et al.
Pubblicazione: (2024)
di: Kim, Minjung, et al.
Pubblicazione: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
di: Lim, Junyoung, et al.
Pubblicazione: (2025)
di: Lim, Junyoung, et al.
Pubblicazione: (2025)
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
di: Kim, Sangwook, et al.
Pubblicazione: (2025)
di: Kim, Sangwook, et al.
Pubblicazione: (2025)
Can Language Models Laugh at YouTube Short-form Videos?
di: Ko, Dayoon, et al.
Pubblicazione: (2023)
di: Ko, Dayoon, et al.
Pubblicazione: (2023)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
di: Kim, Ye-Chan, et al.
Pubblicazione: (2026)
di: Kim, Ye-Chan, et al.
Pubblicazione: (2026)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
di: Jeon, MinJu, et al.
Pubblicazione: (2025)
di: Jeon, MinJu, et al.
Pubblicazione: (2025)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
di: Kim, Si-Woo, et al.
Pubblicazione: (2025)
di: Kim, Si-Woo, et al.
Pubblicazione: (2025)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
di: Piergiovanni, AJ, et al.
Pubblicazione: (2024)
di: Piergiovanni, AJ, et al.
Pubblicazione: (2024)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
di: Lee, Soeun, et al.
Pubblicazione: (2024)
di: Lee, Soeun, et al.
Pubblicazione: (2024)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
di: Choi, Seung hee, et al.
Pubblicazione: (2026)
di: Choi, Seung hee, et al.
Pubblicazione: (2026)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
CIC: A Framework for Culturally-Aware Image Captioning
di: Yun, Youngsik, et al.
Pubblicazione: (2024)
di: Yun, Youngsik, et al.
Pubblicazione: (2024)
Multi-LLM Collaborative Caption Generation in Scientific Documents
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
di: Kim, Hyunjong, et al.
Pubblicazione: (2025)
di: Kim, Hyunjong, et al.
Pubblicazione: (2025)
Gaussian Blending: Rethinking Alpha Blending in 3D Gaussian Splatting
di: Koo, Junseo, et al.
Pubblicazione: (2025)
di: Koo, Junseo, et al.
Pubblicazione: (2025)
ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition
di: Yoa, Seungdong, et al.
Pubblicazione: (2024)
di: Yoa, Seungdong, et al.
Pubblicazione: (2024)
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
di: Lin, Xiaoyu, et al.
Pubblicazione: (2025)
di: Lin, Xiaoyu, et al.
Pubblicazione: (2025)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
di: Lee, Ji Soo, et al.
Pubblicazione: (2025)
di: Lee, Ji Soo, et al.
Pubblicazione: (2025)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
di: Yun, Heeseung, et al.
Pubblicazione: (2025)
di: Yun, Heeseung, et al.
Pubblicazione: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
di: Ahn, Jaewoo, et al.
Pubblicazione: (2025)
di: Ahn, Jaewoo, et al.
Pubblicazione: (2025)
OmniCaptioner: One Captioner to Rule Them All
di: Lu, Yiting, et al.
Pubblicazione: (2025)
di: Lu, Yiting, et al.
Pubblicazione: (2025)
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
di: Kim, Yunsoo, et al.
Pubblicazione: (2025)
di: Kim, Yunsoo, et al.
Pubblicazione: (2025)
Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision
di: Kim, Jinnyeong, et al.
Pubblicazione: (2024)
di: Kim, Jinnyeong, et al.
Pubblicazione: (2024)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
Diffusion Model for Dense Matching
di: Nam, Jisu, et al.
Pubblicazione: (2023)
di: Nam, Jisu, et al.
Pubblicazione: (2023)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
Decoding fMRI Data into Captions using Prefix Language Modeling
di: Shen, Vyacheslav, et al.
Pubblicazione: (2025)
di: Shen, Vyacheslav, et al.
Pubblicazione: (2025)
See or Guess: Counterfactually Regularized Image Captioning
di: Cao, Qian, et al.
Pubblicazione: (2024)
di: Cao, Qian, et al.
Pubblicazione: (2024)
Time-Scaling State-Space Models for Dense Video Captioning
di: Piergiovanni, AJ, et al.
Pubblicazione: (2025)
di: Piergiovanni, AJ, et al.
Pubblicazione: (2025)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
di: Kim, Seonok
Pubblicazione: (2026)
di: Kim, Seonok
Pubblicazione: (2026)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
di: Hwang, Taebaek, et al.
Pubblicazione: (2025)
di: Hwang, Taebaek, et al.
Pubblicazione: (2025)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
di: Oh, Youngmin, et al.
Pubblicazione: (2024)
di: Oh, Youngmin, et al.
Pubblicazione: (2024)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
di: Kim, Hyung Kyu, et al.
Pubblicazione: (2025)
di: Kim, Hyung Kyu, et al.
Pubblicazione: (2025)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
di: Ahn, Jaewoo, et al.
Pubblicazione: (2025)
di: Ahn, Jaewoo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Bi-directional Contextual Attention for 3D Dense Captioning
di: Kim, Minjung, et al.
Pubblicazione: (2024) -
ChartCap: Mitigating Hallucination of Dense Chart Captioning
di: Lim, Junyoung, et al.
Pubblicazione: (2025) -
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
di: Kim, Sangwook, et al.
Pubblicazione: (2025) -
Can Language Models Laugh at YouTube Short-form Videos?
di: Ko, Dayoon, et al.
Pubblicazione: (2023) -
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
di: Bae, Kyungho, et al.
Pubblicazione: (2025)