Guardado en:
| Autores principales: | Park, Seojeong, Choi, Jiho, Baek, Kyungjune, Shim, Hyunjung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2412.20816 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
por: Lee, Seonho, et al.
Publicado: (2025)
por: Lee, Seonho, et al.
Publicado: (2025)
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
por: Park, Junsung, et al.
Publicado: (2024)
por: Park, Junsung, et al.
Publicado: (2024)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
por: Kang, Inha, et al.
Publicado: (2025)
por: Kang, Inha, et al.
Publicado: (2025)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
por: Zhang, Shihang, et al.
Publicado: (2026)
por: Zhang, Shihang, et al.
Publicado: (2026)
Rethinking Direct Preference Optimization in Diffusion Models
por: Kang, Junyong, et al.
Publicado: (2025)
por: Kang, Junyong, et al.
Publicado: (2025)
Precision matters: Precision-aware ensemble for weakly supervised semantic segmentation
por: Park, Junsung, et al.
Publicado: (2024)
por: Park, Junsung, et al.
Publicado: (2024)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
por: Lim, Youngsun, et al.
Publicado: (2024)
por: Lim, Youngsun, et al.
Publicado: (2024)
Grounding Driving VLA via Inverse Kinematics
por: Park, Junsung, et al.
Publicado: (2026)
por: Park, Junsung, et al.
Publicado: (2026)
Label-Augmented Dataset Distillation
por: Kang, Seoungyoon, et al.
Publicado: (2024)
por: Kang, Seoungyoon, et al.
Publicado: (2024)
DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation
por: Kim, Jiwook, et al.
Publicado: (2024)
por: Kim, Jiwook, et al.
Publicado: (2024)
No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather
por: Park, Junsung, et al.
Publicado: (2025)
por: Park, Junsung, et al.
Publicado: (2025)
Sampling Bag of Views for Open-Vocabulary Object Detection
por: Choi, Hojun, et al.
Publicado: (2024)
por: Choi, Hojun, et al.
Publicado: (2024)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
por: Lim, Youngsun, et al.
Publicado: (2024)
por: Lim, Youngsun, et al.
Publicado: (2024)
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
por: Yu, Seungjun, et al.
Publicado: (2025)
por: Yu, Seungjun, et al.
Publicado: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
por: Cao, Zhuo, et al.
Publicado: (2025)
por: Cao, Zhuo, et al.
Publicado: (2025)
PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
por: Choi, Jiho, et al.
Publicado: (2025)
por: Choi, Jiho, et al.
Publicado: (2025)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
por: Kim, Dongseob, et al.
Publicado: (2025)
por: Kim, Dongseob, et al.
Publicado: (2025)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
por: Kim, Ji-Hyeon, et al.
Publicado: (2026)
por: Kim, Ji-Hyeon, et al.
Publicado: (2026)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
por: Yu, An, et al.
Publicado: (2025)
por: Yu, An, et al.
Publicado: (2025)
Memory-Efficient Fine-Tuning for Quantized Diffusion Model
por: Ryu, Hyogon, et al.
Publicado: (2024)
por: Ryu, Hyogon, et al.
Publicado: (2024)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
por: Thanh, Toan Le Ngo, et al.
Publicado: (2025)
por: Thanh, Toan Le Ngo, et al.
Publicado: (2025)
Multi-dimensional Preference Alignment by Conditioning Reward Itself
por: Jang, Jiho, et al.
Publicado: (2025)
por: Jang, Jiho, et al.
Publicado: (2025)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
por: Kwon, Soonwoo, et al.
Publicado: (2025)
por: Kwon, Soonwoo, et al.
Publicado: (2025)
Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict
por: Wu, Chaochen, et al.
Publicado: (2025)
por: Wu, Chaochen, et al.
Publicado: (2025)
SLVideo: A Sign Language Video Moment Retrieval Framework
por: Martins, Gonçalo Vinagre, et al.
Publicado: (2024)
por: Martins, Gonçalo Vinagre, et al.
Publicado: (2024)
Saliency-Guided DETR for Moment Retrieval and Highlight Detection
por: Gordeev, Aleksandr, et al.
Publicado: (2024)
por: Gordeev, Aleksandr, et al.
Publicado: (2024)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
por: An, Na Min, et al.
Publicado: (2025)
por: An, Na Min, et al.
Publicado: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
por: Liu, Xiaolin, et al.
Publicado: (2026)
por: Liu, Xiaolin, et al.
Publicado: (2026)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
por: Park, Jiho, et al.
Publicado: (2025)
por: Park, Jiho, et al.
Publicado: (2025)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
por: Chen, Houlun, et al.
Publicado: (2024)
por: Chen, Houlun, et al.
Publicado: (2024)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
por: Sun, Yunzhuo, et al.
Publicado: (2024)
por: Sun, Yunzhuo, et al.
Publicado: (2024)
Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
por: Choi, Jiho, et al.
Publicado: (2024)
por: Choi, Jiho, et al.
Publicado: (2024)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
por: Lee, Seonho, et al.
Publicado: (2024)
por: Lee, Seonho, et al.
Publicado: (2024)
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
por: Choi, Jiho, et al.
Publicado: (2025)
por: Choi, Jiho, et al.
Publicado: (2025)
I0T: Embedding Standardization Method Towards Zero Modality Gap
por: An, Na Min, et al.
Publicado: (2024)
por: An, Na Min, et al.
Publicado: (2024)
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
por: Yang, Jin, et al.
Publicado: (2024)
por: Yang, Jin, et al.
Publicado: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
por: Um, Sung Jin, et al.
Publicado: (2025)
por: Um, Sung Jin, et al.
Publicado: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
por: Yu, Seungjun, et al.
Publicado: (2025)
por: Yu, Seungjun, et al.
Publicado: (2025)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)
por: Tu, Yunbin, et al.
Publicado: (2024)
Ejemplares similares
-
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
por: Lee, Seonho, et al.
Publicado: (2025) -
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
por: Park, Junsung, et al.
Publicado: (2024) -
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
por: Kang, Inha, et al.
Publicado: (2025) -
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
por: Zhang, Shihang, et al.
Publicado: (2026) -
Rethinking Direct Preference Optimization in Diffusion Models
por: Kang, Junyong, et al.
Publicado: (2025)