Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Luo, Yongdong, Lin, Haojia, Zheng, Xiawu, Jiang, Yigeng, Chao, Fei, Hu, Jie, Jiang, Guannan, Zhang, Songan, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
par: Luo, Yongdong, et autres
Publié: (2024)
par: Luo, Yongdong, et autres
Publié: (2024)
Event-Anchored Frame Selection for Effective Long-Video Understanding
par: Chen, Wang, et autres
Publié: (2026)
par: Chen, Wang, et autres
Publié: (2026)
QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
par: Luo, Yongdong, et autres
Publié: (2025)
par: Luo, Yongdong, et autres
Publié: (2025)
CamoTeacher: Dual-Rotation Consistency Learning for Semi-Supervised Camouflaged Object Detection
par: Lai, Xunfa, et autres
Publié: (2024)
par: Lai, Xunfa, et autres
Publié: (2024)
Multi-branch Collaborative Learning Network for 3D Visual Grounding
par: Qian, Zhipeng, et autres
Publié: (2024)
par: Qian, Zhipeng, et autres
Publié: (2024)
Towards Efficient Automatic Self-Pruning of Large Language Models
par: Huang, Weizhong, et autres
Publié: (2025)
par: Huang, Weizhong, et autres
Publié: (2025)
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
par: Huang, Weizhong, et autres
Publié: (2025)
par: Huang, Weizhong, et autres
Publié: (2025)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
par: Wu, Mingrui, et autres
Publié: (2024)
par: Wu, Mingrui, et autres
Publié: (2024)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
par: Zheng, Xuzhe, et autres
Publié: (2026)
par: Zheng, Xuzhe, et autres
Publié: (2026)
OMPQ: Orthogonal Mixed Precision Quantization
par: Ma, Yuexiao, et autres
Publié: (2021)
par: Ma, Yuexiao, et autres
Publié: (2021)
Depth-Guided Semi-Supervised Instance Segmentation
par: Chen, Xin, et autres
Publié: (2024)
par: Chen, Xin, et autres
Publié: (2024)
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
par: Ma, Qingchuan, et autres
Publié: (2025)
par: Ma, Qingchuan, et autres
Publié: (2025)
ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
par: Liao, Kunpeng, et autres
Publié: (2026)
par: Liao, Kunpeng, et autres
Publié: (2026)
Polybasic Speculative Decoding Through a Theoretical Perspective
par: Wang, Ruilin, et autres
Publié: (2025)
par: Wang, Ruilin, et autres
Publié: (2025)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
par: Song, Jiahe, et autres
Publié: (2025)
par: Song, Jiahe, et autres
Publié: (2025)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
par: Luo, Liqin, et autres
Publié: (2025)
par: Luo, Liqin, et autres
Publié: (2025)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
par: Huang, Haoyu, et autres
Publié: (2026)
par: Huang, Haoyu, et autres
Publié: (2026)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
par: Luo, Gen, et autres
Publié: (2024)
par: Luo, Gen, et autres
Publié: (2024)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
par: Guo, Song, et autres
Publié: (2024)
par: Guo, Song, et autres
Publié: (2024)
Uncovering the Over-smoothing Challenge in Image Super-Resolution: Entropy-based Quantification and Contrastive Optimization
par: Xu, Tianshuo, et autres
Publié: (2022)
par: Xu, Tianshuo, et autres
Publié: (2022)
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
par: Ma, Yiwei, et autres
Publié: (2024)
par: Ma, Yiwei, et autres
Publié: (2024)
Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition
par: Huang, Binyuan, et autres
Publié: (2025)
par: Huang, Binyuan, et autres
Publié: (2025)
Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
par: Zhao, Lirui, et autres
Publié: (2023)
par: Zhao, Lirui, et autres
Publié: (2023)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
par: Chen, Wang, et autres
Publié: (2026)
par: Chen, Wang, et autres
Publié: (2026)
AffineQuant: Affine Transformation Quantization for Large Language Models
par: Ma, Yuexiao, et autres
Publié: (2024)
par: Ma, Yuexiao, et autres
Publié: (2024)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
par: Lian, Shuquan, et autres
Publié: (2025)
par: Lian, Shuquan, et autres
Publié: (2025)
Unified-Width Adaptive Dynamic Network for All-In-One Image Restoration
par: Xu, Yimin, et autres
Publié: (2024)
par: Xu, Yimin, et autres
Publié: (2024)
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
par: Xie, Zhuyang, et autres
Publié: (2024)
par: Xie, Zhuyang, et autres
Publié: (2024)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
par: Dai, Chengpeng, et autres
Publié: (2022)
par: Dai, Chengpeng, et autres
Publié: (2022)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
par: Li, Xiangtai, et autres
Publié: (2025)
par: Li, Xiangtai, et autres
Publié: (2025)
A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
par: Ma, Qingchuan, et autres
Publié: (2026)
par: Ma, Qingchuan, et autres
Publié: (2026)
VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
par: Jiang, Pengfei, et autres
Publié: (2025)
par: Jiang, Pengfei, et autres
Publié: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
par: Xu, Jing, et autres
Publié: (2026)
par: Xu, Jing, et autres
Publié: (2026)
ASBench: Image Anomalies Synthesis Benchmark for Anomaly Detection
par: Zhang, Qunyi, et autres
Publié: (2025)
par: Zhang, Qunyi, et autres
Publié: (2025)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
par: Xu, Boyue, et autres
Publié: (2026)
par: Xu, Boyue, et autres
Publié: (2026)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
par: Fei, Xiang, et autres
Publié: (2025)
par: Fei, Xiang, et autres
Publié: (2025)
A Pedestrian-Vehicle Interaction Benchmark and Annotation Framework for Unstructured Scenes via Uncalibrated Cameras
par: Peng, Haoyang, et autres
Publié: (2026)
par: Peng, Haoyang, et autres
Publié: (2026)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
par: Xie, Tianyu, et autres
Publié: (2026)
par: Xie, Tianyu, et autres
Publié: (2026)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
par: Ge, Shiping, et autres
Publié: (2024)
par: Ge, Shiping, et autres
Publié: (2024)
ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference
par: Huang, Zhaohong, et autres
Publié: (2026)
par: Huang, Zhaohong, et autres
Publié: (2026)
Documents similaires
-
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
par: Luo, Yongdong, et autres
Publié: (2024) -
Event-Anchored Frame Selection for Effective Long-Video Understanding
par: Chen, Wang, et autres
Publié: (2026) -
QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
par: Luo, Yongdong, et autres
Publié: (2025) -
CamoTeacher: Dual-Rotation Consistency Learning for Semi-Supervised Camouflaged Object Detection
par: Lai, Xunfa, et autres
Publié: (2024) -
Multi-branch Collaborative Learning Network for 3D Visual Grounding
par: Qian, Zhipeng, et autres
Publié: (2024)