NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
Fuente:
arXiv
Saved in:
| Main Authors: | Tao, Chaofan, Kwon, Gukyeong, Gunjal, Varad, Yang, Hao, Cai, Zhaowei, Dukler, Yonatan, Swaminathan, Ashwin, Manmatha, R., Taylor, Colin Jon, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixed-Query Transformer: A Unified Image Segmentation Architecture
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023)
by: Golatkar, Aditya, et al.
Published: (2023)
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)
by: Zancato, Luca, et al.
Published: (2024)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024)
by: Li, Xiaolong, et al.
Published: (2024)
The submodularity of the covolume function in global function fields
by: Bang, Gukyeong
Published: (2024)
by: Bang, Gukyeong
Published: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
A Quantitative Evaluation of Score Distillation Sampling Based Text-to-3D
by: Fei, Xiaohan, et al.
Published: (2024)
by: Fei, Xiaohan, et al.
Published: (2024)
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)
by: Golatkar, Aditya, et al.
Published: (2024)
Conjuring Semantic Similarity
by: Liu, Tian Yu, et al.
Published: (2024)
by: Liu, Tian Yu, et al.
Published: (2024)
Fast Sparse View Guided NeRF Update for Object Reconfigurations
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Tangent Transformers for Composition, Privacy and Removal
by: Liu, Tian Yu, et al.
Published: (2023)
by: Liu, Tian Yu, et al.
Published: (2023)
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
by: Dukler, Yonatan, et al.
Published: (2025)
by: Dukler, Yonatan, et al.
Published: (2025)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
AI Agents as Universal Task Solvers
by: Achille, Alessandro, et al.
Published: (2025)
by: Achille, Alessandro, et al.
Published: (2025)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
by: Becker, Evan, et al.
Published: (2024)
by: Becker, Evan, et al.
Published: (2024)
Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion Predictions
by: Zhao, Albert, et al.
Published: (2025)
by: Zhao, Albert, et al.
Published: (2025)
Singular systems of linear forms over global function fields
by: Bang, Gukyeong, et al.
Published: (2024)
by: Bang, Gukyeong, et al.
Published: (2024)
Changes in colour and mechanical properties of wood polypropylene composites on natural weathering
by: Jayashri Gunjal
Published: (2020)
by: Jayashri Gunjal
Published: (2020)
Leveraging Semantic Segmentation Masks with Embeddings for Fine-Grained Form Classification
by: Archibald, Taylor, et al.
Published: (2024)
by: Archibald, Taylor, et al.
Published: (2024)
MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction
by: Tang, Zitian, et al.
Published: (2026)
by: Tang, Zitian, et al.
Published: (2026)
FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation
by: Ma, Wenxuan, et al.
Published: (2026)
by: Ma, Wenxuan, et al.
Published: (2026)
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024)
by: Trager, Matthew, et al.
Published: (2024)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
by: Zou, Shu, et al.
Published: (2025)
by: Zou, Shu, et al.
Published: (2025)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
by: Zhao, Peisen, et al.
Published: (2026)
by: Zhao, Peisen, et al.
Published: (2026)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
by: Kim, Dahun, et al.
Published: (2025)
by: Kim, Dahun, et al.
Published: (2025)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Critical Learning Periods Emerge Even in Deep Linear Networks
by: Kleinman, Michael, et al.
Published: (2023)
by: Kleinman, Michael, et al.
Published: (2023)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)
by: Ram, Dhananjay, et al.
Published: (2025)
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG
by: Hong, Yubin, et al.
Published: (2025)
by: Hong, Yubin, et al.
Published: (2025)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
by: Trager, Matthew, et al.
Published: (2023)
by: Trager, Matthew, et al.
Published: (2023)
Analysis of exciton-polariton condensation under different pumping schemes for 1D and 2D microcavities including the effect of strong correlation between polaritons
by: Pande, Varad R.
Published: (2025)
by: Pande, Varad R.
Published: (2025)
VenusX: Unlocking Fine-Grained Functional Understanding of Proteins
by: Tan, Yang, et al.
Published: (2025)
by: Tan, Yang, et al.
Published: (2025)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)
by: Gunjal, Anisha, et al.
Published: (2024)
Symmetric Monoidal Bicategories and Biextensions
by: Aldrovandi, Ettore, et al.
Published: (2024)
by: Aldrovandi, Ettore, et al.
Published: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
by: Liu, Tian Yu, et al.
Published: (2025)
by: Liu, Tian Yu, et al.
Published: (2025)
Similar Items
-
Mixed-Query Transformer: A Unified Image Segmentation Architecture
by: Wang, Pei, et al.
Published: (2024) -
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024) -
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023) -
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024) -
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)