NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
Fuente:
arXiv
Guardado en:
| Autores principales: | Tao, Chaofan, Kwon, Gukyeong, Gunjal, Varad, Yang, Hao, Cai, Zhaowei, Dukler, Yonatan, Swaminathan, Ashwin, Manmatha, R., Taylor, Colin Jon, Soatto, Stefano |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mixed-Query Transformer: A Unified Image Segmentation Architecture
por: Wang, Pei, et al.
Publicado: (2024)
por: Wang, Pei, et al.
Publicado: (2024)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
por: Kaul, Prannay, et al.
Publicado: (2024)
por: Kaul, Prannay, et al.
Publicado: (2024)
Training Data Protection with Compositional Diffusion Models
por: Golatkar, Aditya, et al.
Publicado: (2023)
por: Golatkar, Aditya, et al.
Publicado: (2023)
On the Scalability of Diffusion-based Text-to-Image Generation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
por: Zancato, Luca, et al.
Publicado: (2024)
por: Zancato, Luca, et al.
Publicado: (2024)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
por: Li, Xiaolong, et al.
Publicado: (2024)
por: Li, Xiaolong, et al.
Publicado: (2024)
The submodularity of the covolume function in global function fields
por: Bang, Gukyeong
Publicado: (2024)
por: Bang, Gukyeong
Publicado: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
A Quantitative Evaluation of Score Distillation Sampling Based Text-to-3D
por: Fei, Xiaohan, et al.
Publicado: (2024)
por: Fei, Xiaohan, et al.
Publicado: (2024)
CPR: Retrieval Augmented Generation for Copyright Protection
por: Golatkar, Aditya, et al.
Publicado: (2024)
por: Golatkar, Aditya, et al.
Publicado: (2024)
Conjuring Semantic Similarity
por: Liu, Tian Yu, et al.
Publicado: (2024)
por: Liu, Tian Yu, et al.
Publicado: (2024)
Fast Sparse View Guided NeRF Update for Object Reconfigurations
por: Lu, Ziqi, et al.
Publicado: (2024)
por: Lu, Ziqi, et al.
Publicado: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
por: Favero, Alessandro, et al.
Publicado: (2024)
por: Favero, Alessandro, et al.
Publicado: (2024)
Tangent Transformers for Composition, Privacy and Removal
por: Liu, Tian Yu, et al.
Publicado: (2023)
por: Liu, Tian Yu, et al.
Publicado: (2023)
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
por: Dukler, Yonatan, et al.
Publicado: (2025)
por: Dukler, Yonatan, et al.
Publicado: (2025)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
por: Biggs, Benjamin, et al.
Publicado: (2024)
por: Biggs, Benjamin, et al.
Publicado: (2024)
AI Agents as Universal Task Solvers
por: Achille, Alessandro, et al.
Publicado: (2025)
por: Achille, Alessandro, et al.
Publicado: (2025)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
por: Becker, Evan, et al.
Publicado: (2024)
por: Becker, Evan, et al.
Publicado: (2024)
Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion Predictions
por: Zhao, Albert, et al.
Publicado: (2025)
por: Zhao, Albert, et al.
Publicado: (2025)
Singular systems of linear forms over global function fields
por: Bang, Gukyeong, et al.
Publicado: (2024)
por: Bang, Gukyeong, et al.
Publicado: (2024)
Changes in colour and mechanical properties of wood polypropylene composites on natural weathering
por: Jayashri Gunjal
Publicado: (2020)
por: Jayashri Gunjal
Publicado: (2020)
Leveraging Semantic Segmentation Masks with Embeddings for Fine-Grained Form Classification
por: Archibald, Taylor, et al.
Publicado: (2024)
por: Archibald, Taylor, et al.
Publicado: (2024)
MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction
por: Tang, Zitian, et al.
Publicado: (2026)
por: Tang, Zitian, et al.
Publicado: (2026)
FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation
por: Ma, Wenxuan, et al.
Publicado: (2026)
por: Ma, Wenxuan, et al.
Publicado: (2026)
Compositional Structures in Neural Embedding and Interaction Decompositions
por: Trager, Matthew, et al.
Publicado: (2024)
por: Trager, Matthew, et al.
Publicado: (2024)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
por: Zou, Shu, et al.
Publicado: (2025)
por: Zou, Shu, et al.
Publicado: (2025)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
por: Kim, Sungnyun, et al.
Publicado: (2024)
por: Kim, Sungnyun, et al.
Publicado: (2024)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
por: Zhao, Peisen, et al.
Publicado: (2026)
por: Zhao, Peisen, et al.
Publicado: (2026)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
por: Kim, Dahun, et al.
Publicado: (2025)
por: Kim, Dahun, et al.
Publicado: (2025)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
por: Pang, Jinlong, et al.
Publicado: (2025)
por: Pang, Jinlong, et al.
Publicado: (2025)
Critical Learning Periods Emerge Even in Deep Linear Networks
por: Kleinman, Michael, et al.
Publicado: (2023)
por: Kleinman, Michael, et al.
Publicado: (2023)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
por: Ram, Dhananjay, et al.
Publicado: (2025)
por: Ram, Dhananjay, et al.
Publicado: (2025)
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG
por: Hong, Yubin, et al.
Publicado: (2025)
por: Hong, Yubin, et al.
Publicado: (2025)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
por: Trager, Matthew, et al.
Publicado: (2023)
por: Trager, Matthew, et al.
Publicado: (2023)
Analysis of exciton-polariton condensation under different pumping schemes for 1D and 2D microcavities including the effect of strong correlation between polaritons
por: Pande, Varad R.
Publicado: (2025)
por: Pande, Varad R.
Publicado: (2025)
VenusX: Unlocking Fine-Grained Functional Understanding of Proteins
por: Tan, Yang, et al.
Publicado: (2025)
por: Tan, Yang, et al.
Publicado: (2025)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
por: Zhang, Zhaoyang, et al.
Publicado: (2023)
por: Zhang, Zhaoyang, et al.
Publicado: (2023)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
por: Gunjal, Anisha, et al.
Publicado: (2024)
por: Gunjal, Anisha, et al.
Publicado: (2024)
Symmetric Monoidal Bicategories and Biextensions
por: Aldrovandi, Ettore, et al.
Publicado: (2024)
por: Aldrovandi, Ettore, et al.
Publicado: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
por: Liu, Tian Yu, et al.
Publicado: (2025)
por: Liu, Tian Yu, et al.
Publicado: (2025)
Ejemplares similares
-
Mixed-Query Transformer: A Unified Image Segmentation Architecture
por: Wang, Pei, et al.
Publicado: (2024) -
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
por: Kaul, Prannay, et al.
Publicado: (2024) -
Training Data Protection with Compositional Diffusion Models
por: Golatkar, Aditya, et al.
Publicado: (2023) -
On the Scalability of Diffusion-based Text-to-Image Generation
por: Li, Hao, et al.
Publicado: (2024) -
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
por: Zancato, Luca, et al.
Publicado: (2024)