Efficiently Disentangling CLIP for Multi-Object Perception
Fuente:
arXiv
Guardado en:
| Autores principales: | Rawlekar, Samyak, Cai, Yujun, Wang, Yiwei, Yang, Ming-Hsuan, Ahuja, Narendra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
por: Rawlekar, Samyak, et al.
Publicado: (2026)
por: Rawlekar, Samyak, et al.
Publicado: (2026)
Learning Implicit Representation for Reconstructing Articulated Objects
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
por: Rawlekar, Samyak, et al.
Publicado: (2024)
por: Rawlekar, Samyak, et al.
Publicado: (2024)
Improving Multi-label Recognition using Class Co-Occurrence Probabilities
por: Rawlekar, Samyak, et al.
Publicado: (2024)
por: Rawlekar, Samyak, et al.
Publicado: (2024)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen
por: Zhang, Hao, et al.
Publicado: (2023)
por: Zhang, Hao, et al.
Publicado: (2023)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
por: Sun, Bowen, et al.
Publicado: (2025)
por: Sun, Bowen, et al.
Publicado: (2025)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
por: Wu, Hang, et al.
Publicado: (2026)
por: Wu, Hang, et al.
Publicado: (2026)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
por: Wang, Zhaochen, et al.
Publicado: (2025)
por: Wang, Zhaochen, et al.
Publicado: (2025)
SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
por: Wu, Fangyu, et al.
Publicado: (2025)
por: Wu, Fangyu, et al.
Publicado: (2025)
PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
por: Li, Sifan, et al.
Publicado: (2025)
por: Li, Sifan, et al.
Publicado: (2025)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
por: Huang, Kuan-Chih, et al.
Publicado: (2023)
por: Huang, Kuan-Chih, et al.
Publicado: (2023)
Ranking-aware adapter for text-driven image ordering with CLIP
por: Yu, Wei-Hsiang, et al.
Publicado: (2024)
por: Yu, Wei-Hsiang, et al.
Publicado: (2024)
RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes
por: Li, Fang, et al.
Publicado: (2025)
por: Li, Fang, et al.
Publicado: (2025)
Self-Calibrating 4D Novel View Synthesis from Monocular Videos Using Gaussian Splatting
por: Li, Fang, et al.
Publicado: (2024)
por: Li, Fang, et al.
Publicado: (2024)
PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object Detection
por: Huang, Kuan-Chih, et al.
Publicado: (2023)
por: Huang, Kuan-Chih, et al.
Publicado: (2023)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
por: Wang, Zhaochen, et al.
Publicado: (2025)
por: Wang, Zhaochen, et al.
Publicado: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
por: Wu, Yike, et al.
Publicado: (2025)
por: Wu, Yike, et al.
Publicado: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
por: Tao, Xingjian, et al.
Publicado: (2026)
por: Tao, Xingjian, et al.
Publicado: (2026)
HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving
por: Xia, Zhongyu, et al.
Publicado: (2025)
por: Xia, Zhongyu, et al.
Publicado: (2025)
Measuring the (Un)Faithfulness of Concept-Based Explanations
por: Kumar, Shubham, et al.
Publicado: (2025)
por: Kumar, Shubham, et al.
Publicado: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
por: Wu, Hang, et al.
Publicado: (2025)
por: Wu, Hang, et al.
Publicado: (2025)
Learning Disentangled Representation for One-shot Progressive Face Swapping
por: Li, Qi, et al.
Publicado: (2022)
por: Li, Qi, et al.
Publicado: (2022)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
por: Dalal, Dwip, et al.
Publicado: (2025)
por: Dalal, Dwip, et al.
Publicado: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
por: Yang, Kaicheng, et al.
Publicado: (2024)
por: Yang, Kaicheng, et al.
Publicado: (2024)
Physically Aware 360$^\circ$ View Generation from a Single Image using Disentangled Scene Embeddings
por: KV, Karthikeya, et al.
Publicado: (2025)
por: KV, Karthikeya, et al.
Publicado: (2025)
SAMOFT: Robust Multi-Object Tracking via Region and Flow
por: Wang, Yanchao, et al.
Publicado: (2026)
por: Wang, Yanchao, et al.
Publicado: (2026)
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
por: Hu, Xiaodan, et al.
Publicado: (2025)
por: Hu, Xiaodan, et al.
Publicado: (2025)
Incremental Object Detection with CLIP
por: Huang, Ziyue, et al.
Publicado: (2023)
por: Huang, Ziyue, et al.
Publicado: (2023)
Cognitive Disentanglement for Referring Multi-Object Tracking
por: Liang, Shaofeng, et al.
Publicado: (2025)
por: Liang, Shaofeng, et al.
Publicado: (2025)
CAPA: Contribution-Aware Pruning and FFN Approximation for Efficient Large Vision-Language Models
por: Jha, Samyak, et al.
Publicado: (2026)
por: Jha, Samyak, et al.
Publicado: (2026)
HENet: Hybrid Encoding for End-to-end Multi-task 3D Perception from Multi-view Cameras
por: Xia, Zhongyu, et al.
Publicado: (2024)
por: Xia, Zhongyu, et al.
Publicado: (2024)
MagicPose4D: Crafting Articulated Models with Appearance and Motion Control
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
por: Xiao, Linhui, et al.
Publicado: (2023)
por: Xiao, Linhui, et al.
Publicado: (2023)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
por: Chung, Jeannie, et al.
Publicado: (2026)
por: Chung, Jeannie, et al.
Publicado: (2026)
Piecewise-Linear Manifolds for Deep Metric Learning
por: Bhatnagar, Shubhang, et al.
Publicado: (2024)
por: Bhatnagar, Shubhang, et al.
Publicado: (2024)
Ejemplares similares
-
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
por: Rawlekar, Samyak, et al.
Publicado: (2026) -
Learning Implicit Representation for Reconstructing Articulated Objects
por: Zhang, Hao, et al.
Publicado: (2024) -
S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
por: Zhang, Hao, et al.
Publicado: (2024) -
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
por: Rawlekar, Samyak, et al.
Publicado: (2024) -
Improving Multi-label Recognition using Class Co-Occurrence Probabilities
por: Rawlekar, Samyak, et al.
Publicado: (2024)