Saved in:
| Main Authors: | Nguyen, Duy-Kien, Assran, Mahmoud, Jain, Unnat, Oswald, Martin R., Snoek, Cees G. M., Chen, Xinlei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.09415 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
by: Nguyen, Duy-Kien, et al.
Published: (2023)
by: Nguyen, Duy-Kien, et al.
Published: (2023)
FVO: Fast Visual Odometry with Transformers
by: Yugay, Vlardimir, et al.
Published: (2025)
by: Yugay, Vlardimir, et al.
Published: (2025)
R-MAE: Regions Meet Masked Autoencoders
by: Nguyen, Duy-Kien, et al.
Published: (2023)
by: Nguyen, Duy-Kien, et al.
Published: (2023)
Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
by: Bao, Yujia, et al.
Published: (2023)
by: Bao, Yujia, et al.
Published: (2023)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
by: Bhowmik, Aritra, et al.
Published: (2023)
by: Bhowmik, Aritra, et al.
Published: (2023)
Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification
by: Avelar, Pedro Henrique da Costa, et al.
Published: (2026)
by: Avelar, Pedro Henrique da Costa, et al.
Published: (2026)
A Pixel Is Worth More Than One 3D Gaussians in Single-View 3D Reconstruction
by: Shen, Jianghao, et al.
Published: (2024)
by: Shen, Jianghao, et al.
Published: (2024)
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
by: Bhowmik, Aritra, et al.
Published: (2024)
by: Bhowmik, Aritra, et al.
Published: (2024)
Low-Resource Vision Challenges for Foundation Models
by: Zhang, Yunhua, et al.
Published: (2024)
by: Zhang, Yunhua, et al.
Published: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
IPO: Interpretable Prompt Optimization for Vision-Language Models
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
Dual Guidance Semi-Supervised Action Detection
by: Singh, Ankit, et al.
Published: (2025)
by: Singh, Ankit, et al.
Published: (2025)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
by: Du, Yingjun, et al.
Published: (2023)
by: Du, Yingjun, et al.
Published: (2023)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion
by: Yan, Xingguang, et al.
Published: (2024)
by: Yan, Xingguang, et al.
Published: (2024)
Segment Any 3D-Part in a Scene from a Sentence
by: Wu, Hongyu, et al.
Published: (2025)
by: Wu, Hongyu, et al.
Published: (2025)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
by: Dorkenwald, Michael, et al.
Published: (2024)
by: Dorkenwald, Michael, et al.
Published: (2024)
Learn to Categorize or Categorize to Learn? Self-Coding for Generalized Category Discovery
by: Rastegar, Sarah, et al.
Published: (2023)
by: Rastegar, Sarah, et al.
Published: (2023)
The Sound of Water: Inferring Physical Properties from Pouring Liquids
by: Bagad, Piyush, et al.
Published: (2024)
by: Bagad, Piyush, et al.
Published: (2024)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
by: Bhowmik, Aritra, et al.
Published: (2025)
by: Bhowmik, Aritra, et al.
Published: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026)
by: Sun, Wenfang, et al.
Published: (2026)
Training-Free Semantic Segmentation via LLM-Supervision
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025)
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025)
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
by: Urbanek, Jack, et al.
Published: (2023)
by: Urbanek, Jack, et al.
Published: (2023)
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
by: Lin, Leo, et al.
Published: (2026)
by: Lin, Leo, et al.
Published: (2026)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
Elastic ViTs from Pretrained Models without Retraining
by: Simoncini, Walter, et al.
Published: (2025)
by: Simoncini, Walter, et al.
Published: (2025)
Any-Shift Prompting for Generalization over Distributions
by: Xiao, Zehao, et al.
Published: (2024)
by: Xiao, Zehao, et al.
Published: (2024)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
by: Cores, Daniel, et al.
Published: (2024)
by: Cores, Daniel, et al.
Published: (2024)
Vision Transformers Need More Than Registers
by: Shi, Cheng, et al.
Published: (2026)
by: Shi, Cheng, et al.
Published: (2026)
MOPA: Modular Object Navigation with PointGoal Agents
by: Raychaudhuri, Sonia, et al.
Published: (2023)
by: Raychaudhuri, Sonia, et al.
Published: (2023)
ViPRA: Video Prediction for Robot Actions
by: Routray, Sandeep, et al.
Published: (2025)
by: Routray, Sandeep, et al.
Published: (2025)
Pixel is a Barrier: Diffusion Models Are More Adversarially Robust Than We Think
by: Xue, Haotian, et al.
Published: (2024)
by: Xue, Haotian, et al.
Published: (2024)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
by: Salehi, Alireza, et al.
Published: (2025)
by: Salehi, Alireza, et al.
Published: (2025)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
by: Rastegar, Sarah, et al.
Published: (2024)
by: Rastegar, Sarah, et al.
Published: (2024)
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Similar Items
-
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
by: Nguyen, Duy-Kien, et al.
Published: (2023) -
FVO: Fast Visual Odometry with Transformers
by: Yugay, Vlardimir, et al.
Published: (2025) -
R-MAE: Regions Meet Masked Autoencoders
by: Nguyen, Duy-Kien, et al.
Published: (2023) -
Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
by: Bao, Yujia, et al.
Published: (2023) -
Union-over-Intersections: Object Detection beyond Winner-Takes-All
by: Bhowmik, Aritra, et al.
Published: (2023)