Improving Accuracy and Generalization for Efficient Visual Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Zaveri, Ram, Patel, Shivang, Gu, Yu, Doretto, Gianfranco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-shot adaptation for morphology-independent cell instance segmentation
by: Zaveri, Ram J., et al.
Published: (2024)
by: Zaveri, Ram J., et al.
Published: (2024)
Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling
by: Singh, Jaskirat, et al.
Published: (2025)
by: Singh, Jaskirat, et al.
Published: (2025)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
Improving Generative Adversarial Network Generalization for Facial Expression Synthesis
by: Akram, Arbish, et al.
Published: (2026)
by: Akram, Arbish, et al.
Published: (2026)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
by: Saleh, Mohamed, et al.
Published: (2026)
by: Saleh, Mohamed, et al.
Published: (2026)
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
by: Sun, Shuyang, et al.
Published: (2023)
by: Sun, Shuyang, et al.
Published: (2023)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)
by: Pang, Xingzhou, et al.
Published: (2026)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
by: Novitskiy, Lev, et al.
Published: (2025)
by: Novitskiy, Lev, et al.
Published: (2025)
Generalized Jersey Number Recognition Using Multi-task Learning With Orientation-guided Weight Refinement
by: Lin, Yung-Hui, et al.
Published: (2024)
by: Lin, Yung-Hui, et al.
Published: (2024)
Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
by: Wang, Jinpeng, et al.
Published: (2025)
by: Wang, Jinpeng, et al.
Published: (2025)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
by: Shen, Chengchao, et al.
Published: (2025)
by: Shen, Chengchao, et al.
Published: (2025)
Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring
by: Huang, Youhan, et al.
Published: (2026)
by: Huang, Youhan, et al.
Published: (2026)
GUESS: Generative Uncertainty Ensemble for Self Supervision
by: Mohamadi, Salman, et al.
Published: (2024)
by: Mohamadi, Salman, et al.
Published: (2024)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
by: Jiang, Xun, et al.
Published: (2026)
by: Jiang, Xun, et al.
Published: (2026)
Evaluating the Impact of Point Cloud Colorization on Semantic Segmentation Accuracy
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
by: Panambur, Tejas, et al.
Published: (2025)
by: Panambur, Tejas, et al.
Published: (2025)
MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
by: Zhao, Binyu, et al.
Published: (2025)
by: Zhao, Binyu, et al.
Published: (2025)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024)
by: Wang, Zhecan, et al.
Published: (2024)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
by: Wang, Ziteng, et al.
Published: (2026)
by: Wang, Ziteng, et al.
Published: (2026)
DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation
by: Liu, Ziyuan, et al.
Published: (2026)
by: Liu, Ziyuan, et al.
Published: (2026)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
by: Yu, Lijun
Published: (2024)
by: Yu, Lijun
Published: (2024)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
by: Wu, Qiong, et al.
Published: (2024)
by: Wu, Qiong, et al.
Published: (2024)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
by: Flynn, John, et al.
Published: (2026)
by: Flynn, John, et al.
Published: (2026)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
by: Sánchez, Jorge, et al.
Published: (2024)
by: Sánchez, Jorge, et al.
Published: (2024)
Regularized Contrastive Partial Multi-view Outlier Detection
by: Wang, Yijia, et al.
Published: (2024)
by: Wang, Yijia, et al.
Published: (2024)
Relating CNN-Transformer Fusion Network for Change Detection
by: Gao, Yuhao, et al.
Published: (2024)
by: Gao, Yuhao, et al.
Published: (2024)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
by: Gao, Lishuai, et al.
Published: (2024)
by: Gao, Lishuai, et al.
Published: (2024)
Competitive Learning for Achieving Content-specific Filters in Video Coding for Machines
by: Zhang, Honglei, et al.
Published: (2024)
by: Zhang, Honglei, et al.
Published: (2024)
MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition
by: Xiang, Peihao, et al.
Published: (2024)
by: Xiang, Peihao, et al.
Published: (2024)
Adversarially Robust Deepfake Detection via Adversarial Feature Similarity Learning
by: Khan, Sarwar
Published: (2024)
by: Khan, Sarwar
Published: (2024)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
CinePile: A Long Video Question Answering Dataset and Benchmark
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
BRep Boundary and Junction Detection for CAD Reverse Engineering
by: Ali, Sk Aziz, et al.
Published: (2024)
by: Ali, Sk Aziz, et al.
Published: (2024)
360VFI: A Dataset and Benchmark for Omnidirectional Video Frame Interpolation
by: Lu, Wenxuan, et al.
Published: (2024)
by: Lu, Wenxuan, et al.
Published: (2024)
Multimodal Transformer With a Low-Computational-Cost Guarantee
by: Park, Sungjin, et al.
Published: (2024)
by: Park, Sungjin, et al.
Published: (2024)
Similar Items
-
Few-shot adaptation for morphology-independent cell instance segmentation
by: Zaveri, Ram J., et al.
Published: (2024) -
Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling
by: Singh, Jaskirat, et al.
Published: (2025) -
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026) -
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025) -
Improving Generative Adversarial Network Generalization for Facial Expression Synthesis
by: Akram, Arbish, et al.
Published: (2026)