Generalized Jersey Number Recognition Using Multi-task Learning With Orientation-guided Weight Refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Yung-Hui, Chang, Yu-Wen, Shih, Huang-Chia, Ogawa, Takahiro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression
von: Ke, Jingcheng, et al.
Veröffentlicht: (2024)
von: Ke, Jingcheng, et al.
Veröffentlicht: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Reinforcing Pre-trained Models Using Counterfactual Images
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
G-Refine: A General Quality Refiner for Text-to-Image Generation
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
Multi-task Prompt Words Learning for Social Media Content Generation
von: Xue, Haochen, et al.
Veröffentlicht: (2024)
von: Xue, Haochen, et al.
Veröffentlicht: (2024)
TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views
von: Hung, Hsiang-Hui, et al.
Veröffentlicht: (2025)
von: Hung, Hsiang-Hui, et al.
Veröffentlicht: (2025)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
Joint Flow And Feature Refinement Using Attention For Video Restoration
von: Merugu, Ranjith, et al.
Veröffentlicht: (2025)
von: Merugu, Ranjith, et al.
Veröffentlicht: (2025)
Multi-task Just Recognizable Difference for Video Coding for Machines: Database, Model, and Coding Application
von: Liu, Junqi, et al.
Veröffentlicht: (2026)
von: Liu, Junqi, et al.
Veröffentlicht: (2026)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
SOSControl: Enhancing Human Motion Generation through Saliency-Aware Symbolic Orientation and Timing Control
von: Au, Ho Yin, et al.
Veröffentlicht: (2025)
von: Au, Ho Yin, et al.
Veröffentlicht: (2025)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
GTATrack: Winner Solution to SoccerTrack 2025 with Deep-EIoU and Global Tracklet Association
von: Jian, Rong-Lin, et al.
Veröffentlicht: (2026)
von: Jian, Rong-Lin, et al.
Veröffentlicht: (2026)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
Scalable Image Coding for Humans and Machines Using Feature Fusion Network
von: Shindo, Takahiro, et al.
Veröffentlicht: (2024)
von: Shindo, Takahiro, et al.
Veröffentlicht: (2024)
A Multi-task Adversarial Attack Against Face Authentication
von: Wang, Hanrui, et al.
Veröffentlicht: (2024)
von: Wang, Hanrui, et al.
Veröffentlicht: (2024)
Memory-Guided View Refinement for Dynamic Human-in-the-loop EQA
von: Lu, Xin, et al.
Veröffentlicht: (2026)
von: Lu, Xin, et al.
Veröffentlicht: (2026)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022)
von: Han, Haochen, et al.
Veröffentlicht: (2022)
Optimized Learned Image Compression for Facial Expression Recognition
von: Li, Xiumei, et al.
Veröffentlicht: (2025)
von: Li, Xiumei, et al.
Veröffentlicht: (2025)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
von: Yu, Hang, et al.
Veröffentlicht: (2025)
von: Yu, Hang, et al.
Veröffentlicht: (2025)
Pedestrian Trajectory Prediction Based on Social Interactions Learning With Random Weights
von: Xie, Jiajia, et al.
Veröffentlicht: (2025)
von: Xie, Jiajia, et al.
Veröffentlicht: (2025)
Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
von: Chang, Haochen, et al.
Veröffentlicht: (2024)
von: Chang, Haochen, et al.
Veröffentlicht: (2024)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
von: Li, Shaokai, et al.
Veröffentlicht: (2024)
von: Li, Shaokai, et al.
Veröffentlicht: (2024)
ReCorD: Reasoning and Correcting Diffusion for HOI Generation
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
Two-step Authentication: Multi-biometric System Using Voice and Facial Recognition
von: Chen, Kuan Wei, et al.
Veröffentlicht: (2026)
von: Chen, Kuan Wei, et al.
Veröffentlicht: (2026)
MISS: Memory-efficient Instance Segmentation Framework By Visual Inductive Priors Flow Propagation
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2024)
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2024)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression
von: Ke, Jingcheng, et al.
Veröffentlicht: (2024) -
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026) -
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
von: Gan, Yaozong, et al.
Veröffentlicht: (2024) -
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
von: Gan, Yaozong, et al.
Veröffentlicht: (2024) -
Reinforcing Pre-trained Models Using Counterfactual Images
von: Li, Xiang, et al.
Veröffentlicht: (2024)