MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Sheludzko, Siarhei, Duka, Dhimitrios, Schiele, Bernt, Kuehne, Hilde, Kukleva, Anna |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
di: Ahmed, Noor, et al.
Pubblicazione: (2024)
di: Ahmed, Noor, et al.
Pubblicazione: (2024)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
di: Shvetsova, Nina, et al.
Pubblicazione: (2023)
di: Shvetsova, Nina, et al.
Pubblicazione: (2023)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
di: Shvetsova, Nina, et al.
Pubblicazione: (2025)
di: Shvetsova, Nina, et al.
Pubblicazione: (2025)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
di: Kukleva, Anna, et al.
Pubblicazione: (2024)
di: Kukleva, Anna, et al.
Pubblicazione: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
di: Das, Anurag, et al.
Pubblicazione: (2026)
di: Das, Anurag, et al.
Pubblicazione: (2026)
VideoGEM: Training-free Action Grounding in Videos
di: Vogel, Felix, et al.
Pubblicazione: (2025)
di: Vogel, Felix, et al.
Pubblicazione: (2025)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
di: Kukleva, Anna, et al.
Pubblicazione: (2025)
di: Kukleva, Anna, et al.
Pubblicazione: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
di: Chaybouti, Sofian, et al.
Pubblicazione: (2025)
di: Chaybouti, Sofian, et al.
Pubblicazione: (2025)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
di: Shi, Shaoshuai, et al.
Pubblicazione: (2023)
di: Shi, Shaoshuai, et al.
Pubblicazione: (2023)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
di: Dong, Jiangxin, et al.
Pubblicazione: (2021)
di: Dong, Jiangxin, et al.
Pubblicazione: (2021)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
di: Gorgun, Ada, et al.
Pubblicazione: (2025)
di: Gorgun, Ada, et al.
Pubblicazione: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
di: Jahagirdar, Soumya Shamarao, et al.
Pubblicazione: (2026)
Decoupled Contrastive Learning for Long-Tailed Recognition
di: Xuan, Shiyu, et al.
Pubblicazione: (2024)
di: Xuan, Shiyu, et al.
Pubblicazione: (2024)
SimNP: Learning Self-Similarity Priors Between Neural Points
di: Wewer, Christopher, et al.
Pubblicazione: (2023)
di: Wewer, Christopher, et al.
Pubblicazione: (2023)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
di: Böhle, Moritz, et al.
Pubblicazione: (2021)
di: Böhle, Moritz, et al.
Pubblicazione: (2021)
Towards Better Understanding Attribution Methods
di: Rao, Sukrut, et al.
Pubblicazione: (2022)
di: Rao, Sukrut, et al.
Pubblicazione: (2022)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
di: Nath, Mriganka, et al.
Pubblicazione: (2026)
di: Nath, Mriganka, et al.
Pubblicazione: (2026)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
di: Böhle, Moritz, et al.
Pubblicazione: (2023)
di: Böhle, Moritz, et al.
Pubblicazione: (2023)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
di: Alshami, Eyad, et al.
Pubblicazione: (2025)
di: Alshami, Eyad, et al.
Pubblicazione: (2025)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
di: Gairola, Siddhartha, et al.
Pubblicazione: (2025)
di: Gairola, Siddhartha, et al.
Pubblicazione: (2025)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
di: Das, Anurag, et al.
Pubblicazione: (2024)
di: Das, Anurag, et al.
Pubblicazione: (2024)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
di: Zheng, Junjie, et al.
Pubblicazione: (2025)
di: Zheng, Junjie, et al.
Pubblicazione: (2025)
Long-Tail Learning with Rebalanced Contrastive Loss
di: De Alvis, Charika, et al.
Pubblicazione: (2023)
di: De Alvis, Charika, et al.
Pubblicazione: (2023)
Aligned Contrastive Loss for Long-Tailed Recognition
di: Ma, Jiali, et al.
Pubblicazione: (2025)
di: Ma, Jiali, et al.
Pubblicazione: (2025)
Sp2360: Sparse-view 360 Scene Reconstruction using Cascaded 2D Diffusion Priors
di: Paul, Soumava, et al.
Pubblicazione: (2024)
di: Paul, Soumava, et al.
Pubblicazione: (2024)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
di: Hu, Xinting, et al.
Pubblicazione: (2025)
di: Hu, Xinting, et al.
Pubblicazione: (2025)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
di: Bousselham, Walid, et al.
Pubblicazione: (2026)
di: Bousselham, Walid, et al.
Pubblicazione: (2026)
MICACL: Multi-Instance Category-Aware Contrastive Learning for Long-Tailed Dynamic Facial Expression Recognition
di: Cui, Feng-Qi, et al.
Pubblicazione: (2025)
di: Cui, Feng-Qi, et al.
Pubblicazione: (2025)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
di: Du, Chaoqun, et al.
Pubblicazione: (2024)
di: Du, Chaoqun, et al.
Pubblicazione: (2024)
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
di: Low, Cheng Yaw, et al.
Pubblicazione: (2025)
di: Low, Cheng Yaw, et al.
Pubblicazione: (2025)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
di: Rao, Sukrut, et al.
Pubblicazione: (2023)
di: Rao, Sukrut, et al.
Pubblicazione: (2023)
Adversarial Training against Location-Optimized Adversarial Patches
di: Rao, Sukrut, et al.
Pubblicazione: (2020)
di: Rao, Sukrut, et al.
Pubblicazione: (2020)
Dual Guidance Semi-Supervised Action Detection
di: Singh, Ankit, et al.
Pubblicazione: (2025)
di: Singh, Ankit, et al.
Pubblicazione: (2025)
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
di: Bousselham, Walid, et al.
Pubblicazione: (2024)
di: Bousselham, Walid, et al.
Pubblicazione: (2024)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
di: Shvetsova, Nina, et al.
Pubblicazione: (2025)
di: Shvetsova, Nina, et al.
Pubblicazione: (2025)
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
di: Bousselham, Walid, et al.
Pubblicazione: (2024)
di: Bousselham, Walid, et al.
Pubblicazione: (2024)
Documenti analoghi
-
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
di: Ahmed, Noor, et al.
Pubblicazione: (2024) -
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
di: Shvetsova, Nina, et al.
Pubblicazione: (2023) -
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
di: Jahagirdar, Soumya, et al.
Pubblicazione: (2026) -
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
di: Shvetsova, Nina, et al.
Pubblicazione: (2025) -
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
di: Kukleva, Anna, et al.
Pubblicazione: (2024)