SeTformer is What You Need for Vision and Language
Fuente:
arXiv
Salvato in:
| Autori principali: | Shamsolmoali, Pourya, Zareapoor, Masoumeh, Granger, Eric, Felsberg, Michael |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
IntRec: Intent-based Retrieval with Contrastive Refinement
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
From Missing Pieces to Masterpieces: Image Completion with Context-Adaptive Diffusion
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2025)
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2025)
Task Switching Without Forgetting via Proximal Decoupling
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
Fractional Correspondence Framework in Detection Transformer
di: Zareapoor, Masoumeh, et al.
Pubblicazione: (2025)
di: Zareapoor, Masoumeh, et al.
Pubblicazione: (2025)
Multi-Domain Learning with Global Expert Mapping
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
Bidirectional Multi-Step Domain Generalization for Visible-Infrared Person Re-Identification
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2024)
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2024)
Finding Structure in Continual Learning
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026)
TD-Paint: Faster Diffusion Inpainting Through Time Aware Pixel Conditioning
di: Mayet, Tsiry, et al.
Pubblicazione: (2024)
di: Mayet, Tsiry, et al.
Pubblicazione: (2024)
From Cross-Modal to Mixed-Modal Visible-Infrared Re-Identification
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2025)
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2025)
Adaptive Generation of Privileged Intermediate Information for Visible-Infrared Person Re-Identification
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2023)
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2023)
Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2025)
di: Alehdaghi, Mahdi, et al.
Pubblicazione: (2025)
Source-Free Domain Adaptation of Weakly-Supervised Object Localization Models for Histology
di: Guichemerre, Alexis, et al.
Pubblicazione: (2024)
di: Guichemerre, Alexis, et al.
Pubblicazione: (2024)
Adaptation of Weakly Supervised Localization in Histopathology by Debiasing Predictions
di: Guichemerre, Alexis, et al.
Pubblicazione: (2026)
di: Guichemerre, Alexis, et al.
Pubblicazione: (2026)
Smart Feature is What You Need
di: Hu, Zhaoxin, et al.
Pubblicazione: (2024)
di: Hu, Zhaoxin, et al.
Pubblicazione: (2024)
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
di: Belal, Atif, et al.
Pubblicazione: (2025)
di: Belal, Atif, et al.
Pubblicazione: (2025)
Frequency Is What You Need: Considering Word Frequency When Text Masking Benefits Vision-Language Model Pre-training
di: Liang, Mingliang, et al.
Pubblicazione: (2024)
di: Liang, Mingliang, et al.
Pubblicazione: (2024)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
di: Song, Junha, et al.
Pubblicazione: (2026)
di: Song, Junha, et al.
Pubblicazione: (2026)
Test-Time Adaptation via Cache Personalization for Facial Expression Recognition in Videos
di: Sharafi, Masoumeh, et al.
Pubblicazione: (2026)
di: Sharafi, Masoumeh, et al.
Pubblicazione: (2026)
CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition
di: Zeeshan, Muhammad Osama, et al.
Pubblicazione: (2026)
di: Zeeshan, Muhammad Osama, et al.
Pubblicazione: (2026)
Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model
di: Xu, Haoran, et al.
Pubblicazione: (2026)
di: Xu, Haoran, et al.
Pubblicazione: (2026)
Perceptual Inductive Bias Is What You Need Before Contrastive Learning
di: Li, Tianqin, et al.
Pubblicazione: (2025)
di: Li, Tianqin, et al.
Pubblicazione: (2025)
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
di: Bui, Phuoc-Nguyen, et al.
Pubblicazione: (2025)
di: Bui, Phuoc-Nguyen, et al.
Pubblicazione: (2025)
You Only Need Less Attention at Each Stage in Vision Transformers
di: Zhang, Shuoxi, et al.
Pubblicazione: (2024)
di: Zhang, Shuoxi, et al.
Pubblicazione: (2024)
Chameleon: Images Are What You Need For Multimodal Learning Robust To Missing Modalities
di: Liaqat, Muhammad Irzam, et al.
Pubblicazione: (2024)
di: Liaqat, Muhammad Irzam, et al.
Pubblicazione: (2024)
Multi-View Representation is What You Need for Point-Cloud Pre-Training
di: Yan, Siming, et al.
Pubblicazione: (2023)
di: Yan, Siming, et al.
Pubblicazione: (2023)
What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
di: Bahng, Muchang, et al.
Pubblicazione: (2025)
di: Bahng, Muchang, et al.
Pubblicazione: (2025)
DiffSF: Diffusion Models for Scene Flow Estimation
di: Zhang, Yushan, et al.
Pubblicazione: (2024)
di: Zhang, Yushan, et al.
Pubblicazione: (2024)
Affine steerers for structured keypoint description
di: Bökman, Georg, et al.
Pubblicazione: (2024)
di: Bökman, Georg, et al.
Pubblicazione: (2024)
DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
di: Edstedt, Johan, et al.
Pubblicazione: (2025)
di: Edstedt, Johan, et al.
Pubblicazione: (2025)
Steerers: A framework for rotation equivariant keypoint descriptors
di: Bökman, Georg, et al.
Pubblicazione: (2023)
di: Bökman, Georg, et al.
Pubblicazione: (2023)
TetraSphere: A Neural Descriptor for O(3)-Invariant Point Cloud Analysis
di: Melnyk, Pavlo, et al.
Pubblicazione: (2022)
di: Melnyk, Pavlo, et al.
Pubblicazione: (2022)
Text is All You Need for Vision-Language Model Jailbreaking
di: Chen, Yihang, et al.
Pubblicazione: (2026)
di: Chen, Yihang, et al.
Pubblicazione: (2026)
FineVision: Open Data Is All You Need
di: Wiedmann, Luis, et al.
Pubblicazione: (2025)
di: Wiedmann, Luis, et al.
Pubblicazione: (2025)
Disentangled Source-Free Personalization for Facial Expression Recognition with Neutral Target Data
di: Sharafi, Masoumeh, et al.
Pubblicazione: (2025)
di: Sharafi, Masoumeh, et al.
Pubblicazione: (2025)
Lite-SAM Is Actually What You Need for Segment Everything
di: Fu, Jianhai, et al.
Pubblicazione: (2024)
di: Fu, Jianhai, et al.
Pubblicazione: (2024)
Masked Generative Transformer Is What You Need for Image Editing
di: Chow, Wei, et al.
Pubblicazione: (2026)
di: Chow, Wei, et al.
Pubblicazione: (2026)
Take Only What You Need: Rank Minimization as an Implicit Forgetting Regularizer in Continual Learning
di: Lu, Haodong, et al.
Pubblicazione: (2024)
di: Lu, Haodong, et al.
Pubblicazione: (2024)
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
di: Zhang, Boqiang, et al.
Pubblicazione: (2024)
di: Zhang, Boqiang, et al.
Pubblicazione: (2024)
Generating 360° Video is What You Need For a 3D Scene
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2025)
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
IntRec: Intent-based Retrieval with Contrastive Refinement
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026) -
HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026) -
From Missing Pieces to Masterpieces: Image Completion with Context-Adaptive Diffusion
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2025) -
Task Switching Without Forgetting via Proximal Decoupling
di: Shamsolmoali, Pourya, et al.
Pubblicazione: (2026) -
Fractional Correspondence Framework in Detection Transformer
di: Zareapoor, Masoumeh, et al.
Pubblicazione: (2025)