Robust Latent Representation Tuning for Image-text Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Hao, Song, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
Feature CAM: Interpretable AI in Image Classification
von: Clement, Frincy, et al.
Veröffentlicht: (2024)
von: Clement, Frincy, et al.
Veröffentlicht: (2024)
Bernini: Latent Semantic Planning for Video Diffusion
von: Bernini Team, et al.
Veröffentlicht: (2026)
von: Bernini Team, et al.
Veröffentlicht: (2026)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
von: Ouyang, Shuyi, et al.
Veröffentlicht: (2024)
von: Ouyang, Shuyi, et al.
Veröffentlicht: (2024)
Tiny Inference-Time Scaling with Latent Verifiers
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2026)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2026)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Robust Fuzzy Multi-view Learning under View Conflict
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
von: Cao, Pu, et al.
Veröffentlicht: (2023)
von: Cao, Pu, et al.
Veröffentlicht: (2023)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
von: Phung, Thu Hang, et al.
Veröffentlicht: (2026)
von: Phung, Thu Hang, et al.
Veröffentlicht: (2026)
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
Harnessing Self-Supervised Features for Art Classification
von: Melis, Federico, et al.
Veröffentlicht: (2026)
von: Melis, Federico, et al.
Veröffentlicht: (2026)
CrisisViT: A Robust Vision Transformer for Crisis Image Classification
von: Long, Zijun, et al.
Veröffentlicht: (2024)
von: Long, Zijun, et al.
Veröffentlicht: (2024)
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
AI-based System for Transforming text and sound to Educational Videos
von: ElAlami, M. E., et al.
Veröffentlicht: (2026)
von: ElAlami, M. E., et al.
Veröffentlicht: (2026)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
von: Ku, Max, et al.
Veröffentlicht: (2024)
von: Ku, Max, et al.
Veröffentlicht: (2024)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
von: Wu, Yi, et al.
Veröffentlicht: (2025)
von: Wu, Yi, et al.
Veröffentlicht: (2025)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
von: Han, Haochen, et al.
Veröffentlicht: (2024)
von: Han, Haochen, et al.
Veröffentlicht: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
von: Dong, Linfeng, et al.
Veröffentlicht: (2025)
von: Dong, Linfeng, et al.
Veröffentlicht: (2025)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024) -
Feature CAM: Interpretable AI in Image Classification
von: Clement, Frincy, et al.
Veröffentlicht: (2024) -
Bernini: Latent Semantic Planning for Video Diffusion
von: Bernini Team, et al.
Veröffentlicht: (2026) -
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023) -
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)