Saved in:
| Main Authors: | Tan, Reuben, Sun, Ximeng, Hu, Ping, Wang, Jui-hsien, Deilamsalehy, Hanieh, Plummer, Bryan A., Russell, Bryan, Saenko, Kate |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.04346 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLAMP: Contrastive LAnguage Model Prompt-tuning
by: Teterwak, Piotr, et al.
Published: (2023)
by: Teterwak, Piotr, et al.
Published: (2023)
Tell Me What's Next: Textual Foresight for Generic UI Representations
by: Burns, Andrea, et al.
Published: (2024)
by: Burns, Andrea, et al.
Published: (2024)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
by: Qraitem, Maan, et al.
Published: (2023)
by: Qraitem, Maan, et al.
Published: (2023)
SLANT: Spurious Logo ANalysis Toolkit
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
Web Artifact Attacks Disrupt Vision Language Models
by: Qraitem, Maan, et al.
Published: (2025)
by: Qraitem, Maan, et al.
Published: (2025)
OP-LoRA: The Blessing of Dimensionality
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
ERM++: An Improved Baseline for Domain Generalization
by: Teterwak, Piotr, et al.
Published: (2023)
by: Teterwak, Piotr, et al.
Published: (2023)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
by: Liu, Aoming, et al.
Published: (2025)
by: Liu, Aoming, et al.
Published: (2025)
Mull-Tokens: Modality-Agnostic Latent Thinking
by: Ray, Arijit, et al.
Published: (2025)
by: Ray, Arijit, et al.
Published: (2025)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
by: Ray, Arijit, et al.
Published: (2024)
by: Ray, Arijit, et al.
Published: (2024)
LNL+K: Enhancing Learning with Noisy Labels Through Noise Source Knowledge Integration
by: Wang, Siqi, et al.
Published: (2023)
by: Wang, Siqi, et al.
Published: (2023)
RECAST: Reparameterized, Compact weight Adaptation for Sequential Tasks
by: Tasnim, Nazia, et al.
Published: (2024)
by: Tasnim, Nazia, et al.
Published: (2024)
Enhancing Feature Diversity Boosts Channel-Adaptive Vision Transformers
by: Pham, Chau, et al.
Published: (2024)
by: Pham, Chau, et al.
Published: (2024)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025)
by: Mishra, Samarth, et al.
Published: (2025)
Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling
by: Li, Nannan, et al.
Published: (2025)
by: Li, Nannan, et al.
Published: (2025)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
by: Petsiuk, Vitali, et al.
Published: (2024)
by: Petsiuk, Vitali, et al.
Published: (2024)
Noise-Aware Generalization: Robustness to In-Domain Noise and Out-of-Domain Generalization
by: Wang, Siqi, et al.
Published: (2025)
by: Wang, Siqi, et al.
Published: (2025)
FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation
by: Ikechukwu, Nicholas, et al.
Published: (2026)
by: Ikechukwu, Nicholas, et al.
Published: (2026)
NewMove: Customizing text-to-video models with novel motions
by: Materzynska, Joanna, et al.
Published: (2023)
by: Materzynska, Joanna, et al.
Published: (2023)
ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel Learning
by: Pham, Chau, et al.
Published: (2025)
by: Pham, Chau, et al.
Published: (2025)
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
by: Qraitem, Maan, et al.
Published: (2026)
by: Qraitem, Maan, et al.
Published: (2026)
Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning
by: Qin, Wenda, et al.
Published: (2025)
by: Qin, Wenda, et al.
Published: (2025)
Federated Adversarial Domain Adaptation
by: Peng, Xingchao, et al.
Published: (2019)
by: Peng, Xingchao, et al.
Published: (2019)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
by: Miller, Kevin, et al.
Published: (2025)
by: Miller, Kevin, et al.
Published: (2025)
PanoFree: Tuning-Free Holistic Multi-view Image Generation with Cross-view Self-Guidance
by: Liu, Aoming, et al.
Published: (2024)
by: Liu, Aoming, et al.
Published: (2024)
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024)
by: Argaw, Dawit Mureja, et al.
Published: (2024)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
by: Tasnim, Nazia, et al.
Published: (2025)
by: Tasnim, Nazia, et al.
Published: (2025)
Generative Timelines for Instructed Visual Assembly
by: Pardo, Alejandro, et al.
Published: (2024)
by: Pardo, Alejandro, et al.
Published: (2024)
Multi-axis Analysis of Image Manipulation Localization
by: Nichols, Keanu, et al.
Published: (2026)
by: Nichols, Keanu, et al.
Published: (2026)
CHAMMI: A benchmark for channel-adaptive models in microscopy imaging
by: Chen, Zitong, et al.
Published: (2023)
by: Chen, Zitong, et al.
Published: (2023)
UniHuman: A Unified Model for Editing Human Images in the Wild
by: Li, Nannan, et al.
Published: (2023)
by: Li, Nannan, et al.
Published: (2023)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
by: Mishra, Samarth, et al.
Published: (2023)
by: Mishra, Samarth, et al.
Published: (2023)
Rethinking Key-frame-based Micro-expression Recognition: A Robust and Accurate Framework Against Key-frame Errors
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
Bidirectional skip-frame prediction for video anomaly detection with intra-domain disparity-driven attention
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
GIFT: Generated Indoor video frames for Texture-less point tracking
by: Huang, Jianzheng, et al.
Published: (2025)
by: Huang, Jianzheng, et al.
Published: (2025)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
Similar Items
-
CLAMP: Contrastive LAnguage Model Prompt-tuning
by: Teterwak, Piotr, et al.
Published: (2023) -
Tell Me What's Next: Textual Foresight for Generic UI Representations
by: Burns, Andrea, et al.
Published: (2024) -
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
by: Qraitem, Maan, et al.
Published: (2023) -
SLANT: Spurious Logo ANalysis Toolkit
by: Qraitem, Maan, et al.
Published: (2024) -
Web Artifact Attacks Disrupt Vision Language Models
by: Qraitem, Maan, et al.
Published: (2025)