Gespeichert in:
| Hauptverfasser: | Cheng, Jiacheng, Shin, Hijung Valentina, Vasconcelos, Nuno, Russell, Bryan, Heilbron, Fabian Caba |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2405.03190 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EditDuet: A Multi-Agent System for Video Non-Linear Editing
von: Sandoval-Castaneda, Marcelo, et al.
Veröffentlicht: (2025)
von: Sandoval-Castaneda, Marcelo, et al.
Veröffentlicht: (2025)
Discovering Divergent Representations between Text-to-Image Models
von: Dunlap, Lisa, et al.
Veröffentlicht: (2025)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2025)
Improving Personalized Search with Regularized Low-Rank Parameter Updates
von: Ryan, Fiona, et al.
Veröffentlicht: (2025)
von: Ryan, Fiona, et al.
Veröffentlicht: (2025)
ResidualViT for Efficient Temporally Dense Video Encoding
von: Soldan, Mattia, et al.
Veröffentlicht: (2025)
von: Soldan, Mattia, et al.
Veröffentlicht: (2025)
Generative Timelines for Instructed Visual Assembly
von: Pardo, Alejandro, et al.
Veröffentlicht: (2024)
von: Pardo, Alejandro, et al.
Veröffentlicht: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
von: Dave, Ishan Rajendrakumar, et al.
Veröffentlicht: (2024)
von: Dave, Ishan Rajendrakumar, et al.
Veröffentlicht: (2024)
CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
von: Phung, Quynh, et al.
Veröffentlicht: (2025)
von: Phung, Quynh, et al.
Veröffentlicht: (2025)
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
Scaling Up Video Summarization Pretraining with Large Language Models
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
Towards Automated Movie Trailer Generation
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models
von: Kwon, Gihyun, et al.
Veröffentlicht: (2024)
von: Kwon, Gihyun, et al.
Veröffentlicht: (2024)
Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image Synthesis
von: Sridhar, Deepak, et al.
Veröffentlicht: (2024)
von: Sridhar, Deepak, et al.
Veröffentlicht: (2024)
SCHEME: Scalable Channel Mixer for Vision Transformers
von: Sridhar, Deepak, et al.
Veröffentlicht: (2023)
von: Sridhar, Deepak, et al.
Veröffentlicht: (2023)
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
von: Sarkar, Aditya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aditya, et al.
Veröffentlicht: (2026)
Diffusion Models with Adaptive Negative Sampling Without External Resources
von: Desai, Alakh, et al.
Veröffentlicht: (2025)
von: Desai, Alakh, et al.
Veröffentlicht: (2025)
Prompt Sliders for Fine-Grained Control, Editing and Erasing of Concepts in Diffusion Models
von: Sridhar, Deepak, et al.
Veröffentlicht: (2024)
von: Sridhar, Deepak, et al.
Veröffentlicht: (2024)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Improving image synthesis with diffusion-negative sampling
von: Desai, Alakh, et al.
Veröffentlicht: (2024)
von: Desai, Alakh, et al.
Veröffentlicht: (2024)
EditAR: Unified Conditional Generation with Autoregressive Models
von: Mu, Jiteng, et al.
Veröffentlicht: (2025)
von: Mu, Jiteng, et al.
Veröffentlicht: (2025)
Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
PSF-Med: Measuring and Explaining Paraphrase Sensitivity in Medical Vision Language Models
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
Linear Alignment of Vision-language Models for Image Captioning
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
Fairness and Bias Mitigation in Computer Vision: A Survey
von: Dehdashtian, Sepehr, et al.
Veröffentlicht: (2024)
von: Dehdashtian, Sepehr, et al.
Veröffentlicht: (2024)
Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption
von: Adachi, Kazuki, et al.
Veröffentlicht: (2025)
von: Adachi, Kazuki, et al.
Veröffentlicht: (2025)
An Attribute-Based Measure of Video Complexity
von: Sarkar, Aditya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aditya, et al.
Veröffentlicht: (2026)
Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
von: Gao, Zhengqing, et al.
Veröffentlicht: (2024)
von: Gao, Zhengqing, et al.
Veröffentlicht: (2024)
Long-Tailed Anomaly Detection with Learnable Class Names
von: Ho, Chih-Hui, et al.
Veröffentlicht: (2024)
von: Ho, Chih-Hui, et al.
Veröffentlicht: (2024)
Diffusion-based Data Augmentation for Object Counting Problems
von: Wang, Zhen, et al.
Veröffentlicht: (2024)
von: Wang, Zhen, et al.
Veröffentlicht: (2024)
PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
von: Mahadev, Rohan, et al.
Veröffentlicht: (2026)
von: Mahadev, Rohan, et al.
Veröffentlicht: (2026)
ProTeCt: Prompt Tuning for Taxonomic Open Set Classification
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2023)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2023)
EgoPrivacy: What Your First-Person Camera Says About You?
von: Li, Yijiang, et al.
Veröffentlicht: (2025)
von: Li, Yijiang, et al.
Veröffentlicht: (2025)
AdaptSplat: Adapting Vision Foundation Models for Feed-Forward 3D Gaussian Splatting
von: Xing, Mingwei, et al.
Veröffentlicht: (2026)
von: Xing, Mingwei, et al.
Veröffentlicht: (2026)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025)
von: Won, John, et al.
Veröffentlicht: (2025)
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
von: Lou, Meng, et al.
Veröffentlicht: (2026)
von: Lou, Meng, et al.
Veröffentlicht: (2026)
IntroStyle: Training-Free Introspective Style Attribution using Diffusion Features
von: Kumar, Anand, et al.
Veröffentlicht: (2024)
von: Kumar, Anand, et al.
Veröffentlicht: (2024)
Vision encoders should be image size agnostic and task driven
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2025)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2025)
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
von: Khan, Sohail Ahmed, et al.
Veröffentlicht: (2024)
von: Khan, Sohail Ahmed, et al.
Veröffentlicht: (2024)
Anomaly Detection by Adapting a pre-trained Vision Language Model
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation
von: Zhang, Xiaoran, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EditDuet: A Multi-Agent System for Video Non-Linear Editing
von: Sandoval-Castaneda, Marcelo, et al.
Veröffentlicht: (2025) -
Discovering Divergent Representations between Text-to-Image Models
von: Dunlap, Lisa, et al.
Veröffentlicht: (2025) -
Improving Personalized Search with Regularized Low-Rank Parameter Updates
von: Ryan, Fiona, et al.
Veröffentlicht: (2025) -
ResidualViT for Efficient Temporally Dense Video Encoding
von: Soldan, Mattia, et al.
Veröffentlicht: (2025) -
Generative Timelines for Instructed Visual Assembly
von: Pardo, Alejandro, et al.
Veröffentlicht: (2024)