Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Vatsal, Gwilliam, Matthew, Kohavi, Gefen, Verma, Eshan, Ulbricht, Daniel, Shrivastava, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026)
by: Agarwal, Vatsal, et al.
Published: (2026)
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)
by: Jayatilaka, Gihan, et al.
Published: (2025)
Accelerate High-Quality Diffusion Models with Inner Loop Feedback
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
Revealing the Utilized Rank of Subspaces of Learning in Neural Networks
by: Garg, Isha, et al.
Published: (2024)
by: Garg, Isha, et al.
Published: (2024)
Characterizing Motion Encoding in Video Diffusion Timesteps
by: Baherwani, Vatsal, et al.
Published: (2025)
by: Baherwani, Vatsal, et al.
Published: (2025)
NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields
by: Zhu, Eric, et al.
Published: (2024)
by: Zhu, Eric, et al.
Published: (2024)
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
by: Maiya, Shishira R, et al.
Published: (2024)
by: Maiya, Shishira R, et al.
Published: (2024)
Cubify Anything: Scaling Indoor 3D Object Detection
by: Lazarow, Justin, et al.
Published: (2024)
by: Lazarow, Justin, et al.
Published: (2024)
LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation
by: Swaminathan, Archana, et al.
Published: (2024)
by: Swaminathan, Archana, et al.
Published: (2024)
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024)
by: Padmanabhan, Namitha, et al.
Published: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
by: Gwilliam, Matthew, et al.
Published: (2023)
by: Gwilliam, Matthew, et al.
Published: (2023)
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)
by: Walmer, Matthew, et al.
Published: (2026)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation
by: Oorloff, Trevine, et al.
Published: (2025)
by: Oorloff, Trevine, et al.
Published: (2025)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
by: Daxberger, Erik, et al.
Published: (2025)
by: Daxberger, Erik, et al.
Published: (2025)
Scale Space Diffusion
by: Mukhopadhyay, Soumik, et al.
Published: (2026)
by: Mukhopadhyay, Soumik, et al.
Published: (2026)
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
by: Ren, Yixuan, et al.
Published: (2025)
by: Ren, Yixuan, et al.
Published: (2025)
MedRAT: Unpaired Medical Report Generation via Auxiliary Tasks
by: Hirsch, Elad, et al.
Published: (2024)
by: Hirsch, Elad, et al.
Published: (2024)
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
by: Saini, Nirat, et al.
Published: (2024)
by: Saini, Nirat, et al.
Published: (2024)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
by: He, Bo, et al.
Published: (2024)
by: He, Bo, et al.
Published: (2024)
Analyzing the Feature Extractor Networks for Face Image Synthesis
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Adaptive Deep Iris Feature Extractor at Arbitrary Resolutions
by: Shoji, Yuho, et al.
Published: (2024)
by: Shoji, Yuho, et al.
Published: (2024)
Adversarial Patch for 3D Local Feature Extractor
by: Pao, Yu Wen, et al.
Published: (2024)
by: Pao, Yu Wen, et al.
Published: (2024)
V-VIPE: Variational View Invariant Pose Embedding
by: Levy, Mara, et al.
Published: (2024)
by: Levy, Mara, et al.
Published: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
DDPM-CD: Denoising Diffusion Probabilistic Models as Feature Extractors for Change Detection
by: Bandara, Wele Gedara Chaminda, et al.
Published: (2022)
by: Bandara, Wele Gedara Chaminda, et al.
Published: (2022)
DiffVQA: Video Quality Assessment Using Diffusion Feature Extractor
by: Chen, Wei-Ting, et al.
Published: (2025)
by: Chen, Wei-Ting, et al.
Published: (2025)
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
by: Goel, Arushi, et al.
Published: (2026)
by: Goel, Arushi, et al.
Published: (2026)
High-resolution efficient image generation from WiFi CSI using a pretrained latent diffusion model
by: Ramesh, Eshan, et al.
Published: (2025)
by: Ramesh, Eshan, et al.
Published: (2025)
An Illumination-Robust Feature Extractor Augmented by Relightable 3D Reconstruction
by: Zhao, Shunyi, et al.
Published: (2024)
by: Zhao, Shunyi, et al.
Published: (2024)
Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference
by: Yuan, Cheng, et al.
Published: (2025)
by: Yuan, Cheng, et al.
Published: (2025)
EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS
by: Girish, Sharath, et al.
Published: (2023)
by: Girish, Sharath, et al.
Published: (2023)
Similar Items
-
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026) -
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025) -
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026) -
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026) -
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)