Do text-free diffusion models learn discriminative visual representations?
Fuente:
arXiv
Saved in:
| Main Authors: | Mukhopadhyay, Soumik, Gwilliam, Matthew, Yamaguchi, Yosuke, Agarwal, Vatsal, Padmanabhan, Namitha, Swaminathan, Archana, Zhou, Tianyi, Ohya, Jun, Shrivastava, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026)
by: Agarwal, Vatsal, et al.
Published: (2026)
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024)
by: Padmanabhan, Namitha, et al.
Published: (2024)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)
by: Agarwal, Vatsal, et al.
Published: (2025)
LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation
by: Swaminathan, Archana, et al.
Published: (2024)
by: Swaminathan, Archana, et al.
Published: (2024)
Scale Space Diffusion
by: Mukhopadhyay, Soumik, et al.
Published: (2026)
by: Mukhopadhyay, Soumik, et al.
Published: (2026)
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)
by: Jayatilaka, Gihan, et al.
Published: (2025)
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024)
by: Kumar, Pulkit, et al.
Published: (2024)
NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields
by: Zhu, Eric, et al.
Published: (2024)
by: Zhu, Eric, et al.
Published: (2024)
Characterizing Motion Encoding in Video Diffusion Timesteps
by: Baherwani, Vatsal, et al.
Published: (2025)
by: Baherwani, Vatsal, et al.
Published: (2025)
Accelerate High-Quality Diffusion Models with Inner Loop Feedback
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
by: Maiya, Shishira R, et al.
Published: (2024)
by: Maiya, Shishira R, et al.
Published: (2024)
VeriGraph: Scene Graphs for Execution Verifiable Robot Planning
by: Ekpo, Daniel, et al.
Published: (2024)
by: Ekpo, Daniel, et al.
Published: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
by: Gwilliam, Matthew, et al.
Published: (2023)
by: Gwilliam, Matthew, et al.
Published: (2023)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
by: Garg, Abhinav, et al.
Published: (2024)
by: Garg, Abhinav, et al.
Published: (2024)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Anisotropic magnetization dynamics in Fe5GeTe2 at room temperature
by: Bera, Alapan, et al.
Published: (2024)
by: Bera, Alapan, et al.
Published: (2024)
Disorder-driven Weyl-Kondo Semimetal Phase in WTe$_2$
by: Manna, Arpan, et al.
Published: (2025)
by: Manna, Arpan, et al.
Published: (2025)
Angle dependent hysteretic magnetotransport in MnBi2Te4 nanoflakes
by: Das, Tithiparna, et al.
Published: (2026)
by: Das, Tithiparna, et al.
Published: (2026)
Metamagnetic quantum criticality in the antiferromagnetic topological insulator MnBi$_2$Te$_4$
by: Das, Tithiparna, et al.
Published: (2025)
by: Das, Tithiparna, et al.
Published: (2025)
Chiral orbital current driven topological Hall effect in Mn3Si2Te6
by: Das, Arnab, et al.
Published: (2025)
by: Das, Arnab, et al.
Published: (2025)
Tuning the chiral orbital currents in a colossal magnetoresistive nodal line ferrimagnet
by: Das, Arnab, et al.
Published: (2025)
by: Das, Arnab, et al.
Published: (2025)
Spin-reorientation driven topological Hall effect in Fe4GeTe2
by: Bera, Alapan, et al.
Published: (2025)
by: Bera, Alapan, et al.
Published: (2025)
Common topological origin of longitudinal and transverse magnetoresistance in Fe3GeTe2
by: Bera, Alapan, et al.
Published: (2025)
by: Bera, Alapan, et al.
Published: (2025)
Study of Magnetoresistance Plateau as a Probe of Spin‐Wave Excitations in Fe4$_4$GeTe2$_2$
by: Alapan Bera, et al.
Published: (2026)
by: Alapan Bera, et al.
Published: (2026)
V-VIPE: Variational View Invariant Pose Embedding
by: Levy, Mara, et al.
Published: (2024)
by: Levy, Mara, et al.
Published: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Charm quark and QGP interactions through the spectra and anisotropic flow of D$^0$ over the widest p$_\text{T}$ interval using event-shape engineering at CMS
by: Chandra, Soumik
Published: (2025)
by: Chandra, Soumik
Published: (2025)
A stabilizer code model with non-invertible symmetries: Strange fractons, confinement, and non-commutative and non-Abelian fusion rules
by: Kibe, Tanay, et al.
Published: (2023)
by: Kibe, Tanay, et al.
Published: (2023)
Learning complete and explainable visual representations from itemized text supervision
by: Lyu, Yiwei, et al.
Published: (2025)
by: Lyu, Yiwei, et al.
Published: (2025)
Remarks on the locality of generalized global symmetries
by: Gwilliam, Owen
Published: (2025)
by: Gwilliam, Owen
Published: (2025)
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)
by: Walmer, Matthew, et al.
Published: (2026)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
by: Agarwal, Abhinav
Published: (2026)
by: Agarwal, Abhinav
Published: (2026)
LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Kinematic Study of Molecular Gas in Cometary Globule -- LBN 437
by: Aardra, S., et al.
Published: (2025)
by: Aardra, S., et al.
Published: (2025)
Similar Items
-
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026) -
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025) -
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026) -
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024) -
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)