Scaling Properties of Diffusion Models for Perceptual Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ravishankar, Rahul, Patel, Zeeshan, Rajasegaran, Jathushan, Malik, Jitendra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Empirical Study of Autoregressive Pre-training from Videos
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
Gaussian Masked Autoencoders
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
Synthesizing Moving People with 3D Control
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Poly-Autoregressive Prediction for Modeling Interactions
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025)
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025)
Tracking by Predicting 3-D Gaussians Over Time
von: Baranwal, Tanish, et al.
Veröffentlicht: (2025)
von: Baranwal, Tanish, et al.
Veröffentlicht: (2025)
Synergy and Synchrony in Couple Dances
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024)
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024)
FewShotNeRF: Meta-Learning-based Novel View Synthesis for Rapid Scene-Specific Adaptation
von: Sivakumar, Piraveen, et al.
Veröffentlicht: (2024)
von: Sivakumar, Piraveen, et al.
Veröffentlicht: (2024)
Latent Guidance in Diffusion Models for Perceptual Evaluations
von: Saini, Shreshth, et al.
Veröffentlicht: (2025)
von: Saini, Shreshth, et al.
Veröffentlicht: (2025)
BudgetFusion: Perceptually-Guided Adaptive Diffusion Models
von: Li, Qinchan, et al.
Veröffentlicht: (2024)
von: Li, Qinchan, et al.
Veröffentlicht: (2024)
Perceptual Evaluation of GANs and Diffusion Models for Generating X-rays
von: Schuit, Gregory, et al.
Veröffentlicht: (2025)
von: Schuit, Gregory, et al.
Veröffentlicht: (2025)
Diffusion Model with Perceptual Loss
von: Lin, Shanchuan, et al.
Veröffentlicht: (2023)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2023)
Humanoid Locomotion as Next Token Prediction
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
Scaling-up Perceptual Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2025)
von: Jia, Ziheng, et al.
Veröffentlicht: (2025)
ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
Denoising Task Routing for Diffusion Models
von: Park, Byeongjun, et al.
Veröffentlicht: (2023)
von: Park, Byeongjun, et al.
Veröffentlicht: (2023)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
von: Tian, Ran, et al.
Veröffentlicht: (2023)
von: Tian, Ran, et al.
Veröffentlicht: (2023)
Probing Perceptual Constancy in Large Vision-Language Models
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
Your Pre-trained Diffusion Model Secretly Knows Restoration
von: Rajagopalan, Sudarshan, et al.
Veröffentlicht: (2026)
von: Rajagopalan, Sudarshan, et al.
Veröffentlicht: (2026)
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
von: Carrión, Héctor, et al.
Veröffentlicht: (2025)
von: Carrión, Héctor, et al.
Veröffentlicht: (2025)
Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
von: Ma, Qianli, et al.
Veröffentlicht: (2024)
von: Ma, Qianli, et al.
Veröffentlicht: (2024)
Diffusion Model in Latent Space for Medical Image Segmentation Task
von: Ngoc, Huynh Trinh, et al.
Veröffentlicht: (2025)
von: Ngoc, Huynh Trinh, et al.
Veröffentlicht: (2025)
Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities
von: Chavhan, Ruchika, et al.
Veröffentlicht: (2025)
von: Chavhan, Ruchika, et al.
Veröffentlicht: (2025)
OrthoDiffusion: A Generalizable Multi-Task Diffusion Foundation Model for Musculoskeletal MRI Interpretation
von: Lan, Tian, et al.
Veröffentlicht: (2026)
von: Lan, Tian, et al.
Veröffentlicht: (2026)
Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
von: Geng, Zichen, et al.
Veröffentlicht: (2025)
von: Geng, Zichen, et al.
Veröffentlicht: (2025)
Gradient-Free Classifier Guidance for Diffusion Model Sampling
von: Shenoy, Rahul, et al.
Veröffentlicht: (2024)
von: Shenoy, Rahul, et al.
Veröffentlicht: (2024)
Scale Space Diffusion
von: Mukhopadhyay, Soumik, et al.
Veröffentlicht: (2026)
von: Mukhopadhyay, Soumik, et al.
Veröffentlicht: (2026)
Latent Feature-Guided Diffusion Models for Shadow Removal
von: Mei, Kangfu, et al.
Veröffentlicht: (2023)
von: Mei, Kangfu, et al.
Veröffentlicht: (2023)
Predicting and Enhancing the Fairness of DNNs with the Curvature of Perceptual Manifolds
von: Ma, Yanbiao, et al.
Veröffentlicht: (2023)
von: Ma, Yanbiao, et al.
Veröffentlicht: (2023)
Perceptual Group Tokenizer: Building Perception with Iterative Grouping
von: Deng, Zhiwei, et al.
Veröffentlicht: (2023)
von: Deng, Zhiwei, et al.
Veröffentlicht: (2023)
Upsample Guidance: Scale Up Diffusion Models without Training
von: Hwang, Juno, et al.
Veröffentlicht: (2024)
von: Hwang, Juno, et al.
Veröffentlicht: (2024)
Estimating Body and Hand Motion in an Ego-sensed World
von: Yi, Brent, et al.
Veröffentlicht: (2024)
von: Yi, Brent, et al.
Veröffentlicht: (2024)
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
von: Bang, Junseo, et al.
Veröffentlicht: (2025)
von: Bang, Junseo, et al.
Veröffentlicht: (2025)
HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
von: Kwon, Young D., et al.
Veröffentlicht: (2025)
von: Kwon, Young D., et al.
Veröffentlicht: (2025)
Symmetrization Weighted Binary Cross-Entropy: Modeling Perceptual Asymmetry for Human-Consistent Neural Edge Detection
von: Shu, Hao
Veröffentlicht: (2025)
von: Shu, Hao
Veröffentlicht: (2025)
Hand-Object Interaction Pretraining from Videos
von: Singh, Himanshu Gaurav, et al.
Veröffentlicht: (2024)
von: Singh, Himanshu Gaurav, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Empirical Study of Autoregressive Pre-training from Videos
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025) -
Gaussian Masked Autoencoders
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025) -
Synthesizing Moving People with 3D Control
von: Li, Boyi, et al.
Veröffentlicht: (2024) -
Poly-Autoregressive Prediction for Modeling Interactions
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025) -
Tracking by Predicting 3-D Gaussians Over Time
von: Baranwal, Tanish, et al.
Veröffentlicht: (2025)