How to Spin an Object: First, Get the Shape Right
Fuente:
arXiv
Saved in:
| Main Authors: | Kabra, Rishabh, Hudson, Drew A., van Steenkiste, Sjoerd, Carreira, Joao, Mitra, Niloy J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026)
by: Kabra, Rishabh, et al.
Published: (2026)
Moving Off-the-Grid: Scene-Grounded Video Representations
by: van Steenkiste, Sjoerd, et al.
Published: (2024)
by: van Steenkiste, Sjoerd, et al.
Published: (2024)
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023)
by: Kabra, Rishabh, et al.
Published: (2023)
DORSal: Diffusion for Object-centric Representations of Scenes et al
by: Jabri, Allan, et al.
Published: (2023)
by: Jabri, Allan, et al.
Published: (2023)
Direct Motion Models for Assessing Generated Videos
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
by: Seitzer, Maximilian, et al.
Published: (2023)
by: Seitzer, Maximilian, et al.
Published: (2023)
MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
Frozen Forecasting: A Unified Evaluation
by: Walker, Jacob C, et al.
Published: (2025)
by: Walker, Jacob C, et al.
Published: (2025)
LoST: Level of Semantics Tokenization for 3D Shapes
by: Dutt, Niladri Shekhar, et al.
Published: (2026)
by: Dutt, Niladri Shekhar, et al.
Published: (2026)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
Scaling 4D Representations
by: Carreira, João, et al.
Published: (2024)
by: Carreira, João, et al.
Published: (2024)
OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts
by: Xiao, Shiting, et al.
Published: (2025)
by: Xiao, Shiting, et al.
Published: (2025)
RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
by: Anciukevičius, Titas, et al.
Published: (2022)
by: Anciukevičius, Titas, et al.
Published: (2022)
Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
by: Majee, Anay, et al.
Published: (2025)
by: Majee, Anay, et al.
Published: (2025)
From Image to Video: An Empirical Study of Diffusion Representations
by: Vélez, Pedro, et al.
Published: (2025)
by: Vélez, Pedro, et al.
Published: (2025)
Towards Robust Cross-Dataset Object Detection Generalization under Domain Specificity
by: Chakraborty, Ritabrata, et al.
Published: (2026)
by: Chakraborty, Ritabrata, et al.
Published: (2026)
Beyond Attention Heatmaps: How to Get Better Explanations for Multiple Instance Learning Models in Histopathology
by: Idaji, Mina Jamshidi, et al.
Published: (2026)
by: Idaji, Mina Jamshidi, et al.
Published: (2026)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
Unified Knowledge Distillation Framework: Fine-Grained Alignment and Geometric Relationship Preservation for Deep Face Recognition
by: Mishra, Durgesh, et al.
Published: (2025)
by: Mishra, Durgesh, et al.
Published: (2025)
SHaSaM: Submodular Hard Sample Mining for Fair Facial Attribute Recognition
by: Majee, Anay, et al.
Published: (2026)
by: Majee, Anay, et al.
Published: (2026)
EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition
by: Gupta, Rishabh, et al.
Published: (2026)
by: Gupta, Rishabh, et al.
Published: (2026)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
Multimodal Object Query Initialization for 3D Object Detection
by: van Geerenstein, Mathijs R., et al.
Published: (2023)
by: van Geerenstein, Mathijs R., et al.
Published: (2023)
Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features
by: Dutt, Niladri Shekhar, et al.
Published: (2023)
by: Dutt, Niladri Shekhar, et al.
Published: (2023)
BLiSS: Bootstrapped Linear Shape Space
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
How Weight Resampling and Optimizers Shape the Dynamics of Continual Learning and Forgetting in Neural Networks
by: Frati, Lapo, et al.
Published: (2025)
by: Frati, Lapo, et al.
Published: (2025)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
by: Loeffler, Christoffer, et al.
Published: (2022)
by: Loeffler, Christoffer, et al.
Published: (2022)
AI-CNet3D: An Anatomically-Informed Cross-Attention Network with Multi-Task Consistency Fine-tuning for 3D Glaucoma Classification
by: Kenia, Roshan, et al.
Published: (2025)
by: Kenia, Roshan, et al.
Published: (2025)
Do Your Best and Get Enough Rest for Continual Learning
by: Kang, Hankyul, et al.
Published: (2025)
by: Kang, Hankyul, et al.
Published: (2025)
Diverse Part Synthesis for 3D Shape Creation
by: Guan, Yanran, et al.
Published: (2024)
by: Guan, Yanran, et al.
Published: (2024)
SCoRe: Submodular Combinatorial Representation Learning
by: Majee, Anay, et al.
Published: (2023)
by: Majee, Anay, et al.
Published: (2023)
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024)
by: Doersch, Carl, et al.
Published: (2024)
From Data Statistics to Feature Geometry: How Correlations Shape Superposition
by: Prieto, Lucas, et al.
Published: (2026)
by: Prieto, Lucas, et al.
Published: (2026)
Cross-View World Models
by: Sharma, Rishabh, et al.
Published: (2026)
by: Sharma, Rishabh, et al.
Published: (2026)
Abstract Art Interpretation Using ControlNet
by: Srivastava, Rishabh, et al.
Published: (2024)
by: Srivastava, Rishabh, et al.
Published: (2024)
Enhancing Tea Leaf Disease Recognition with Attention Mechanisms and Grad-CAM Visualization
by: Shikdar, Omar Faruq, et al.
Published: (2025)
by: Shikdar, Omar Faruq, et al.
Published: (2025)
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024)
by: Petty, Jackson, et al.
Published: (2024)
Similar Items
-
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024) -
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026) -
Moving Off-the-Grid: Scene-Grounded Video Representations
by: van Steenkiste, Sjoerd, et al.
Published: (2024) -
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023) -
DORSal: Diffusion for Object-centric Representations of Scenes et al
by: Jabri, Allan, et al.
Published: (2023)