How PARTs assemble into wholes: Learning the relative composition of images
Fuente:
arXiv
Saved in:
| Main Authors: | Ayoughi, Melika, Abnar, Samira, Huang, Chen, Sandino, Chris, Lala, Sayeri, Dhekane, Eeshan Gunesh, Busbridge, Dan, Zhai, Shuangfei, Thilak, Vimal, Susskind, Josh, Mettes, Pascal, Groth, Paul, Goh, Hanlin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minimizing Hyperbolic Embedding Distortion with LLM-Guided Hierarchy Restructuring
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025)
by: Abnar, Samira, et al.
Published: (2025)
Label-Efficient Sleep Staging Using Transformers Pre-trained with Position Prediction
by: Lala, Sayeri, et al.
Published: (2024)
by: Lala, Sayeri, et al.
Published: (2024)
Learning the relative composition of EEG signals using pairwise relative shift pretraining
by: Sandino, Christopher, et al.
Published: (2025)
by: Sandino, Christopher, et al.
Published: (2025)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
by: de Seyssel, Maureen, et al.
Published: (2025)
by: de Seyssel, Maureen, et al.
Published: (2025)
Poly-View Contrastive Learning
by: Shidani, Amitis, et al.
Published: (2024)
by: Shidani, Amitis, et al.
Published: (2024)
Continual Hyperbolic Learning of Instances and Classes
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026)
by: Ramapuram, Jason, et al.
Published: (2026)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Thou Shalt Not Prompt: Zero-Shot Human Activity Recognition in Smart Homes via Language Modeling of Sensor Data & Activities
by: Dhekane, Sourish Gunesh, et al.
Published: (2025)
by: Dhekane, Sourish Gunesh, et al.
Published: (2025)
Transfer Learning in Human Activity Recognition: A Survey
by: Dhekane, Sourish Gunesh, et al.
Published: (2024)
by: Dhekane, Sourish Gunesh, et al.
Published: (2024)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
TAD-SIE: Sample Size Estimation for Clinical Randomized Controlled Trials using a Trend-Adaptive Design with a Synthetic-Intervention-Based Estimator
by: Lala, Sayeri, et al.
Published: (2024)
by: Lala, Sayeri, et al.
Published: (2024)
METRIK: Measurement-Efficient Randomized Controlled Trials using Transformers with Input Masking
by: Lala, Sayeri, et al.
Published: (2024)
by: Lala, Sayeri, et al.
Published: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023)
by: Gu, Jiatao, et al.
Published: (2023)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
CONFINE: Conformal Prediction for Interpretable Neural Networks
by: Huang, Linhui, et al.
Published: (2024)
by: Huang, Linhui, et al.
Published: (2024)
PLANNER: Generating Diversified Paragraph via Latent Language Diffusion Model
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
Normalizing Trajectory Models
by: Gu, Jiatao, et al.
Published: (2026)
by: Gu, Jiatao, et al.
Published: (2026)
TADA: Improved Diffusion Sampling with Training-free Augmented Dynamics
by: Chen, Tianrong, et al.
Published: (2025)
by: Chen, Tianrong, et al.
Published: (2025)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024)
by: Zhang, Dinghuai, et al.
Published: (2024)
DISCOVER: Identifying Patterns of Daily Living in Human Activities from Smart Home Data
by: Karpekov, Alexander, et al.
Published: (2025)
by: Karpekov, Alexander, et al.
Published: (2025)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
by: Ramapuram, Jason, et al.
Published: (2024)
by: Ramapuram, Jason, et al.
Published: (2024)
Normalizing Flows with Iterative Denoising
by: Chen, Tianrong, et al.
Published: (2026)
by: Chen, Tianrong, et al.
Published: (2026)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
Layout Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)
by: Thukral, Megha, et al.
Published: (2024)
by: Thukral, Megha, et al.
Published: (2024)
Exclusive Self Attention
by: Zhai, Shuangfei
Published: (2026)
by: Zhai, Shuangfei
Published: (2026)
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
Vanishing Gradients in Reinforcement Finetuning of Language Models
by: Razin, Noam, et al.
Published: (2023)
by: Razin, Noam, et al.
Published: (2023)
Low-distortion and GPU-compatible Tree Embeddings in Hyperbolic Space
by: van Spengler, Max, et al.
Published: (2025)
by: van Spengler, Max, et al.
Published: (2025)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
by: Berthelot, David, et al.
Published: (2026)
by: Berthelot, David, et al.
Published: (2026)
Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
by: Mago, Gowreesh, et al.
Published: (2025)
by: Mago, Gowreesh, et al.
Published: (2025)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
Generative Modeling with Phase Stochastic Bridges
by: Chen, Tianrong, et al.
Published: (2023)
by: Chen, Tianrong, et al.
Published: (2023)
Similar Items
-
Minimizing Hyperbolic Embedding Distortion with LLM-Guided Hierarchy Restructuring
by: Ayoughi, Melika, et al.
Published: (2025) -
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025) -
Label-Efficient Sleep Staging Using Transformers Pre-trained with Position Prediction
by: Lala, Sayeri, et al.
Published: (2024) -
Learning the relative composition of EEG signals using pairwise relative shift pretraining
by: Sandino, Christopher, et al.
Published: (2025) -
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
by: de Seyssel, Maureen, et al.
Published: (2025)