SF20K Competition 2025: Summary and findings
Fuente:
arXiv
Saved in:
| Main Authors: | Ghermi, Ridouane, Wang, Xi, Kalogeiton, Vicky, Laptev, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long Story Short: Story-level Video Understanding from 20K Short Films
by: Ghermi, Ridouane, et al.
Published: (2024)
by: Ghermi, Ridouane, et al.
Published: (2024)
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
by: Xia, Yingjie, et al.
Published: (2025)
by: Xia, Yingjie, et al.
Published: (2025)
E.T. the Exceptional Trajectories: Text-to-camera-trajectory generation with character awareness
by: Courant, Robin, et al.
Published: (2024)
by: Courant, Robin, et al.
Published: (2024)
AKiRa: Augmentation Kit on Rays for optical video generation
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Pulp Motion: Framing-aware multimodal camera and human motion generation
by: Courant, Robin, et al.
Published: (2025)
by: Courant, Robin, et al.
Published: (2025)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
Diffusion Reinforcement Learning via Centered Reward Distillation
by: Zhu, Yuanzhi, et al.
Published: (2026)
by: Zhu, Yuanzhi, et al.
Published: (2026)
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
by: Le, Minh-Quan, et al.
Published: (2025)
by: Le, Minh-Quan, et al.
Published: (2025)
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
by: Mordacq, Julie, et al.
Published: (2025)
by: Mordacq, Julie, et al.
Published: (2025)
Bridging Text and Image for Artist Style Transfer via Contrastive Learning
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
One-step Diffusion Models with Bregman Density Ratio Matching
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
by: Boudier, Luc, et al.
Published: (2025)
by: Boudier, Luc, et al.
Published: (2025)
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
by: Dufour, Nicolas, et al.
Published: (2025)
by: Dufour, Nicolas, et al.
Published: (2025)
ADAPT: Multimodal Learning for Detecting Physiological Changes under Missing Modalities
by: Mordacq, Julie, et al.
Published: (2024)
by: Mordacq, Julie, et al.
Published: (2024)
LEAD: Latent Realignment for Human Motion Diffusion
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
Analysis of Classifier-Free Guidance Weight Schedulers
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Name Your Style: An Arbitrary Artist-aware Image Style Transfer
by: Liu, Zhi-Song, et al.
Published: (2022)
by: Liu, Zhi-Song, et al.
Published: (2022)
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2025)
by: Benigmim, Yasser, et al.
Published: (2025)
Collaborating Foundation Models for Domain Generalized Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2023)
by: Benigmim, Yasser, et al.
Published: (2023)
FakeParts: a New Family of AI-Generated DeepFakes
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
ScanEdit: Hierarchically-Guided Functional 3D Scan Editing
by: Boudjoghra, Mohamed el amine, et al.
Published: (2025)
by: Boudjoghra, Mohamed el amine, et al.
Published: (2025)
InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
by: Zhang, Yangsong, et al.
Published: (2025)
by: Zhang, Yangsong, et al.
Published: (2025)
Mitigating Object Hallucination via Concentric Causal Attention
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
by: Souček, Tomáš, et al.
Published: (2023)
by: Souček, Tomáš, et al.
Published: (2023)
AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars
by: Fazylov, Ramazan, et al.
Published: (2025)
by: Fazylov, Ramazan, et al.
Published: (2025)
How far can we go with ImageNet for Text-to-Image generation?
by: Degeorge, L., et al.
Published: (2025)
by: Degeorge, L., et al.
Published: (2025)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024)
by: Souček, Tomáš, et al.
Published: (2024)
Perception Test 2025: Challenge Summary and a Unified VQA Extension
by: Heyward, Joseph, et al.
Published: (2026)
by: Heyward, Joseph, et al.
Published: (2026)
Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
by: Vitek, Matej, et al.
Published: (2025)
by: Vitek, Matej, et al.
Published: (2025)
GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic Fields
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025
by: Ma, Jingzhe, et al.
Published: (2026)
by: Ma, Jingzhe, et al.
Published: (2026)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
DiffSF: Diffusion Models for Scene Flow Estimation
by: Zhang, Yushan, et al.
Published: (2024)
by: Zhang, Yushan, et al.
Published: (2024)
tSF: Transformer-based Semantic Filter for Few-Shot Learning
by: Lai, Jinxiang, et al.
Published: (2022)
by: Lai, Jinxiang, et al.
Published: (2022)
Similar Items
-
Long Story Short: Story-level Video Understanding from 20K Short Films
by: Ghermi, Ridouane, et al.
Published: (2024) -
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
by: Xia, Yingjie, et al.
Published: (2025) -
E.T. the Exceptional Trajectories: Text-to-camera-trajectory generation with character awareness
by: Courant, Robin, et al.
Published: (2024) -
AKiRa: Augmentation Kit on Rays for optical video generation
by: Wang, Xi, et al.
Published: (2024) -
Pulp Motion: Framing-aware multimodal camera and human motion generation
by: Courant, Robin, et al.
Published: (2025)