Scaling 4D Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Carreira, João, Gokay, Dilara, King, Michael, Zhang, Chuhan, Rocco, Ignacio, Mahendran, Aravindh, Keck, Thomas Albert, Heyward, Joseph, Koppula, Skanda, Pot, Etienne, Erdogan, Goker, Hasson, Yana, Yang, Yi, Greff, Klaus, Moing, Guillaume Le, van Steenkiste, Sjoerd, Zoran, Daniel, Hudson, Drew A., Vélez, Pedro, Polanía, Luisa, Friedman, Luke, Duvarney, Chris, Goroshin, Ross, Allen, Kelsey, Walker, Jacob, Kabra, Rishabh, Aboussouan, Eric, Sun, Jennifer, Kipf, Thomas, Doersch, Carl, Pătrăucean, Viorica, Damen, Dima, Luc, Pauline, Sajjadi, Mehdi S. M., Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perception Test 2025: Challenge Summary and a Unified VQA Extension
by: Heyward, Joseph, et al.
Published: (2026)
by: Heyward, Joseph, et al.
Published: (2026)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
by: Zholus, Artem, et al.
Published: (2025)
by: Zholus, Artem, et al.
Published: (2025)
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024)
by: Doersch, Carl, et al.
Published: (2024)
Moving Off-the-Grid: Scene-Grounded Video Representations
by: van Steenkiste, Sjoerd, et al.
Published: (2024)
by: van Steenkiste, Sjoerd, et al.
Published: (2024)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
Learning from Streaming Video with Orthogonal Gradients
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
by: Papalampidi, Pinelopi, et al.
Published: (2023)
by: Papalampidi, Pinelopi, et al.
Published: (2023)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024)
by: Koppula, Skanda, et al.
Published: (2024)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
by: Seitzer, Maximilian, et al.
Published: (2023)
by: Seitzer, Maximilian, et al.
Published: (2023)
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
TRecViT: A Recurrent Video Transformer
by: Pătrăucean, Viorica, et al.
Published: (2024)
by: Pătrăucean, Viorica, et al.
Published: (2024)
DORSal: Diffusion for Object-centric Representations of Scenes et al
by: Jabri, Allan, et al.
Published: (2023)
by: Jabri, Allan, et al.
Published: (2023)
Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
by: Zhang, Chuhan, et al.
Published: (2025)
by: Zhang, Chuhan, et al.
Published: (2025)
Direct Motion Models for Assessing Generated Videos
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026)
by: Kabra, Rishabh, et al.
Published: (2026)
How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Causal Evidence that Language Models use Confidence to Drive Behavior
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
From Image to Video: An Empirical Study of Diffusion Representations
by: Vélez, Pedro, et al.
Published: (2025)
by: Vélez, Pedro, et al.
Published: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
LayerLock: Non-collapsing Representation Learning with Progressive Freezing
by: Erdogan, Goker, et al.
Published: (2025)
by: Erdogan, Goker, et al.
Published: (2025)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Combustion Behaviour of Single Silicon Particles in Different Oxidizing Environments
by: Herman, et al.
Published: (2025)
by: Herman, et al.
Published: (2025)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Memory Consolidation Enables Long-Context Video Understanding
by: Balažević, Ivana, et al.
Published: (2024)
by: Balažević, Ivana, et al.
Published: (2024)
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
by: Kumaran, Dharshan, et al.
Published: (2025)
by: Kumaran, Dharshan, et al.
Published: (2025)
Frozen Forecasting: A Unified Evaluation
by: Walker, Jacob C, et al.
Published: (2025)
by: Walker, Jacob C, et al.
Published: (2025)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
by: Chalk, Jacob, et al.
Published: (2024)
by: Chalk, Jacob, et al.
Published: (2024)
Epic-Sounds: A Large-scale Dataset of Actions That Sound
by: Huh, Jaesung, et al.
Published: (2023)
by: Huh, Jaesung, et al.
Published: (2023)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Reseña de Kaufmant, Marie-Eugénie, «Le cheval au théâtre dans l’Espagne du Siècle d’Or. Fondements idéologiques et mécanismes d’une poétique dans la comedia nueva», Binges, Éditions Orbis Tertius, 2018, 576 pp. ISBN: 978-2-36783-113-8
by: Morgane Kappès-Le Moing
Published: (2019)
by: Morgane Kappès-Le Moing
Published: (2019)
How to Spin an Object: First, Get the Shape Right
by: Kabra, Rishabh, et al.
Published: (2024)
by: Kabra, Rishabh, et al.
Published: (2024)
TAPNext++: What's Next for Tracking Any Point (TAP)?
by: Jung, Sebastian, et al.
Published: (2026)
by: Jung, Sebastian, et al.
Published: (2026)
Home in A Hybrid World
by: Pot, Martin
Published: (2022)
by: Pot, Martin
Published: (2022)
Osons être des littéraires
by: Oliver Pot
Published: (2016)
by: Oliver Pot
Published: (2016)
Avaliação do pensamento crítico em contexto escolar: uma perspectiva emergente em psicologia
by: Viorica Alich
Published: (2016)
by: Viorica Alich
Published: (2016)
Making sense of the protests in Turkey (and Brazil): contesting neo-liberal urbanism in ‘Rebel Cities’
by: Bulent Gokay
Published: (2015)
by: Bulent Gokay
Published: (2015)
Similar Items
-
Perception Test 2025: Challenge Summary and a Unified VQA Extension
by: Heyward, Joseph, et al.
Published: (2026) -
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024) -
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
by: Zholus, Artem, et al.
Published: (2025) -
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024) -
Moving Off-the-Grid: Scene-Grounded Video Representations
by: van Steenkiste, Sjoerd, et al.
Published: (2024)