SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
Fuente:
arXiv
Guardado en:
| Autores principales: | Hasson, Yana, Luc, Pauline, Momeni, Liliane, Ovsjanikov, Maks, Moing, Guillaume Le, Kuznetsova, Alina, Ktena, Ira, Sun, Jennifer J., Koppula, Skanda, Gokay, Dilara, Heyward, Joseph, Pot, Etienne, Zisserman, Andrew |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BootsTAP: Bootstrapped Training for Tracking-Any-Point
por: Doersch, Carl, et al.
Publicado: (2024)
por: Doersch, Carl, et al.
Publicado: (2024)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
por: Koppula, Skanda, et al.
Publicado: (2024)
por: Koppula, Skanda, et al.
Publicado: (2024)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
por: Kabra, Rishabh, et al.
Publicado: (2026)
por: Kabra, Rishabh, et al.
Publicado: (2026)
Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
por: Zhang, Chuhan, et al.
Publicado: (2025)
por: Zhang, Chuhan, et al.
Publicado: (2025)
Unique Lives, Shared World: Learning from Single-Life Videos
por: Han, Tengda, et al.
Publicado: (2025)
por: Han, Tengda, et al.
Publicado: (2025)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
por: Papalampidi, Pinelopi, et al.
Publicado: (2023)
por: Papalampidi, Pinelopi, et al.
Publicado: (2023)
Scaling 4D Representations
por: Carreira, João, et al.
Publicado: (2024)
por: Carreira, João, et al.
Publicado: (2024)
Smoothed Graph Contrastive Learning via Seamless Proximity Integration
por: Behmanesh, Maysam, et al.
Publicado: (2024)
por: Behmanesh, Maysam, et al.
Publicado: (2024)
Shape Non-rigid Kinematics (SNK): A Zero-Shot Method for Non-Rigid Shape Matching via Unsupervised Functional Map Regularized Reconstruction
por: Attaiki, Souhaib, et al.
Publicado: (2024)
por: Attaiki, Souhaib, et al.
Publicado: (2024)
Memory-Scalable and Simplified Functional Map Learning
por: Magnet, Robin, et al.
Publicado: (2024)
por: Magnet, Robin, et al.
Publicado: (2024)
FILTR: Extracting Topological Features from Pretrained 3D Models
por: Martinez, Louis, et al.
Publicado: (2026)
por: Martinez, Louis, et al.
Publicado: (2026)
Learning from Streaming Video with Orthogonal Gradients
por: Han, Tengda, et al.
Publicado: (2025)
por: Han, Tengda, et al.
Publicado: (2025)
Deformation Recovery: Localized Learning for Detail-Preserving Deformations
por: Sundararaman, Ramana, et al.
Publicado: (2024)
por: Sundararaman, Ramana, et al.
Publicado: (2024)
Generative Drifting is Secretly Score Matching: a Spectral and Variational Perspective
por: Turan, Erkan, et al.
Publicado: (2026)
por: Turan, Erkan, et al.
Publicado: (2026)
Self-Supervised Dual Contouring
por: Sundararaman, Ramana, et al.
Publicado: (2024)
por: Sundararaman, Ramana, et al.
Publicado: (2024)
FourieRF: Few-Shot NeRFs via Progressive Fourier Frequency Control
por: Gomez, Diego, et al.
Publicado: (2025)
por: Gomez, Diego, et al.
Publicado: (2025)
Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D Features
por: Wimmer, Thomas, et al.
Publicado: (2023)
por: Wimmer, Thomas, et al.
Publicado: (2023)
Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
por: Behmanesh, Maysam, et al.
Publicado: (2025)
por: Behmanesh, Maysam, et al.
Publicado: (2025)
To Supervise or Not to Supervise: Understanding and Addressing the Key Challenges of Point Cloud Transfer Learning
por: Hadgi, Souhail, et al.
Publicado: (2024)
por: Hadgi, Souhail, et al.
Publicado: (2024)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
por: Jang, Youngjoon, et al.
Publicado: (2025)
por: Jang, Youngjoon, et al.
Publicado: (2025)
Learning from One Continuous Video Stream
por: Carreira, João, et al.
Publicado: (2023)
por: Carreira, João, et al.
Publicado: (2023)
DiffuMatch: Category-Agnostic Spectral Diffusion Priors for Robust Non-rigid Shape Matching
por: Pierson, Emery, et al.
Publicado: (2025)
por: Pierson, Emery, et al.
Publicado: (2025)
PoNQ: a Neural QEM-based Mesh Representation
por: Maruani, Nissim, et al.
Publicado: (2024)
por: Maruani, Nissim, et al.
Publicado: (2024)
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
por: Jiang, Zifan, et al.
Publicado: (2025)
por: Jiang, Zifan, et al.
Publicado: (2025)
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
por: Jang, Youngjoon, et al.
Publicado: (2025)
por: Jang, Youngjoon, et al.
Publicado: (2025)
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
por: Chen, Victoria Yue, et al.
Publicado: (2026)
por: Chen, Victoria Yue, et al.
Publicado: (2026)
DeBaRA: Denoising-Based 3D Room Arrangement Generation
por: Maillard, Léopold, et al.
Publicado: (2024)
por: Maillard, Léopold, et al.
Publicado: (2024)
LACONIC: A 3D Layout Adapter for Controllable Image Creation
por: Maillard, Léopold, et al.
Publicado: (2025)
por: Maillard, Léopold, et al.
Publicado: (2025)
From Blobs to Spokes: High-Fidelity Surface Reconstruction via Oriented Gaussians
por: Gomez, Diego, et al.
Publicado: (2026)
por: Gomez, Diego, et al.
Publicado: (2026)
PaNDaS: Learnable Deformation Modeling with Localized Control
por: Besnier, Thomas, et al.
Publicado: (2024)
por: Besnier, Thomas, et al.
Publicado: (2024)
Dynamic Reflections: Probing Video Representations with Text Alignment
por: Zhu, Tyler, et al.
Publicado: (2025)
por: Zhu, Tyler, et al.
Publicado: (2025)
AtomSurf : Surface Representation for Learning on Protein Structures
por: Mallet, Vincent, et al.
Publicado: (2023)
por: Mallet, Vincent, et al.
Publicado: (2023)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
por: Heyward, Joseph, et al.
Publicado: (2024)
por: Heyward, Joseph, et al.
Publicado: (2024)
Memory Consolidation Enables Long-Context Video Understanding
por: Balažević, Ivana, et al.
Publicado: (2024)
por: Balažević, Ivana, et al.
Publicado: (2024)
A Tale of Two Languages: Large-Vocabulary Continuous Sign Language Recognition from Spoken Language Supervision
por: Raude, Charles, et al.
Publicado: (2024)
por: Raude, Charles, et al.
Publicado: (2024)
Moving Off-the-Grid: Scene-Grounded Video Representations
por: van Steenkiste, Sjoerd, et al.
Publicado: (2024)
por: van Steenkiste, Sjoerd, et al.
Publicado: (2024)
Medical Context Distorts Decisions in Clinical Vision Language Models
por: Restrepo, David, et al.
Publicado: (2026)
por: Restrepo, David, et al.
Publicado: (2026)
On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
por: Restrepo, David, et al.
Publicado: (2025)
por: Restrepo, David, et al.
Publicado: (2025)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
por: Attaiki, Souhaib, et al.
Publicado: (2024)
por: Attaiki, Souhaib, et al.
Publicado: (2024)
MILo: Mesh-In-the-Loop Gaussian Splatting for Detailed and Efficient Surface Reconstruction
por: Guédon, Antoine, et al.
Publicado: (2025)
por: Guédon, Antoine, et al.
Publicado: (2025)
Ejemplares similares
-
BootsTAP: Bootstrapped Training for Tracking-Any-Point
por: Doersch, Carl, et al.
Publicado: (2024) -
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
por: Koppula, Skanda, et al.
Publicado: (2024) -
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
por: Kabra, Rishabh, et al.
Publicado: (2026) -
Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
por: Zhang, Chuhan, et al.
Publicado: (2025) -
Unique Lives, Shared World: Learning from Single-Life Videos
por: Han, Tengda, et al.
Publicado: (2025)