Learning Video Representations without Natural Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Xueyang, Chen, Xinlei, Gandelsman, Yossi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Time Training on Video Streams
by: Wang, Renhao, et al.
Published: (2023)
by: Wang, Renhao, et al.
Published: (2023)
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
by: Ekin, Yigit, et al.
Published: (2026)
by: Ekin, Yigit, et al.
Published: (2026)
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
by: Bu, Edmund, et al.
Published: (2025)
by: Bu, Edmund, et al.
Published: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
by: Jiang, Nick, et al.
Published: (2024)
by: Jiang, Nick, et al.
Published: (2024)
An Empirical Study of Autoregressive Pre-training from Videos
by: Rajasegaran, Jathushan, et al.
Published: (2025)
by: Rajasegaran, Jathushan, et al.
Published: (2025)
Interpreting the Second-Order Effects of Neurons in CLIP
by: Gandelsman, Yossi, et al.
Published: (2024)
by: Gandelsman, Yossi, et al.
Published: (2024)
The More You See in 2D, the More You Perceive in 3D
by: Han, Xinyang, et al.
Published: (2024)
by: Han, Xinyang, et al.
Published: (2024)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
Learning Natural Consistency Representation for Face Forgery Video Detection
by: Zhang, Daichi, et al.
Published: (2024)
by: Zhang, Daichi, et al.
Published: (2024)
Teaching Humans Subtle Differences with DIFFusion
by: Chiquier, Mia, et al.
Published: (2025)
by: Chiquier, Mia, et al.
Published: (2025)
GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation
by: Lin, Lang, et al.
Published: (2025)
by: Lin, Lang, et al.
Published: (2025)
Video Depth without Video Models
by: Ke, Bingxin, et al.
Published: (2024)
by: Ke, Bingxin, et al.
Published: (2024)
Synthesizing Moving People with 3D Control
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
by: Yan, Tianyi, et al.
Published: (2025)
by: Yan, Tianyi, et al.
Published: (2025)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
Multi-entity Video Transformers for Fine-Grained Video Representation Learning
by: Walmer, Matthew, et al.
Published: (2023)
by: Walmer, Matthew, et al.
Published: (2023)
Jailbreaking Vision-Language Models Through the Visual Modality
by: Azulay, Aharon, et al.
Published: (2026)
by: Azulay, Aharon, et al.
Published: (2026)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)
by: Si, Chenyang, et al.
Published: (2025)
SEAL: Semantic Attention Learning for Long Video Representation
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
by: Wang, Lirui, et al.
Published: (2025)
by: Wang, Lirui, et al.
Published: (2025)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
SPKLIP: Aligning Spike Video Streams with Natural Language
by: Gao, Yongchang, et al.
Published: (2025)
by: Gao, Yongchang, et al.
Published: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Clapper: Compact Learning and Video Representation in VLMs
by: Kong, Lingyu, et al.
Published: (2025)
by: Kong, Lingyu, et al.
Published: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Still-Moving: Customized Video Generation without Customized Video Data
by: Chefer, Hila, et al.
Published: (2024)
by: Chefer, Hila, et al.
Published: (2024)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
by: Dou, Zi-Yi, et al.
Published: (2024)
by: Dou, Zi-Yi, et al.
Published: (2024)
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
by: Thoker, Fida Mohammad, et al.
Published: (2025)
by: Thoker, Fida Mohammad, et al.
Published: (2025)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Similar Items
-
Test-Time Training on Video Streams
by: Wang, Renhao, et al.
Published: (2023) -
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025) -
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
by: Ekin, Yigit, et al.
Published: (2026) -
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
by: Bu, Edmund, et al.
Published: (2025) -
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)