Video Understanding by Design: How Datasets Shape Architectures and Insights
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Lei, Koniusz, Piotr, Gao, Yongsheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Graph Your Own Prompt
por: Ding, Xi, et al.
Publicado: (2025)
por: Ding, Xi, et al.
Publicado: (2025)
Learning Time in Static Classifiers
por: Ding, Xi, et al.
Publicado: (2025)
por: Ding, Xi, et al.
Publicado: (2025)
Subspace Kernel Learning on Tensor Sequences
por: Wang, Lei, et al.
Publicado: (2026)
por: Wang, Lei, et al.
Publicado: (2026)
Uncertainty-DTW for Sequences and Visual Tokens
por: Wang, Lei, et al.
Publicado: (2026)
por: Wang, Lei, et al.
Publicado: (2026)
Feature Hallucination for Self-supervised Action Recognition
por: Wang, Lei, et al.
Publicado: (2025)
por: Wang, Lei, et al.
Publicado: (2025)
Motion meets Attention: Video Motion Prompts
por: Chen, Qixiang, et al.
Publicado: (2024)
por: Chen, Qixiang, et al.
Publicado: (2024)
Adaptive Multi-head Contrastive Learning
por: Wang, Lei, et al.
Publicado: (2023)
por: Wang, Lei, et al.
Publicado: (2023)
Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint Alignment
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
por: Ding, Dexuan, et al.
Publicado: (2024)
por: Ding, Dexuan, et al.
Publicado: (2024)
Pre-training with Random Orthogonal Projection Image Modeling
por: Haghighat, Maryam, et al.
Publicado: (2023)
por: Haghighat, Maryam, et al.
Publicado: (2023)
Possibilistic Predictive Uncertainty for Deep Learning
por: Ni, Yao, et al.
Publicado: (2026)
por: Ni, Yao, et al.
Publicado: (2026)
Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection
por: Wang, Lei, et al.
Publicado: (2026)
por: Wang, Lei, et al.
Publicado: (2026)
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
por: Lin, Shen, et al.
Publicado: (2026)
por: Lin, Shen, et al.
Publicado: (2026)
Hierarchically Robust Zero-shot Vision-language Models
por: Dong, Junhao, et al.
Publicado: (2026)
por: Dong, Junhao, et al.
Publicado: (2026)
Trust-Aware Joint Feature-Prediction Discrepancy for Robust Domain Adaptation
por: Ding, Xi, et al.
Publicado: (2026)
por: Ding, Xi, et al.
Publicado: (2026)
Insights from the Use of Previously Unseen Neural Architecture Search Datasets
por: Geada, Rob, et al.
Publicado: (2024)
por: Geada, Rob, et al.
Publicado: (2024)
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
por: Ni, Yao, et al.
Publicado: (2025)
por: Ni, Yao, et al.
Publicado: (2025)
CHAIN: Enhancing Generalization in Data-Efficient GANs via lipsCHitz continuity constrAIned Normalization
por: Ni, Yao, et al.
Publicado: (2024)
por: Ni, Yao, et al.
Publicado: (2024)
CSA-Net: Channel-wise Spatially Autocorrelated Attention Networks
por: Nikzad, Nick, et al.
Publicado: (2024)
por: Nikzad, Nick, et al.
Publicado: (2024)
Do Language Models Understand Time?
por: Ding, Xi, et al.
Publicado: (2024)
por: Ding, Xi, et al.
Publicado: (2024)
Architecture, Dataset and Model-Scale Agnostic Data-free Meta-Learning
por: Hu, Zixuan, et al.
Publicado: (2023)
por: Hu, Zixuan, et al.
Publicado: (2023)
HourVideo: 1-Hour Video-Language Understanding
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2024)
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2024)
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
por: Chen, Yi, et al.
Publicado: (2025)
por: Chen, Yi, et al.
Publicado: (2025)
PRISM: Diversifying Dataset Distillation by Decoupling Architectural Priors
por: Moser, Brian B., et al.
Publicado: (2025)
por: Moser, Brian B., et al.
Publicado: (2025)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
por: Liao, Yi, et al.
Publicado: (2024)
por: Liao, Yi, et al.
Publicado: (2024)
TraNCE: Transformative Non-linear Concept Explainer for CNNs
por: Akpudo, Ugochukwu Ejike, et al.
Publicado: (2025)
por: Akpudo, Ugochukwu Ejike, et al.
Publicado: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
por: Song, Enxin, et al.
Publicado: (2025)
por: Song, Enxin, et al.
Publicado: (2025)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
por: Li, Xiaolong, et al.
Publicado: (2025)
por: Li, Xiaolong, et al.
Publicado: (2025)
Scale Efficient Training for Large Datasets
por: Zhou, Qing, et al.
Publicado: (2025)
por: Zhou, Qing, et al.
Publicado: (2025)
WiTUnet: A U-Shaped Architecture Integrating CNN and Transformer for Improved Feature Alignment and Local Information Fusion
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
MVR: Multi-view Video Reward Shaping for Reinforcement Learning
por: Luo, Lirui, et al.
Publicado: (2026)
por: Luo, Lirui, et al.
Publicado: (2026)
VIDEOP2R: Video Understanding from Perception to Reasoning
por: Jiang, Yifan, et al.
Publicado: (2025)
por: Jiang, Yifan, et al.
Publicado: (2025)
AutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding
por: Huang, Zihan, et al.
Publicado: (2024)
por: Huang, Zihan, et al.
Publicado: (2024)
Adaptive Keyframe Sampling for Long Video Understanding
por: Tang, Xi, et al.
Publicado: (2025)
por: Tang, Xi, et al.
Publicado: (2025)
Elucidating the Design Space of Dataset Condensation
por: Shao, Shitong, et al.
Publicado: (2024)
por: Shao, Shitong, et al.
Publicado: (2024)
PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularization
por: Ni, Yao, et al.
Publicado: (2024)
por: Ni, Yao, et al.
Publicado: (2024)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
por: Yuan, Yuqian, et al.
Publicado: (2024)
por: Yuan, Yuqian, et al.
Publicado: (2024)
From Data Statistics to Feature Geometry: How Correlations Shape Superposition
por: Prieto, Lucas, et al.
Publicado: (2026)
por: Prieto, Lucas, et al.
Publicado: (2026)
When Spatial meets Temporal in Action Recognition
por: Chen, Huilin, et al.
Publicado: (2024)
por: Chen, Huilin, et al.
Publicado: (2024)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
por: Zhu, Zirui, et al.
Publicado: (2025)
por: Zhu, Zirui, et al.
Publicado: (2025)
Ejemplares similares
-
Graph Your Own Prompt
por: Ding, Xi, et al.
Publicado: (2025) -
Learning Time in Static Classifiers
por: Ding, Xi, et al.
Publicado: (2025) -
Subspace Kernel Learning on Tensor Sequences
por: Wang, Lei, et al.
Publicado: (2026) -
Uncertainty-DTW for Sequences and Visual Tokens
por: Wang, Lei, et al.
Publicado: (2026) -
Feature Hallucination for Self-supervised Action Recognition
por: Wang, Lei, et al.
Publicado: (2025)