How Important are Videos for Training Video LLMs?
Fuente:
arXiv
Guardado en:
| Autores principales: | Lydakis, George, Hermans, Alexander, Athar, Ali, de Geus, Daan, Leibe, Bastian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
por: Nekrasov, Alexey, et al.
Publicado: (2025)
por: Nekrasov, Alexey, et al.
Publicado: (2025)
DONUT: A Decoder-Only Model for Trajectory Prediction
por: Knoche, Markus, et al.
Publicado: (2025)
por: Knoche, Markus, et al.
Publicado: (2025)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024)
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
por: Knaebel, Karim, et al.
Publicado: (2025)
por: Knaebel, Karim, et al.
Publicado: (2025)
Your ViT is Secretly an Image Segmentation Model
por: Kerssies, Tommie, et al.
Publicado: (2025)
por: Kerssies, Tommie, et al.
Publicado: (2025)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
por: Norouzi, Narges, et al.
Publicado: (2026)
por: Norouzi, Narges, et al.
Publicado: (2026)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
por: Yilmaz, Kadir, et al.
Publicado: (2026)
por: Yilmaz, Kadir, et al.
Publicado: (2026)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
por: Knaebel, Karim, et al.
Publicado: (2023)
por: Knaebel, Karim, et al.
Publicado: (2023)
SurGe: Improved Surface Geometry in Point Maps
por: Knaebel, Karim, et al.
Publicado: (2026)
por: Knaebel, Karim, et al.
Publicado: (2026)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
por: Cavagnero, Niccolò, et al.
Publicado: (2026)
por: Cavagnero, Niccolò, et al.
Publicado: (2026)
Point-VOS: Pointing Up Video Object Segmentation
por: Zulfikar, Idil Esen, et al.
Publicado: (2024)
por: Zulfikar, Idil Esen, et al.
Publicado: (2024)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
por: de Geus, Daan, et al.
Publicado: (2024)
por: de Geus, Daan, et al.
Publicado: (2024)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
por: Piekenbrinck, Jens, et al.
Publicado: (2025)
por: Piekenbrinck, Jens, et al.
Publicado: (2025)
OoDIS: Anomaly Instance Segmentation and Detection Benchmark
por: Nekrasov, Alexey, et al.
Publicado: (2024)
por: Nekrasov, Alexey, et al.
Publicado: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
por: Kerssies, Tommie, et al.
Publicado: (2024)
por: Kerssies, Tommie, et al.
Publicado: (2024)
An Ordinal Regression Framework for a Deep Learning Based Severity Assessment for Chest Radiographs
por: Wienholt, Patrick, et al.
Publicado: (2024)
por: Wienholt, Patrick, et al.
Publicado: (2024)
Look Gauss, No Pose: Novel View Synthesis using Gaussian Splatting without Accurate Pose Initialization
por: Schmidt, Christian, et al.
Publicado: (2024)
por: Schmidt, Christian, et al.
Publicado: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
por: Norouzi, Narges, et al.
Publicado: (2024)
por: Norouzi, Narges, et al.
Publicado: (2024)
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
por: Athar, Ali, et al.
Publicado: (2024)
por: Athar, Ali, et al.
Publicado: (2024)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
por: Yilmaz, Kadir, et al.
Publicado: (2023)
por: Yilmaz, Kadir, et al.
Publicado: (2023)
An Empirical Study on How Video-LLMs Answer Video Questions
por: Gou, Chenhui, et al.
Publicado: (2025)
por: Gou, Chenhui, et al.
Publicado: (2025)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
por: Kerssies, Tommie, et al.
Publicado: (2024)
por: Kerssies, Tommie, et al.
Publicado: (2024)
LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation
por: Liu, Chang, et al.
Publicado: (2025)
por: Liu, Chang, et al.
Publicado: (2025)
MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
por: Uslu, Jan-Lucas, et al.
Publicado: (2024)
por: Uslu, Jan-Lucas, et al.
Publicado: (2024)
OCCUQ: Exploring Efficient Uncertainty Quantification for 3D Occupancy Prediction
por: Heidrich, Severin, et al.
Publicado: (2025)
por: Heidrich, Severin, et al.
Publicado: (2025)
Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
por: Wang, Chung-Shien Brian, et al.
Publicado: (2025)
por: Wang, Chung-Shien Brian, et al.
Publicado: (2025)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
por: Englert, Brunó B., et al.
Publicado: (2024)
por: Englert, Brunó B., et al.
Publicado: (2024)
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
por: Linder, Timm, et al.
Publicado: (2025)
por: Linder, Timm, et al.
Publicado: (2025)
Cyto R-CNN and CytoNuke Dataset: Towards reliable whole-cell segmentation in bright-field histological images
por: Raufeisen, Johannes, et al.
Publicado: (2024)
por: Raufeisen, Johannes, et al.
Publicado: (2024)
Interactive4D: Interactive 4D LiDAR Segmentation
por: Fradlin, Ilya, et al.
Publicado: (2024)
por: Fradlin, Ilya, et al.
Publicado: (2024)
Rig3DGS: Creating Controllable Portraits from Casual Monocular Videos
por: Rivero, Alfredo, et al.
Publicado: (2024)
por: Rivero, Alfredo, et al.
Publicado: (2024)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
por: Shvetsova, Nina, et al.
Publicado: (2023)
por: Shvetsova, Nina, et al.
Publicado: (2023)
Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
por: Nekrasov, Alexey, et al.
Publicado: (2025)
por: Nekrasov, Alexey, et al.
Publicado: (2025)
How to Design and Train Your Implicit Neural Representation for Video Compression
por: Gwilliam, Matthew, et al.
Publicado: (2025)
por: Gwilliam, Matthew, et al.
Publicado: (2025)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
por: Wang, Yueqian, et al.
Publicado: (2024)
por: Wang, Yueqian, et al.
Publicado: (2024)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
por: Song, Baiyang, et al.
Publicado: (2026)
por: Song, Baiyang, et al.
Publicado: (2026)
Enhancing Video Transformers for Action Understanding with VLM-aided Training
por: Lu, Hui, et al.
Publicado: (2024)
por: Lu, Hui, et al.
Publicado: (2024)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
por: Fateh, Fawad Javed, et al.
Publicado: (2024)
por: Fateh, Fawad Javed, et al.
Publicado: (2024)
TempCompass: Do Video LLMs Really Understand Videos?
por: Liu, Yuanxin, et al.
Publicado: (2024)
por: Liu, Yuanxin, et al.
Publicado: (2024)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
por: Kerssies, Tommie, et al.
Publicado: (2026)
por: Kerssies, Tommie, et al.
Publicado: (2026)
Ejemplares similares
-
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
por: Nekrasov, Alexey, et al.
Publicado: (2025) -
DONUT: A Decoder-Only Model for Trajectory Prediction
por: Knoche, Markus, et al.
Publicado: (2025) -
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024) -
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
por: Knaebel, Karim, et al.
Publicado: (2025) -
Your ViT is Secretly an Image Segmentation Model
por: Kerssies, Tommie, et al.
Publicado: (2025)