Online pre-training with long-form videos
Fuente:
arXiv
Saved in:
| Main Authors: | Kato, Itsuki, Kamiya, Kodai, Tamaki, Toru |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-model learning by sequential reading of untrimmed videos for action recognition
by: Kamiya, Kodai, et al.
Published: (2024)
by: Kamiya, Kodai, et al.
Published: (2024)
Shift and matching queries for video semantic segmentation
by: Mizuno, Tsubasa, et al.
Published: (2024)
by: Mizuno, Tsubasa, et al.
Published: (2024)
Fine-grained length controllable video captioning with ordinal embeddings
by: Nitta, Tomoya, et al.
Published: (2024)
by: Nitta, Tomoya, et al.
Published: (2024)
M3DDM+: An improved video outpainting by a modified masking strategy
by: Murakawa, Takuya, et al.
Published: (2026)
by: Murakawa, Takuya, et al.
Published: (2026)
Reflective Dialogue between Teacher and Solver Agents for Video Question Answering
by: Murakawa, Takuya, et al.
Published: (2026)
by: Murakawa, Takuya, et al.
Published: (2026)
Query matching for spatio-temporal action detection with query-based object detector
by: Hori, Shimon, et al.
Published: (2024)
by: Hori, Shimon, et al.
Published: (2024)
MoExDA: Domain Adaptation for Edge-based Action Recognition
by: Sugimoto, Takuya, et al.
Published: (2025)
by: Sugimoto, Takuya, et al.
Published: (2025)
BFMD: A Full-Match Badminton Dense Dataset for Dense Shot Captioning
by: Ding, Ning, et al.
Published: (2026)
by: Ding, Ning, et al.
Published: (2026)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
Disentangling Static and Dynamic Information for Reducing Static Bias in Action Recognition
by: Kobayashi, Masato, et al.
Published: (2025)
by: Kobayashi, Masato, et al.
Published: (2025)
Action tube generation by person query matching for spatio-temporal action detection
by: Omi, Kazuki, et al.
Published: (2025)
by: Omi, Kazuki, et al.
Published: (2025)
Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
by: Takama, Yusaku, et al.
Published: (2025)
by: Takama, Yusaku, et al.
Published: (2025)
Can masking background and object reduce static bias for zero-shot action recognition?
by: Fukuzawa, Takumi, et al.
Published: (2025)
by: Fukuzawa, Takumi, et al.
Published: (2025)
From Macro to Micro: Boosting micro-expression recognition via pre-training on macro-expression videos
by: Li, Hanting, et al.
Published: (2024)
by: Li, Hanting, et al.
Published: (2024)
Primitive Geometry Segment Pre-training for 3D Medical Image Segmentation
by: Tadokoro, Ryu, et al.
Published: (2024)
by: Tadokoro, Ryu, et al.
Published: (2024)
Inpainting-Driven Mask Optimization for Object Removal
by: Shimosato, Kodai, et al.
Published: (2024)
by: Shimosato, Kodai, et al.
Published: (2024)
AMEGO: Active Memory from long EGOcentric videos
by: Goletto, Gabriele, et al.
Published: (2024)
by: Goletto, Gabriele, et al.
Published: (2024)
Koala: Key frame-conditioned long video-LLM
by: Tan, Reuben, et al.
Published: (2024)
by: Tan, Reuben, et al.
Published: (2024)
Balancing long- and short-term dynamics for the modeling of saliency in videos
by: Wulff, Theodor, et al.
Published: (2025)
by: Wulff, Theodor, et al.
Published: (2025)
RobustFormer: Noise-Robust Pre-training for images and videos
by: Bastola, Ashish, et al.
Published: (2024)
by: Bastola, Ashish, et al.
Published: (2024)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
by: Venkataramanan, Shashanka, et al.
Published: (2023)
by: Venkataramanan, Shashanka, et al.
Published: (2023)
Anatomical grounding pre-training for medical phrase grounding
by: Zhang, Wenjun, et al.
Published: (2025)
by: Zhang, Wenjun, et al.
Published: (2025)
Cross-video Identity Correlating for Person Re-identification Pre-training
by: Zuo, Jialong, et al.
Published: (2024)
by: Zuo, Jialong, et al.
Published: (2024)
Novel Anomaly Detection Scenarios and Evaluation Metrics to Address the Ambiguity in the Definition of Normal Samples
by: Saito, Reiji, et al.
Published: (2026)
by: Saito, Reiji, et al.
Published: (2026)
Online 3D reconstruction and dense tracking in endoscopic videos
by: Hayoz, Michel, et al.
Published: (2024)
by: Hayoz, Michel, et al.
Published: (2024)
MGI: Multimodal Contrastive pre-training of Genomic and Medical Imaging
by: Zhou, Jiaying, et al.
Published: (2024)
by: Zhou, Jiaying, et al.
Published: (2024)
Anatomically-guided masked autoencoder pre-training for aneurysm detection
by: Ceballos-Arroyo, Alberto Mario, et al.
Published: (2025)
by: Ceballos-Arroyo, Alberto Mario, et al.
Published: (2025)
Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
by: Hirata, Kodai, et al.
Published: (2025)
by: Hirata, Kodai, et al.
Published: (2025)
Large-scale unsupervised audio pre-training for video-to-speech synthesis
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
From pre-training to downstream performance: Does domain-specific pre-training make sense?
by: Krones, Felix
Published: (2026)
by: Krones, Felix
Published: (2026)
General surgery vision transformer: A video pre-trained foundation model for general surgery
by: Schmidgall, Samuel, et al.
Published: (2024)
by: Schmidgall, Samuel, et al.
Published: (2024)
Segmenting the motion components of a video: A long-term unsupervised model
by: Meunier, Etienne, et al.
Published: (2023)
by: Meunier, Etienne, et al.
Published: (2023)
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
MVOC: a training-free multiple video object composition method with diffusion models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
by: Plou, Carlos, et al.
Published: (2025)
by: Plou, Carlos, et al.
Published: (2025)
Self-supervised transformer-based pre-training method with General Plant Infection dataset
by: Wang, Zhengle, et al.
Published: (2024)
by: Wang, Zhengle, et al.
Published: (2024)
Limitations of NERF with pre-trained Vision Features for Few-Shot 3D Reconstruction
by: Sanjyal, Ankit
Published: (2025)
by: Sanjyal, Ankit
Published: (2025)
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERT
by: Sengupta, Saurav, et al.
Published: (2023)
by: Sengupta, Saurav, et al.
Published: (2023)
Vascular anatomy-aware self-supervised pre-training for X-ray angiogram analysis
by: Huang, De-Xing, et al.
Published: (2026)
by: Huang, De-Xing, et al.
Published: (2026)
Audio-visual training for improved grounding in video-text LLMs
by: Sagare, Shivprasad, et al.
Published: (2024)
by: Sagare, Shivprasad, et al.
Published: (2024)
Similar Items
-
Multi-model learning by sequential reading of untrimmed videos for action recognition
by: Kamiya, Kodai, et al.
Published: (2024) -
Shift and matching queries for video semantic segmentation
by: Mizuno, Tsubasa, et al.
Published: (2024) -
Fine-grained length controllable video captioning with ordinal embeddings
by: Nitta, Tomoya, et al.
Published: (2024) -
M3DDM+: An improved video outpainting by a modified masking strategy
by: Murakawa, Takuya, et al.
Published: (2026) -
Reflective Dialogue between Teacher and Solver Agents for Video Question Answering
by: Murakawa, Takuya, et al.
Published: (2026)