Action-conditioned video data improves predictability
Fuente:
arXiv
Salvato in:
| Autori principali: | Sarkar, Meenakshi, Ghose, Debasish |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Video Generation with Learned Action Prior
di: Sarkar, Meenakshi, et al.
Pubblicazione: (2024)
di: Sarkar, Meenakshi, et al.
Pubblicazione: (2024)
Can 3D point cloud data improve automated body condition score prediction in dairy cattle?
di: Tang, Zhou, et al.
Pubblicazione: (2026)
di: Tang, Zhou, et al.
Pubblicazione: (2026)
Meta Episodic learning with Dynamic Task Sampling for CLIP-based Point Cloud Classification
di: Ghose, Shuvozit, et al.
Pubblicazione: (2024)
di: Ghose, Shuvozit, et al.
Pubblicazione: (2024)
Koala: Key frame-conditioned long video-LLM
di: Tan, Reuben, et al.
Pubblicazione: (2024)
di: Tan, Reuben, et al.
Pubblicazione: (2024)
Don't Pause! Every prediction matters in a streaming video
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2026)
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2026)
A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters
di: Desai, Shasvat, et al.
Pubblicazione: (2025)
di: Desai, Shasvat, et al.
Pubblicazione: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
M3DDM+: An improved video outpainting by a modified masking strategy
di: Murakawa, Takuya, et al.
Pubblicazione: (2026)
di: Murakawa, Takuya, et al.
Pubblicazione: (2026)
Exploiting temporal information to detect conversational groups in videos and predict the next speaker
di: Tosato, Lucrezia, et al.
Pubblicazione: (2024)
di: Tosato, Lucrezia, et al.
Pubblicazione: (2024)
Autoregression-free video prediction using diffusion model for mitigating error propagation
di: Ko, Woonho, et al.
Pubblicazione: (2025)
di: Ko, Woonho, et al.
Pubblicazione: (2025)
AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance
di: Klein, Benjamin, et al.
Pubblicazione: (2026)
di: Klein, Benjamin, et al.
Pubblicazione: (2026)
CLIP-based Point Cloud Classification via Point Cloud to Image Translation
di: Ghose, Shuvozit, et al.
Pubblicazione: (2024)
di: Ghose, Shuvozit, et al.
Pubblicazione: (2024)
Bidirectional skip-frame prediction for video anomaly detection with intra-domain disparity-driven attention
di: Lyu, Jiahao, et al.
Pubblicazione: (2024)
di: Lyu, Jiahao, et al.
Pubblicazione: (2024)
Audio-visual training for improved grounding in video-text LLMs
di: Sagare, Shivprasad, et al.
Pubblicazione: (2024)
di: Sagare, Shivprasad, et al.
Pubblicazione: (2024)
Image Segmentation with transformers: An Overview, Challenges and Future
di: Chetia, Deepjyoti, et al.
Pubblicazione: (2025)
di: Chetia, Deepjyoti, et al.
Pubblicazione: (2025)
Robust soybean seed yield estimation using high-throughput ground robot videos
di: Feng, Jiale, et al.
Pubblicazione: (2024)
di: Feng, Jiale, et al.
Pubblicazione: (2024)
LiteVLA-H: Dual-Rate Vision-Language-Action Inference for Onboard Aerial Guidance and Semantic Perception
di: williams, Justin, et al.
Pubblicazione: (2026)
di: williams, Justin, et al.
Pubblicazione: (2026)
Foul prediction with estimated poses from soccer broadcast video
di: Fang, Jiale, et al.
Pubblicazione: (2024)
di: Fang, Jiale, et al.
Pubblicazione: (2024)
Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
di: Sun, Jian, et al.
Pubblicazione: (2026)
di: Sun, Jian, et al.
Pubblicazione: (2026)
EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-based Vision
di: Torbunov, Dmitrii, et al.
Pubblicazione: (2024)
di: Torbunov, Dmitrii, et al.
Pubblicazione: (2024)
3D Gaussian Splatting with Normal Information for Mesh Extraction and Improved Rendering
di: Krishnan, Meenakshi, et al.
Pubblicazione: (2025)
di: Krishnan, Meenakshi, et al.
Pubblicazione: (2025)
Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions
di: Hariat, Marwane, et al.
Pubblicazione: (2026)
di: Hariat, Marwane, et al.
Pubblicazione: (2026)
Test-time augmentation improves efficiency in conformal prediction
di: Shanmugam, Divya, et al.
Pubblicazione: (2025)
di: Shanmugam, Divya, et al.
Pubblicazione: (2025)
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
di: Xu, Kechun, et al.
Pubblicazione: (2025)
di: Xu, Kechun, et al.
Pubblicazione: (2025)
Action-guided generation of 3D functionality segmentation data
di: Corsetti, Jaime, et al.
Pubblicazione: (2025)
di: Corsetti, Jaime, et al.
Pubblicazione: (2025)
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks
di: Ghose, Partho, et al.
Pubblicazione: (2026)
di: Ghose, Partho, et al.
Pubblicazione: (2026)
Adaptive local boundary conditions to improve Deformable Image Registration
di: Inacio, Eloïse, et al.
Pubblicazione: (2024)
di: Inacio, Eloïse, et al.
Pubblicazione: (2024)
DLM-VMTL:A Double Layer Mapper for heterogeneous data video Multi-task prompt learning
di: Bo, Zeyi, et al.
Pubblicazione: (2024)
di: Bo, Zeyi, et al.
Pubblicazione: (2024)
Real time anomalies detection on video
di: Poirier, Fabien
Pubblicazione: (2024)
di: Poirier, Fabien
Pubblicazione: (2024)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts
di: Chhipa, Prakash Chandra, et al.
Pubblicazione: (2024)
di: Chhipa, Prakash Chandra, et al.
Pubblicazione: (2024)
LCM: Log Conformal Maps for Robust Representation Learning to Mitigate Perspective Distortion
di: Chippa, Meenakshi Subhash, et al.
Pubblicazione: (2024)
di: Chippa, Meenakshi Subhash, et al.
Pubblicazione: (2024)
How good are deep learning methods for automated road safety analysis using video data? An experimental study
di: Liu, Qingwu, et al.
Pubblicazione: (2025)
di: Liu, Qingwu, et al.
Pubblicazione: (2025)
Video prediction using score-based conditional density estimation
di: Fiquet, Pierre-Étienne H., et al.
Pubblicazione: (2024)
di: Fiquet, Pierre-Étienne H., et al.
Pubblicazione: (2024)
Action-Guided Attention for Video Action Anticipation
di: Tai, Tsung-Ming, et al.
Pubblicazione: (2026)
di: Tai, Tsung-Ming, et al.
Pubblicazione: (2026)
Utilizing dataset affinity prediction in object detection to assess training data
di: Becker, Stefan, et al.
Pubblicazione: (2023)
di: Becker, Stefan, et al.
Pubblicazione: (2023)
Online pre-training with long-form videos
di: Kato, Itsuki, et al.
Pubblicazione: (2024)
di: Kato, Itsuki, et al.
Pubblicazione: (2024)
Towards motion from video diffusion models
di: Janson, Paul, et al.
Pubblicazione: (2024)
di: Janson, Paul, et al.
Pubblicazione: (2024)
Shift and matching queries for video semantic segmentation
di: Mizuno, Tsubasa, et al.
Pubblicazione: (2024)
di: Mizuno, Tsubasa, et al.
Pubblicazione: (2024)
Deep video representation learning: a survey
di: Ravanbakhsh, Elham, et al.
Pubblicazione: (2024)
di: Ravanbakhsh, Elham, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Video Generation with Learned Action Prior
di: Sarkar, Meenakshi, et al.
Pubblicazione: (2024) -
Can 3D point cloud data improve automated body condition score prediction in dairy cattle?
di: Tang, Zhou, et al.
Pubblicazione: (2026) -
Meta Episodic learning with Dynamic Task Sampling for CLIP-based Point Cloud Classification
di: Ghose, Shuvozit, et al.
Pubblicazione: (2024) -
Koala: Key frame-conditioned long video-LLM
di: Tan, Reuben, et al.
Pubblicazione: (2024) -
Don't Pause! Every prediction matters in a streaming video
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2026)