Action-conditioned video data improves predictability
Fuente:
arXiv
Guardado en:
| Autores principales: | Sarkar, Meenakshi, Ghose, Debasish |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Video Generation with Learned Action Prior
por: Sarkar, Meenakshi, et al.
Publicado: (2024)
por: Sarkar, Meenakshi, et al.
Publicado: (2024)
Can 3D point cloud data improve automated body condition score prediction in dairy cattle?
por: Tang, Zhou, et al.
Publicado: (2026)
por: Tang, Zhou, et al.
Publicado: (2026)
Meta Episodic learning with Dynamic Task Sampling for CLIP-based Point Cloud Classification
por: Ghose, Shuvozit, et al.
Publicado: (2024)
por: Ghose, Shuvozit, et al.
Publicado: (2024)
Koala: Key frame-conditioned long video-LLM
por: Tan, Reuben, et al.
Publicado: (2024)
por: Tan, Reuben, et al.
Publicado: (2024)
Don't Pause! Every prediction matters in a streaming video
por: Chatterjee, Dibyadip, et al.
Publicado: (2026)
por: Chatterjee, Dibyadip, et al.
Publicado: (2026)
A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters
por: Desai, Shasvat, et al.
Publicado: (2025)
por: Desai, Shasvat, et al.
Publicado: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
M3DDM+: An improved video outpainting by a modified masking strategy
por: Murakawa, Takuya, et al.
Publicado: (2026)
por: Murakawa, Takuya, et al.
Publicado: (2026)
Exploiting temporal information to detect conversational groups in videos and predict the next speaker
por: Tosato, Lucrezia, et al.
Publicado: (2024)
por: Tosato, Lucrezia, et al.
Publicado: (2024)
Autoregression-free video prediction using diffusion model for mitigating error propagation
por: Ko, Woonho, et al.
Publicado: (2025)
por: Ko, Woonho, et al.
Publicado: (2025)
AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance
por: Klein, Benjamin, et al.
Publicado: (2026)
por: Klein, Benjamin, et al.
Publicado: (2026)
CLIP-based Point Cloud Classification via Point Cloud to Image Translation
por: Ghose, Shuvozit, et al.
Publicado: (2024)
por: Ghose, Shuvozit, et al.
Publicado: (2024)
Bidirectional skip-frame prediction for video anomaly detection with intra-domain disparity-driven attention
por: Lyu, Jiahao, et al.
Publicado: (2024)
por: Lyu, Jiahao, et al.
Publicado: (2024)
Audio-visual training for improved grounding in video-text LLMs
por: Sagare, Shivprasad, et al.
Publicado: (2024)
por: Sagare, Shivprasad, et al.
Publicado: (2024)
Image Segmentation with transformers: An Overview, Challenges and Future
por: Chetia, Deepjyoti, et al.
Publicado: (2025)
por: Chetia, Deepjyoti, et al.
Publicado: (2025)
Robust soybean seed yield estimation using high-throughput ground robot videos
por: Feng, Jiale, et al.
Publicado: (2024)
por: Feng, Jiale, et al.
Publicado: (2024)
LiteVLA-H: Dual-Rate Vision-Language-Action Inference for Onboard Aerial Guidance and Semantic Perception
por: williams, Justin, et al.
Publicado: (2026)
por: williams, Justin, et al.
Publicado: (2026)
Foul prediction with estimated poses from soccer broadcast video
por: Fang, Jiale, et al.
Publicado: (2024)
por: Fang, Jiale, et al.
Publicado: (2024)
Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
por: Sun, Jian, et al.
Publicado: (2026)
por: Sun, Jian, et al.
Publicado: (2026)
EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-based Vision
por: Torbunov, Dmitrii, et al.
Publicado: (2024)
por: Torbunov, Dmitrii, et al.
Publicado: (2024)
3D Gaussian Splatting with Normal Information for Mesh Extraction and Improved Rendering
por: Krishnan, Meenakshi, et al.
Publicado: (2025)
por: Krishnan, Meenakshi, et al.
Publicado: (2025)
Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions
por: Hariat, Marwane, et al.
Publicado: (2026)
por: Hariat, Marwane, et al.
Publicado: (2026)
Test-time augmentation improves efficiency in conformal prediction
por: Shanmugam, Divya, et al.
Publicado: (2025)
por: Shanmugam, Divya, et al.
Publicado: (2025)
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
por: Xu, Kechun, et al.
Publicado: (2025)
por: Xu, Kechun, et al.
Publicado: (2025)
Action-guided generation of 3D functionality segmentation data
por: Corsetti, Jaime, et al.
Publicado: (2025)
por: Corsetti, Jaime, et al.
Publicado: (2025)
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks
por: Ghose, Partho, et al.
Publicado: (2026)
por: Ghose, Partho, et al.
Publicado: (2026)
Adaptive local boundary conditions to improve Deformable Image Registration
por: Inacio, Eloïse, et al.
Publicado: (2024)
por: Inacio, Eloïse, et al.
Publicado: (2024)
DLM-VMTL:A Double Layer Mapper for heterogeneous data video Multi-task prompt learning
por: Bo, Zeyi, et al.
Publicado: (2024)
por: Bo, Zeyi, et al.
Publicado: (2024)
Real time anomalies detection on video
por: Poirier, Fabien
Publicado: (2024)
por: Poirier, Fabien
Publicado: (2024)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
por: Venkataramanan, Shashanka, et al.
Publicado: (2023)
por: Venkataramanan, Shashanka, et al.
Publicado: (2023)
Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts
por: Chhipa, Prakash Chandra, et al.
Publicado: (2024)
por: Chhipa, Prakash Chandra, et al.
Publicado: (2024)
LCM: Log Conformal Maps for Robust Representation Learning to Mitigate Perspective Distortion
por: Chippa, Meenakshi Subhash, et al.
Publicado: (2024)
por: Chippa, Meenakshi Subhash, et al.
Publicado: (2024)
How good are deep learning methods for automated road safety analysis using video data? An experimental study
por: Liu, Qingwu, et al.
Publicado: (2025)
por: Liu, Qingwu, et al.
Publicado: (2025)
Video prediction using score-based conditional density estimation
por: Fiquet, Pierre-Étienne H., et al.
Publicado: (2024)
por: Fiquet, Pierre-Étienne H., et al.
Publicado: (2024)
Action-Guided Attention for Video Action Anticipation
por: Tai, Tsung-Ming, et al.
Publicado: (2026)
por: Tai, Tsung-Ming, et al.
Publicado: (2026)
Utilizing dataset affinity prediction in object detection to assess training data
por: Becker, Stefan, et al.
Publicado: (2023)
por: Becker, Stefan, et al.
Publicado: (2023)
Online pre-training with long-form videos
por: Kato, Itsuki, et al.
Publicado: (2024)
por: Kato, Itsuki, et al.
Publicado: (2024)
Towards motion from video diffusion models
por: Janson, Paul, et al.
Publicado: (2024)
por: Janson, Paul, et al.
Publicado: (2024)
Shift and matching queries for video semantic segmentation
por: Mizuno, Tsubasa, et al.
Publicado: (2024)
por: Mizuno, Tsubasa, et al.
Publicado: (2024)
Deep video representation learning: a survey
por: Ravanbakhsh, Elham, et al.
Publicado: (2024)
por: Ravanbakhsh, Elham, et al.
Publicado: (2024)
Ejemplares similares
-
Video Generation with Learned Action Prior
por: Sarkar, Meenakshi, et al.
Publicado: (2024) -
Can 3D point cloud data improve automated body condition score prediction in dairy cattle?
por: Tang, Zhou, et al.
Publicado: (2026) -
Meta Episodic learning with Dynamic Task Sampling for CLIP-based Point Cloud Classification
por: Ghose, Shuvozit, et al.
Publicado: (2024) -
Koala: Key frame-conditioned long video-LLM
por: Tan, Reuben, et al.
Publicado: (2024) -
Don't Pause! Every prediction matters in a streaming video
por: Chatterjee, Dibyadip, et al.
Publicado: (2026)