VIOLA: Towards Video In-Context Learning with Minimal Annotations
Fuente:
arXiv
Saved in:
| Main Authors: | Fujii, Ryo, Saito, Hideo, Hachiuma, Ryo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025)
by: Fujii, Ryo, et al.
Published: (2025)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
by: Nagano, Koshiro, et al.
Published: (2026)
by: Nagano, Koshiro, et al.
Published: (2026)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
by: Ishikawa, Reina, et al.
Published: (2025)
by: Ishikawa, Reina, et al.
Published: (2025)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
EgoSurgery-HTS: A Dataset for Egocentric Hand-Tool Segmentation in Open Surgery Videos
by: Darjana, Nathan, et al.
Published: (2025)
by: Darjana, Nathan, et al.
Published: (2025)
EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Video CLIP Model for Multi-View Echocardiography Interpretation
by: Takizawa, Ryo, et al.
Published: (2025)
by: Takizawa, Ryo, et al.
Published: (2025)
Towards Temporal Change Explanations from Bi-Temporal Satellite Images
by: Tsujimoto, Ryo, et al.
Published: (2024)
by: Tsujimoto, Ryo, et al.
Published: (2024)
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
by: Zhou, Yutong, et al.
Published: (2024)
by: Zhou, Yutong, et al.
Published: (2024)
Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
by: Xiao, Zeqi, et al.
Published: (2025)
by: Xiao, Zeqi, et al.
Published: (2025)
Affinity-Graph-Guided Contractive Learning for Pretext-Free Medical Image Segmentation with Minimal Annotation
by: Cheng, Zehua, et al.
Published: (2024)
by: Cheng, Zehua, et al.
Published: (2024)
From Images to Insights: Explainable Biodiversity Monitoring with Plain Language Habitat Explanations
by: Zhou, Yutong, et al.
Published: (2025)
by: Zhou, Yutong, et al.
Published: (2025)
Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis
by: Li, Gary Y., et al.
Published: (2025)
by: Li, Gary Y., et al.
Published: (2025)
DeVAn: Dense Video Annotation for Video-Language Models
by: Liu, Tingkai, et al.
Published: (2023)
by: Liu, Tingkai, et al.
Published: (2023)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
by: Omasa, Takamitsu, et al.
Published: (2025)
by: Omasa, Takamitsu, et al.
Published: (2025)
Guess the Unified Model: How Much Can We Recover from Generated Images?
by: Cekinmez, Jasin, et al.
Published: (2026)
by: Cekinmez, Jasin, et al.
Published: (2026)
Examining Deployment and Refinement of the VIOLA-AI Intracranial Hemorrhage Model Using an Interactive NeoMedSys Platform
by: Liu, Qinghui, et al.
Published: (2025)
by: Liu, Qinghui, et al.
Published: (2025)
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
One Patient's Annotation is Another One's Initialization: Towards Zero-Shot Surgical Video Segmentation with Cross-Patient Initialization
by: Mousavi, Seyed Amir, et al.
Published: (2025)
by: Mousavi, Seyed Amir, et al.
Published: (2025)
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
by: Sun, Guolei, et al.
Published: (2022)
by: Sun, Guolei, et al.
Published: (2022)
Understanding Annotation Error Propagation and Learning an Adaptive Policy for Expert Intervention in Barrett's Video Segmentation
by: Rasanjalee, Lokesha, et al.
Published: (2026)
by: Rasanjalee, Lokesha, et al.
Published: (2026)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
by: Li, Yuanhang, et al.
Published: (2025)
by: Li, Yuanhang, et al.
Published: (2025)
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Disturbance-Free Surgical Video Generation from Multi-Camera Shadowless Lamps for Open Surgery
by: Kato, Yuna, et al.
Published: (2025)
by: Kato, Yuna, et al.
Published: (2025)
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
by: Muttaqien, Muhammad A., et al.
Published: (2025)
by: Muttaqien, Muhammad A., et al.
Published: (2025)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
by: Lin, Yijing, et al.
Published: (2025)
by: Lin, Yijing, et al.
Published: (2025)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025)
by: Mai, Ziyang, et al.
Published: (2025)
Context-Aware Temporal Embedding of Objects in Video Data
by: Farhan, Ahnaf, et al.
Published: (2024)
by: Farhan, Ahnaf, et al.
Published: (2024)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Similar Items
-
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025) -
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024) -
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024) -
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024) -
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
by: Nagano, Koshiro, et al.
Published: (2026)