GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yuxuan, Gao, Difei, Yu, Licheng, Lei, Stan Weixian, Feiszli, Matt, Shou, Mike Zheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Factorized Learning for Temporally Grounded Video-Language Models
di: Zeng, Wenzheng, et al.
Pubblicazione: (2025)
di: Zeng, Wenzheng, et al.
Pubblicazione: (2025)
ViT-Lens: Towards Omni-modal Representations
di: Lei, Weixian, et al.
Pubblicazione: (2023)
di: Lei, Weixian, et al.
Pubblicazione: (2023)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
di: Hu, Siyuan, et al.
Pubblicazione: (2024)
di: Hu, Siyuan, et al.
Pubblicazione: (2024)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
di: Lei, Weixian, et al.
Pubblicazione: (2023)
di: Lei, Weixian, et al.
Pubblicazione: (2023)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
D-AR: Diffusion via Autoregressive Models
di: Gao, Ziteng, et al.
Pubblicazione: (2025)
di: Gao, Ziteng, et al.
Pubblicazione: (2025)
ADen: Adaptive Density Representations for Sparse-view Camera Pose Estimation
di: Tang, Hao, et al.
Pubblicazione: (2024)
di: Tang, Hao, et al.
Pubblicazione: (2024)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2024)
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2024)
P-Flow: Prompting Visual Effects Generation
di: Zhao, Rui, et al.
Pubblicazione: (2026)
di: Zhao, Rui, et al.
Pubblicazione: (2026)
Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
di: Hu, Juan, et al.
Pubblicazione: (2024)
di: Hu, Juan, et al.
Pubblicazione: (2024)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
di: Zhang, David Junhao, et al.
Pubblicazione: (2023)
di: Zhang, David Junhao, et al.
Pubblicazione: (2023)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
di: Tang, Yolo Yunlong, et al.
Pubblicazione: (2023)
di: Tang, Yolo Yunlong, et al.
Pubblicazione: (2023)
ShowUI-Aloha: Human-Taught GUI Agent
di: Zhang, Yichun, et al.
Pubblicazione: (2026)
di: Zhang, Yichun, et al.
Pubblicazione: (2026)
Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
ROICtrl: Boosting Instance Control for Visual Generation
di: Gu, Yuchao, et al.
Pubblicazione: (2024)
di: Gu, Yuchao, et al.
Pubblicazione: (2024)
Parrot Captions Teach CLIP to Spot Text
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
di: Yang, Zhiwei, et al.
Pubblicazione: (2025)
di: Yang, Zhiwei, et al.
Pubblicazione: (2025)
Automated Movie Generation via Multi-Agent CoT Planning
di: Wu, Weijia, et al.
Pubblicazione: (2025)
di: Wu, Weijia, et al.
Pubblicazione: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
di: Zhao, Rui, et al.
Pubblicazione: (2025)
di: Zhao, Rui, et al.
Pubblicazione: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
di: Song, Yiren, et al.
Pubblicazione: (2025)
di: Song, Yiren, et al.
Pubblicazione: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
di: Ran, Lingmin, et al.
Pubblicazione: (2025)
di: Ran, Lingmin, et al.
Pubblicazione: (2025)
Ego-centric Predictive Model Conditioned on Hand Trajectories
di: Zhang, Binjie, et al.
Pubblicazione: (2025)
di: Zhang, Binjie, et al.
Pubblicazione: (2025)
Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
di: Zheng, Ziwei, et al.
Pubblicazione: (2024)
di: Zheng, Ziwei, et al.
Pubblicazione: (2024)
GUI Action Narrator: Where and When Did That Action Take Place?
di: Wu, Qinchen, et al.
Pubblicazione: (2024)
di: Wu, Qinchen, et al.
Pubblicazione: (2024)
Learning Video Context as Interleaved Multimodal Sequences
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
Factorized Visual Tokenization and Generation
di: Bai, Zechen, et al.
Pubblicazione: (2024)
di: Bai, Zechen, et al.
Pubblicazione: (2024)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
di: Bai, Zechen, et al.
Pubblicazione: (2025)
di: Bai, Zechen, et al.
Pubblicazione: (2025)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
di: Song, Quanjian, et al.
Pubblicazione: (2025)
di: Song, Quanjian, et al.
Pubblicazione: (2025)
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
di: Yang, Pei, et al.
Pubblicazione: (2026)
di: Yang, Pei, et al.
Pubblicazione: (2026)
VideoLLM-online: Online Video Large Language Model for Streaming Video
di: Chen, Joya, et al.
Pubblicazione: (2024)
di: Chen, Joya, et al.
Pubblicazione: (2024)
Mitty: Diffusion-based Human-to-Robot Video Generation
di: Song, Yiren, et al.
Pubblicazione: (2025)
di: Song, Yiren, et al.
Pubblicazione: (2025)
OmniPSD: Layered PSD Generation with Diffusion Transformer
di: Liu, Cheng, et al.
Pubblicazione: (2025)
di: Liu, Cheng, et al.
Pubblicazione: (2025)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
di: Song, Yiren, et al.
Pubblicazione: (2026)
di: Song, Yiren, et al.
Pubblicazione: (2026)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
di: Yang, Pei, et al.
Pubblicazione: (2025)
di: Yang, Pei, et al.
Pubblicazione: (2025)
ICON: Incremental CONfidence for Joint Pose and Radiance Field Optimization
di: Wang, Weiyao, et al.
Pubblicazione: (2024)
di: Wang, Weiyao, et al.
Pubblicazione: (2024)
3x2: 3D Object Part Segmentation by 2D Semantic Correspondences
di: Thai, Anh, et al.
Pubblicazione: (2024)
di: Thai, Anh, et al.
Pubblicazione: (2024)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
di: Mao, Weijia, et al.
Pubblicazione: (2025)
di: Mao, Weijia, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Factorized Learning for Temporally Grounded Video-Language Models
di: Zeng, Wenzheng, et al.
Pubblicazione: (2025) -
ViT-Lens: Towards Omni-modal Representations
di: Lei, Weixian, et al.
Pubblicazione: (2023) -
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
di: Hu, Siyuan, et al.
Pubblicazione: (2024) -
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025) -
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
di: Lei, Weixian, et al.
Pubblicazione: (2023)