GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Yuxuan, Gao, Difei, Yu, Licheng, Lei, Stan Weixian, Feiszli, Matt, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Publié: |
2022
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Factorized Learning for Temporally Grounded Video-Language Models
par: Zeng, Wenzheng, et autres
Publié: (2025)
par: Zeng, Wenzheng, et autres
Publié: (2025)
ViT-Lens: Towards Omni-modal Representations
par: Lei, Weixian, et autres
Publié: (2023)
par: Lei, Weixian, et autres
Publié: (2023)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
par: Hu, Siyuan, et autres
Publié: (2024)
par: Hu, Siyuan, et autres
Publié: (2024)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
par: Lin, Kevin Qinghong, et autres
Publié: (2025)
par: Lin, Kevin Qinghong, et autres
Publié: (2025)
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
par: Lei, Weixian, et autres
Publié: (2023)
par: Lei, Weixian, et autres
Publié: (2023)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
D-AR: Diffusion via Autoregressive Models
par: Gao, Ziteng, et autres
Publié: (2025)
par: Gao, Ziteng, et autres
Publié: (2025)
ADen: Adaptive Density Representations for Sparse-view Camera Pose Estimation
par: Tang, Hao, et autres
Publié: (2024)
par: Tang, Hao, et autres
Publié: (2024)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
par: Zhao, Henry Hengyuan, et autres
Publié: (2024)
par: Zhao, Henry Hengyuan, et autres
Publié: (2024)
P-Flow: Prompting Visual Effects Generation
par: Zhao, Rui, et autres
Publié: (2026)
par: Zhao, Rui, et autres
Publié: (2026)
Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
par: Hu, Juan, et autres
Publié: (2024)
par: Hu, Juan, et autres
Publié: (2024)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
par: Zhang, David Junhao, et autres
Publié: (2023)
par: Zhang, David Junhao, et autres
Publié: (2023)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
par: Tang, Yolo Yunlong, et autres
Publié: (2023)
par: Tang, Yolo Yunlong, et autres
Publié: (2023)
ShowUI-Aloha: Human-Taught GUI Agent
par: Zhang, Yichun, et autres
Publié: (2026)
par: Zhang, Yichun, et autres
Publié: (2026)
Grounded Video Caption Generation
par: Kazakos, Evangelos, et autres
Publié: (2024)
par: Kazakos, Evangelos, et autres
Publié: (2024)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
par: Vo, Dinh-Khoi, et autres
Publié: (2025)
ROICtrl: Boosting Instance Control for Visual Generation
par: Gu, Yuchao, et autres
Publié: (2024)
par: Gu, Yuchao, et autres
Publié: (2024)
Parrot Captions Teach CLIP to Spot Text
par: Lin, Yiqi, et autres
Publié: (2023)
par: Lin, Yiqi, et autres
Publié: (2023)
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
par: Yang, Zhiwei, et autres
Publié: (2025)
par: Yang, Zhiwei, et autres
Publié: (2025)
Automated Movie Generation via Multi-Agent CoT Planning
par: Wu, Weijia, et autres
Publié: (2025)
par: Wu, Weijia, et autres
Publié: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
par: Zhao, Rui, et autres
Publié: (2025)
par: Zhao, Rui, et autres
Publié: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
par: Song, Yiren, et autres
Publié: (2025)
par: Song, Yiren, et autres
Publié: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
par: Ran, Lingmin, et autres
Publié: (2025)
par: Ran, Lingmin, et autres
Publié: (2025)
Ego-centric Predictive Model Conditioned on Hand Trajectories
par: Zhang, Binjie, et autres
Publié: (2025)
par: Zhang, Binjie, et autres
Publié: (2025)
Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
par: Zheng, Ziwei, et autres
Publié: (2024)
par: Zheng, Ziwei, et autres
Publié: (2024)
GUI Action Narrator: Where and When Did That Action Take Place?
par: Wu, Qinchen, et autres
Publié: (2024)
par: Wu, Qinchen, et autres
Publié: (2024)
Learning Video Context as Interleaved Multimodal Sequences
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
Factorized Visual Tokenization and Generation
par: Bai, Zechen, et autres
Publié: (2024)
par: Bai, Zechen, et autres
Publié: (2024)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
par: Bai, Zechen, et autres
Publié: (2025)
par: Bai, Zechen, et autres
Publié: (2025)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
par: Song, Quanjian, et autres
Publié: (2025)
par: Song, Quanjian, et autres
Publié: (2025)
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
par: Yang, Pei, et autres
Publié: (2026)
par: Yang, Pei, et autres
Publié: (2026)
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
Mitty: Diffusion-based Human-to-Robot Video Generation
par: Song, Yiren, et autres
Publié: (2025)
par: Song, Yiren, et autres
Publié: (2025)
OmniPSD: Layered PSD Generation with Diffusion Transformer
par: Liu, Cheng, et autres
Publié: (2025)
par: Liu, Cheng, et autres
Publié: (2025)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
par: Song, Yiren, et autres
Publié: (2026)
par: Song, Yiren, et autres
Publié: (2026)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
par: Yang, Pei, et autres
Publié: (2025)
par: Yang, Pei, et autres
Publié: (2025)
ICON: Incremental CONfidence for Joint Pose and Radiance Field Optimization
par: Wang, Weiyao, et autres
Publié: (2024)
par: Wang, Weiyao, et autres
Publié: (2024)
3x2: 3D Object Part Segmentation by 2D Semantic Correspondences
par: Thai, Anh, et autres
Publié: (2024)
par: Thai, Anh, et autres
Publié: (2024)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
par: Mao, Weijia, et autres
Publié: (2025)
par: Mao, Weijia, et autres
Publié: (2025)
Documents similaires
-
Factorized Learning for Temporally Grounded Video-Language Models
par: Zeng, Wenzheng, et autres
Publié: (2025) -
ViT-Lens: Towards Omni-modal Representations
par: Lei, Weixian, et autres
Publié: (2023) -
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
par: Hu, Siyuan, et autres
Publié: (2024) -
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
par: Lin, Kevin Qinghong, et autres
Publié: (2025) -
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
par: Lei, Weixian, et autres
Publié: (2023)