CurEvo: Curriculum-Guided Self-Evolution for Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Guiyi, Yu, Junqing, Chen, Yi-Ping Phoebe, Chen, Xu, Yang, Wei, Song, Zikai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction
by: Ye, Liliang, et al.
Published: (2026)
by: Ye, Liliang, et al.
Published: (2026)
Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
by: Song, Zikai, et al.
Published: (2026)
by: Song, Zikai, et al.
Published: (2026)
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
by: Hu, Yangliu, et al.
Published: (2025)
by: Hu, Yangliu, et al.
Published: (2025)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
Autogenic Language Embedding for Coherent Point Tracking
by: Song, Zikai, et al.
Published: (2024)
by: Song, Zikai, et al.
Published: (2024)
GateMOT: Q-Gated Attention for Dense Object Tracking
by: Lv, Mingjin, et al.
Published: (2026)
by: Lv, Mingjin, et al.
Published: (2026)
Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Optimized View and Geometry Distillation from Multi-view Diffuser
by: Zhang, Youjia, et al.
Published: (2023)
by: Zhang, Youjia, et al.
Published: (2023)
Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
by: Li, Weiqing, et al.
Published: (2026)
by: Li, Weiqing, et al.
Published: (2026)
Ref-GS: Directional Factorization for 2D Gaussian Splatting
by: Zhang, Youjia, et al.
Published: (2024)
by: Zhang, Youjia, et al.
Published: (2024)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
Cross-Modality Masked Learning for Survival Prediction in ICI Treated NSCLC Patients
by: Xing, Qilong, et al.
Published: (2025)
by: Xing, Qilong, et al.
Published: (2025)
Video-Zero: Self-Evolution Video Understanding
by: Zhang, Ruixu, et al.
Published: (2026)
by: Zhang, Ruixu, et al.
Published: (2026)
MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation
by: Xing, Qilong, et al.
Published: (2025)
by: Xing, Qilong, et al.
Published: (2025)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
by: Jung, Minjoon, et al.
Published: (2026)
by: Jung, Minjoon, et al.
Published: (2026)
CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding
by: Jeon, Yongjun, et al.
Published: (2026)
by: Jeon, Yongjun, et al.
Published: (2026)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Progressive Text-to-Image Diffusion with Soft Latent Direction
by: Ye, YuTeng, et al.
Published: (2023)
by: Ye, YuTeng, et al.
Published: (2023)
TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing
by: Xu, Teng, et al.
Published: (2024)
by: Xu, Teng, et al.
Published: (2024)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation
by: Xing, Qilong, et al.
Published: (2025)
by: Xing, Qilong, et al.
Published: (2025)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
IP-MOT: Instance Prompt Learning for Cross-Domain Multi-Object Tracking
by: Luo, Run, et al.
Published: (2024)
by: Luo, Run, et al.
Published: (2024)
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
by: Wang, Xiangfeng, et al.
Published: (2025)
by: Wang, Xiangfeng, et al.
Published: (2025)
EvoVLA: Self-Evolving Vision-Language-Action Model
by: Liu, Zeting, et al.
Published: (2025)
by: Liu, Zeting, et al.
Published: (2025)
Generative Prior-Guided Neural Interface Reconstruction for 3D Electrical Impedance Tomography
by: Liu, Haibo, et al.
Published: (2025)
by: Liu, Haibo, et al.
Published: (2025)
Med-Evo: Test-time Self-evolution for Medical Multimodal Large Language Models
by: Xu, Dunyuan, et al.
Published: (2026)
by: Xu, Dunyuan, et al.
Published: (2026)
EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding
by: Huang, Muye, et al.
Published: (2024)
by: Huang, Muye, et al.
Published: (2024)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
by: Chen, Xiuwei, et al.
Published: (2025)
by: Chen, Xiuwei, et al.
Published: (2025)
PointSmile: Point Self-supervised Learning via Curriculum Mutual Information
by: Li, Xin, et al.
Published: (2023)
by: Li, Xin, et al.
Published: (2023)
Towards Long Video Understanding via Fine-detailed Video Story Generation
by: You, Zeng, et al.
Published: (2024)
by: You, Zeng, et al.
Published: (2024)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
by: Xie, Ming, et al.
Published: (2026)
by: Xie, Ming, et al.
Published: (2026)
Prototype-Guided Curriculum Learning for Zero-Shot Learning
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
by: Yin, Yufei, et al.
Published: (2026)
by: Yin, Yufei, et al.
Published: (2026)
DiffusionTrack: Diffusion Model For Multi-Object Tracking
by: Luo, Run, et al.
Published: (2023)
by: Luo, Run, et al.
Published: (2023)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
by: Ma, Wentao, et al.
Published: (2025)
by: Ma, Wentao, et al.
Published: (2025)
Similar Items
-
OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction
by: Ye, Liliang, et al.
Published: (2026) -
Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
by: Song, Zikai, et al.
Published: (2026) -
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
by: Hu, Yangliu, et al.
Published: (2025) -
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025) -
Autogenic Language Embedding for Coherent Point Tracking
by: Song, Zikai, et al.
Published: (2024)