Enregistré dans:
| Auteurs principaux: | Yu, Yating, Cao, Congqi, Wang, Zhaoying, Meng, Weihua, Li, Jie, Li, Yuxin, Wei, Zihao, Shen, Zhongpei, Zhang, Jiajun |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2511.00613 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
par: Yu, Yating, et autres
Publié: (2025)
par: Yu, Yating, et autres
Publié: (2025)
Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
par: Cao, Congqi, et autres
Publié: (2025)
par: Cao, Congqi, et autres
Publié: (2025)
Prototypical Learning Guided Context-Aware Segmentation Network for Few-Shot Anomaly Detection
par: Jiang, Yuxin, et autres
Publié: (2025)
par: Jiang, Yuxin, et autres
Publié: (2025)
SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning
par: Zhao, Zihao, et autres
Publié: (2026)
par: Zhao, Zihao, et autres
Publié: (2026)
Autoregressive Denoising Score Matching is a Good Video Anomaly Detector
par: Zhang, Hanwen, et autres
Publié: (2025)
par: Zhang, Hanwen, et autres
Publié: (2025)
SAW-Bench: Learning Situated Awareness in the Real World
par: Li, Chuhan, et autres
Publié: (2026)
par: Li, Chuhan, et autres
Publié: (2026)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
par: Li, Yifei, et autres
Publié: (2025)
par: Li, Yifei, et autres
Publié: (2025)
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP
par: Yu, Yating, et autres
Publié: (2024)
par: Yu, Yating, et autres
Publié: (2024)
Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
par: Cao, Congqi, et autres
Publié: (2024)
par: Cao, Congqi, et autres
Publié: (2024)
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
par: Wen, Zimo, et autres
Publié: (2026)
par: Wen, Zimo, et autres
Publié: (2026)
Multi-View Reconstruction with Global Context for 3D Anomaly Detection
par: Sun, Yihan, et autres
Publié: (2025)
par: Sun, Yihan, et autres
Publié: (2025)
Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method
par: Huang, Chao, et autres
Publié: (2026)
par: Huang, Chao, et autres
Publié: (2026)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
par: Wu, Dan, et autres
Publié: (2023)
par: Wu, Dan, et autres
Publié: (2023)
Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition
par: Cao, Congqi, et autres
Publié: (2025)
par: Cao, Congqi, et autres
Publié: (2025)
ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions
par: Liu, Zihao, et autres
Publié: (2026)
par: Liu, Zihao, et autres
Publié: (2026)
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
par: Li, Keyu, et autres
Publié: (2026)
par: Li, Keyu, et autres
Publié: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
par: Lin, Junming, et autres
Publié: (2024)
par: Lin, Junming, et autres
Publié: (2024)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
par: Dai, Zunkai, et autres
Publié: (2026)
par: Dai, Zunkai, et autres
Publié: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
par: Yang, Jie, et autres
Publié: (2026)
par: Yang, Jie, et autres
Publié: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
par: Hu, Lingxiang, et autres
Publié: (2026)
par: Hu, Lingxiang, et autres
Publié: (2026)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
par: Zhu, Hengchuan, et autres
Publié: (2025)
par: Zhu, Hengchuan, et autres
Publié: (2025)
Towards Real-World HDR Video Reconstruction: A Large-Scale Benchmark Dataset and A Two-Stage Alignment Network
par: Shu, Yong, et autres
Publié: (2024)
par: Shu, Yong, et autres
Publié: (2024)
Towards Aerial Collaborative Stereo: Real-Time Cross-Camera Feature Association and Relative Pose Estimation for UAVs
par: Wang, Zhaoying, et autres
Publié: (2024)
par: Wang, Zhaoying, et autres
Publié: (2024)
VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
par: Zhu, Liyun, et autres
Publié: (2025)
par: Zhu, Liyun, et autres
Publié: (2025)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
par: Zhu, Jie, et autres
Publié: (2026)
par: Zhu, Jie, et autres
Publié: (2026)
ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
par: Li, Xiaozhe, et autres
Publié: (2025)
par: Li, Xiaozhe, et autres
Publié: (2025)
Retina‐Like Neuromorphic Visual Sensor for Sensing Broad‐Spectrum Ultraviolet Light (Advanced Optical Materials 35/2024)
par: Zhaoying Xi, et autres
Publié: (2024)
par: Zhaoying Xi, et autres
Publié: (2024)
CauCLIP: Bridging the Sim-to-Real Gap in Surgical Video Understanding via Causality-Inspired Vision-Language Modeling
par: He, Yuxin, et autres
Publié: (2026)
par: He, Yuxin, et autres
Publié: (2026)
Hawk: Learning to Understand Open-World Video Anomalies
par: Tang, Jiaqi, et autres
Publié: (2024)
par: Tang, Jiaqi, et autres
Publié: (2024)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
par: Luo, Sha, et autres
Publié: (2026)
par: Luo, Sha, et autres
Publié: (2026)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
par: Li, Bozheng, et autres
Publié: (2025)
par: Li, Bozheng, et autres
Publié: (2025)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
par: Li, Jia, et autres
Publié: (2026)
par: Li, Jia, et autres
Publié: (2026)
Can LLMs Understand Time Series Anomalies?
par: Zhou, Zihao, et autres
Publié: (2024)
par: Zhou, Zihao, et autres
Publié: (2024)
Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?
par: Jiang, Hongbo, et autres
Publié: (2026)
par: Jiang, Hongbo, et autres
Publié: (2026)
Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting
par: Yao, Yueyang, et autres
Publié: (2025)
par: Yao, Yueyang, et autres
Publié: (2025)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
par: Liu, Zihao, et autres
Publié: (2025)
par: Liu, Zihao, et autres
Publié: (2025)
WorldModelBench: Judging Video Generation Models As World Models
par: Li, Dacheng, et autres
Publié: (2025)
par: Li, Dacheng, et autres
Publié: (2025)
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
par: Tang, Yuanbo, et autres
Publié: (2026)
par: Tang, Yuanbo, et autres
Publié: (2026)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
par: Xun, Shuhang, et autres
Publié: (2025)
par: Xun, Shuhang, et autres
Publié: (2025)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
par: Gao, Zeyu, et autres
Publié: (2025)
par: Gao, Zeyu, et autres
Publié: (2025)
Documents similaires
-
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
par: Yu, Yating, et autres
Publié: (2025) -
Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
par: Cao, Congqi, et autres
Publié: (2025) -
Prototypical Learning Guided Context-Aware Segmentation Network for Few-Shot Anomaly Detection
par: Jiang, Yuxin, et autres
Publié: (2025) -
SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning
par: Zhao, Zihao, et autres
Publié: (2026) -
Autoregressive Denoising Score Matching is a Good Video Anomaly Detector
par: Zhang, Hanwen, et autres
Publié: (2025)