On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pang, Zhanzhong, Chatterjee, Dibyadip, Sener, Fadime, Yao, Angela |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
par: Pang, Zhanzhong, et autres
Publié: (2026)
par: Pang, Zhanzhong, et autres
Publié: (2026)
Don't Pause! Every prediction matters in a streaming video
par: Chatterjee, Dibyadip, et autres
Publié: (2026)
par: Chatterjee, Dibyadip, et autres
Publié: (2026)
Context-Enhanced Memory-Refined Transformer for Online Action Detection
par: Pang, Zhanzhong, et autres
Publié: (2025)
par: Pang, Zhanzhong, et autres
Publié: (2025)
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
par: Pang, Zhanzhong, et autres
Publié: (2025)
par: Pang, Zhanzhong, et autres
Publié: (2025)
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
par: Pang, Zhanzhong, et autres
Publié: (2024)
par: Pang, Zhanzhong, et autres
Publié: (2024)
On the Utility of 3D Hand Poses for Action Recognition
par: Shamil, Md Salman, et autres
Publié: (2024)
par: Shamil, Md Salman, et autres
Publié: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
par: Kukleva, Anna, et autres
Publié: (2024)
par: Kukleva, Anna, et autres
Publié: (2024)
Rethinking Domain Generalization: Discriminability and Generalizability
par: Long, Shaocong, et autres
Publié: (2023)
par: Long, Shaocong, et autres
Publié: (2023)
SneakPeek: Future-Guided Instructional Streaming Video Generation
par: Hong, Cheeun, et autres
Publié: (2025)
par: Hong, Cheeun, et autres
Publié: (2025)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
par: Zhu, Xiaorong, et autres
Publié: (2025)
par: Zhu, Xiaorong, et autres
Publié: (2025)
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
par: Fan, Zicong, et autres
Publié: (2025)
par: Fan, Zicong, et autres
Publié: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
par: Lin, Jingli, et autres
Publié: (2025)
par: Lin, Jingli, et autres
Publié: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
par: Zhang, Gengyuan, et autres
Publié: (2025)
par: Zhang, Gengyuan, et autres
Publié: (2025)
Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
par: Si, Guangzong, et autres
Publié: (2025)
par: Si, Guangzong, et autres
Publié: (2025)
Condensing Action Segmentation Datasets via Generative Network Inversion
par: Ding, Guodong, et autres
Publié: (2025)
par: Ding, Guodong, et autres
Publié: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
par: He, Yuping, et autres
Publié: (2025)
par: He, Yuping, et autres
Publié: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
par: Qiu, Han, et autres
Publié: (2025)
par: Qiu, Han, et autres
Publié: (2025)
Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception
par: Pang, Ziqi, et autres
Publié: (2025)
par: Pang, Ziqi, et autres
Publié: (2025)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
par: Gao, Minghe, et autres
Publié: (2024)
par: Gao, Minghe, et autres
Publié: (2024)
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
par: Wang, Ziyi, et autres
Publié: (2026)
par: Wang, Ziyi, et autres
Publié: (2026)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
par: Chang, Chun-Peng, et autres
Publié: (2024)
par: Chang, Chun-Peng, et autres
Publié: (2024)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
par: Lv, Qi, et autres
Publié: (2025)
par: Lv, Qi, et autres
Publié: (2025)
GAIA: Rethinking Action Quality Assessment for AI-Generated Videos
par: Chen, Zijian, et autres
Publié: (2024)
par: Chen, Zijian, et autres
Publié: (2024)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
par: Sun, Peiwen, et autres
Publié: (2026)
par: Sun, Peiwen, et autres
Publié: (2026)
ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
par: Yang, Liu, et autres
Publié: (2025)
par: Yang, Liu, et autres
Publié: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
par: Sun, Yanpeng, et autres
Publié: (2025)
par: Sun, Yanpeng, et autres
Publié: (2025)
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
par: Li, Bao, et autres
Publié: (2025)
par: Li, Bao, et autres
Publié: (2025)
Coherent Temporal Synthesis for Incremental Action Segmentation
par: Ding, Guodong, et autres
Publié: (2024)
par: Ding, Guodong, et autres
Publié: (2024)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
par: Zhao, Qiyan, et autres
Publié: (2026)
par: Zhao, Qiyan, et autres
Publié: (2026)
OnlineTAS: An Online Baseline for Temporal Action Segmentation
par: Zhong, Qing, et autres
Publié: (2024)
par: Zhong, Qing, et autres
Publié: (2024)
Bridging Degradation Discrimination and Generation for Universal Image Restoration
par: Hu, JiaKui, et autres
Publié: (2026)
par: Hu, JiaKui, et autres
Publié: (2026)
Rethinking Prior Information Generation with CLIP for Few-Shot Segmentation
par: Wang, Jin, et autres
Publié: (2024)
par: Wang, Jin, et autres
Publié: (2024)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
par: Liu, Xianjie, et autres
Publié: (2026)
par: Liu, Xianjie, et autres
Publié: (2026)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
par: Zhang, Shan, et autres
Publié: (2025)
par: Zhang, Shan, et autres
Publié: (2025)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
par: Li, Yun, et autres
Publié: (2025)
par: Li, Yun, et autres
Publié: (2025)
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
par: Wang, Yonghui, et autres
Publié: (2024)
par: Wang, Yonghui, et autres
Publié: (2024)
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
par: Xing, Jiazheng, et autres
Publié: (2026)
par: Xing, Jiazheng, et autres
Publié: (2026)
Rethinking generalization of classifiers in separable classes scenarios and over-parameterized regimes
par: Martinetz, Julius, et autres
Publié: (2024)
par: Martinetz, Julius, et autres
Publié: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
par: Lu, Yujie, et autres
Publié: (2024)
par: Lu, Yujie, et autres
Publié: (2024)
Documents similaires
-
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
par: Pang, Zhanzhong, et autres
Publié: (2026) -
Don't Pause! Every prediction matters in a streaming video
par: Chatterjee, Dibyadip, et autres
Publié: (2026) -
Context-Enhanced Memory-Refined Transformer for Online Action Detection
par: Pang, Zhanzhong, et autres
Publié: (2025) -
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
par: Pang, Zhanzhong, et autres
Publié: (2025) -
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
par: Pang, Zhanzhong, et autres
Publié: (2024)