Saved in:
| Main Authors: | Ponbagavathi, Thinesh Thiyakesan, Yang, Chengzheng, Roitberg, Alina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.07996 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Order Matters: On Parameter-Efficient Image-to-Video Probing for Recognizing Nearly Symmetric Actions
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
Probing Fine-Grained Action Understanding and Cross-View Generalization of Foundation Models
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2024)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2024)
Deep Learning for Metabolic Rate Estimation from Biosignals: A Comparative Study of Architectures and Signal Selection
by: Babakhani, Sarvenaz, et al.
Published: (2025)
by: Babakhani, Sarvenaz, et al.
Published: (2025)
Towards Synthetic Data Generation for Improved Pain Recognition in Videos under Patient Constraints
by: Nasimzada, Jonas, et al.
Published: (2024)
by: Nasimzada, Jonas, et al.
Published: (2024)
Towards Activated Muscle Group Estimation in the Wild
by: Peng, Kunyu, et al.
Published: (2023)
by: Peng, Kunyu, et al.
Published: (2023)
Exploring Few-Shot Adaptation for Activity Recognition on Diverse Domains
by: Peng, Kunyu, et al.
Published: (2023)
by: Peng, Kunyu, et al.
Published: (2023)
AdaptiveClick: Clicks-aware Transformer with Adaptive Focal Loss for Interactive Image Segmentation
by: Lin, Jiacheng, et al.
Published: (2023)
by: Lin, Jiacheng, et al.
Published: (2023)
Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
TransKD: Transformer Knowledge Distillation for Efficient Semantic Segmentation
by: Liu, Ruiping, et al.
Published: (2022)
by: Liu, Ruiping, et al.
Published: (2022)
Interactive Multi-Turn Retrieval for Health Videos
by: Wu, Chengzheng, et al.
Published: (2026)
by: Wu, Chengzheng, et al.
Published: (2026)
Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
by: Peng, Jihua, et al.
Published: (2025)
by: Peng, Jihua, et al.
Published: (2025)
Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic Segmentation
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
Time2General: Learning Spatiotemporal Invariant Representations for Domain-Generalization Video Semantic Segmentation
by: Chen, Siyu, et al.
Published: (2026)
by: Chen, Siyu, et al.
Published: (2026)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Black-box Attacks on Image Activity Prediction and its Natural Language Explanations
by: Baia, Alina Elena, et al.
Published: (2023)
by: Baia, Alina Elena, et al.
Published: (2023)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Skeleton-Based Human Action Recognition with Noisy Labels
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Navigating Hallucinations for Reasoning of Unintentional Activities
by: Grover, Shresth, et al.
Published: (2024)
by: Grover, Shresth, et al.
Published: (2024)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
PromptGAR: Flexible Promptive Group Activity Recognition
by: Jin, Zhangyu, et al.
Published: (2025)
by: Jin, Zhangyu, et al.
Published: (2025)
Pixels or Positions? Benchmarking Modalities in Group Activity Recognition
by: Karki, Drishya, et al.
Published: (2025)
by: Karki, Drishya, et al.
Published: (2025)
Group Relative Attention Guidance for Image Editing
by: Zhang, Xuanpu, et al.
Published: (2025)
by: Zhang, Xuanpu, et al.
Published: (2025)
Group Relative Policy Optimization for Image Captioning
by: Liang, Xu
Published: (2025)
by: Liang, Xu
Published: (2025)
Group-DINOmics: Incorporating People Dynamics into DINO for Self-supervised Group Activity Feature Learning
by: Tezuka, Ryuki, et al.
Published: (2026)
by: Tezuka, Ryuki, et al.
Published: (2026)
Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
Referring Atomic Video Action Recognition
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
Learning Group Activity Features Through Person Attribute Prediction
by: Nakatani, Chihiro, et al.
Published: (2024)
by: Nakatani, Chihiro, et al.
Published: (2024)
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
by: Meng, Shibei, et al.
Published: (2026)
by: Meng, Shibei, et al.
Published: (2026)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
by: Xu, Guoliang, et al.
Published: (2024)
by: Xu, Guoliang, et al.
Published: (2024)
Structure-Aware Correspondence Learning for Relative Pose Estimation
by: Chen, Yihan, et al.
Published: (2025)
by: Chen, Yihan, et al.
Published: (2025)
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
by: Wen, Di, et al.
Published: (2025)
by: Wen, Di, et al.
Published: (2025)
Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks
by: Yang, Zhichao, et al.
Published: (2026)
by: Yang, Zhichao, et al.
Published: (2026)
RelTopo: Multi-Level Relational Modeling for Driving Scene Topology Reasoning
by: Luo, Yueru, et al.
Published: (2025)
by: Luo, Yueru, et al.
Published: (2025)
SGA-INTERACT: A 3D Skeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic
by: Yang, Yuchen, et al.
Published: (2025)
by: Yang, Yuchen, et al.
Published: (2025)
Skeleton-based Group Activity Recognition via Spatial-Temporal Panoramic Graph
by: Li, Zhengcen, et al.
Published: (2024)
by: Li, Zhengcen, et al.
Published: (2024)
Towards More Practical Group Activity Detection: A New Benchmark and Model
by: Kim, Dongkeun, et al.
Published: (2023)
by: Kim, Dongkeun, et al.
Published: (2023)
Similar Items
-
Order Matters: On Parameter-Efficient Image-to-Video Probing for Recognizing Nearly Symmetric Actions
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025) -
T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025) -
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026) -
Probing Fine-Grained Action Understanding and Cross-View Generalization of Foundation Models
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2024) -
Deep Learning for Metabolic Rate Estimation from Biosignals: A Comparative Study of Architectures and Signal Selection
by: Babakhani, Sarvenaz, et al.
Published: (2025)