Dynamic Group Detection using VLM-augmented Temporal Groupness Graph
Fuente:
arXiv
Saved in:
| Main Authors: | Yokoyama, Kaname, Nakatani, Chihiro, Ukita, Norimichi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Group-DINOmics: Incorporating People Dynamics into DINO for Self-supervised Group Activity Feature Learning
by: Tezuka, Ryuki, et al.
Published: (2026)
by: Tezuka, Ryuki, et al.
Published: (2026)
Learning Group Activity Features Through Person Attribute Prediction
by: Nakatani, Chihiro, et al.
Published: (2024)
by: Nakatani, Chihiro, et al.
Published: (2024)
End-to-End Shared Attention Estimation via Group Detection with Feedback Refinement
by: Nakatani, Chihiro, et al.
Published: (2026)
by: Nakatani, Chihiro, et al.
Published: (2026)
Human-in-the-loop Adaptation in Group Activity Feature Learning for Team Sports Video Retrieval
by: Nakatani, Chihiro, et al.
Published: (2026)
by: Nakatani, Chihiro, et al.
Published: (2026)
Size-Variable Virtual Try-On with Physical Clothes Size
by: Yamashita, Yohei, et al.
Published: (2024)
by: Yamashita, Yohei, et al.
Published: (2024)
Time-series Initialization and Conditioning for Video-agnostic Stabilization of Video Super-Resolution using Recurrent Networks
by: Mori, Hiroshi, et al.
Published: (2024)
by: Mori, Hiroshi, et al.
Published: (2024)
Data-Driven Stochastic Motion Evaluation and Optimization with Image by Spatially-Aligned Temporal Encoding
by: Oba, Takeru, et al.
Published: (2023)
by: Oba, Takeru, et al.
Published: (2023)
Efficient Cost-and-Quality Controllable Arbitrary-scale Super-resolution with Fourier Constraints
by: Akita, Kazutoshi, et al.
Published: (2025)
by: Akita, Kazutoshi, et al.
Published: (2025)
Inpainting-Driven Mask Optimization for Object Removal
by: Shimosato, Kodai, et al.
Published: (2024)
by: Shimosato, Kodai, et al.
Published: (2024)
SAMIDARE: Advanced Tracking-by-Segmentation for Dense Scenarios
by: Hirano, Shozaburo, et al.
Published: (2026)
by: Hirano, Shozaburo, et al.
Published: (2026)
MMCM: Multimodality-aware Metric using Clustering-based Modes for Probabilistic Human Motion Prediction
by: Tokoro, Kyotaro, et al.
Published: (2025)
by: Tokoro, Kyotaro, et al.
Published: (2025)
Test-time Cost-and-Quality Controllable Arbitrary-Scale Super-Resolution with Variable Fourier Components
by: Akita, Kazutoshi, et al.
Published: (2024)
by: Akita, Kazutoshi, et al.
Published: (2024)
Human Motion Prediction via Test-domain-aware Adaptation with Easily-available Human Motions Estimated from Videos
by: Shimbo, Katsuki, et al.
Published: (2025)
by: Shimbo, Katsuki, et al.
Published: (2025)
Depth Estimation fusing Image and Radar Measurements with Uncertain Directions
by: Kotani, Masaya, et al.
Published: (2024)
by: Kotani, Masaya, et al.
Published: (2024)
Multi-Person Pose Estimation Evaluation Using Optimal Transportation and Improved Pose Matching
by: Moriki, Takato, et al.
Published: (2026)
by: Moriki, Takato, et al.
Published: (2026)
Burst Super-Resolution with Diffusion Models for Improving Perceptual Quality
by: Tokoro, Kyotaro, et al.
Published: (2024)
by: Tokoro, Kyotaro, et al.
Published: (2024)
Joint Learning of Blind Super-Resolution and Crack Segmentation for Realistic Degraded Images
by: Kondo, Yuki, et al.
Published: (2023)
by: Kondo, Yuki, et al.
Published: (2023)
Selective Social-Interaction via Individual Importance for Fast Human Trajectory Prediction
by: Urano, Yota, et al.
Published: (2025)
by: Urano, Yota, et al.
Published: (2025)
CacheFlow: Fast Human Motion Prediction by Cached Normalizing Flow
by: Maeda, Takahiro, et al.
Published: (2025)
by: Maeda, Takahiro, et al.
Published: (2025)
Multimodal Active Measurement for Human Mesh Recovery in Close Proximity
by: Maeda, Takahiro, et al.
Published: (2023)
by: Maeda, Takahiro, et al.
Published: (2023)
Efficient Burst Super-Resolution with One-step Diffusion
by: Kawai, Kento, et al.
Published: (2025)
by: Kawai, Kento, et al.
Published: (2025)
Physical Plausibility-aware Trajectory Prediction via Locomotion Embodiment
by: Taketsugu, Hiromu, et al.
Published: (2025)
by: Taketsugu, Hiromu, et al.
Published: (2025)
NTIRE 2023 Image Shadow Removal Challenge Technical Report: Team IIM_TTI
by: Kondo, Yuki, et al.
Published: (2024)
by: Kondo, Yuki, et al.
Published: (2024)
Unified and Dynamic Graph for Temporal Character Grouping in Long Videos
by: Shu, Xiujun, et al.
Published: (2023)
by: Shu, Xiujun, et al.
Published: (2023)
VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
by: Shen, Wenhao, et al.
Published: (2026)
by: Shen, Wenhao, et al.
Published: (2026)
Skeleton-based Group Activity Recognition via Spatial-Temporal Panoramic Graph
by: Li, Zhengcen, et al.
Published: (2024)
by: Li, Zhengcen, et al.
Published: (2024)
Random Walk on Pixel Manifolds for Anomaly Segmentation of Complex Driving Scenes
by: Zeng, Zelong, et al.
Published: (2024)
by: Zeng, Zelong, et al.
Published: (2024)
GGMotion: Group Graph Dynamics-Kinematics Networks for Human Motion Prediction
by: Wan, Shuaijin
Published: (2025)
by: Wan, Shuaijin
Published: (2025)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
EgoGroups: A Benchmark For Detecting Social Groups of People in the Wild
by: Murrugarra-Llerena, Jeffri, et al.
Published: (2026)
by: Murrugarra-Llerena, Jeffri, et al.
Published: (2026)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024)
by: Pang, Zhanzhong, et al.
Published: (2024)
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
by: Huang, Mingzhe, et al.
Published: (2026)
by: Huang, Mingzhe, et al.
Published: (2026)
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models
by: Makarov, Vladislav, et al.
Published: (2026)
by: Makarov, Vladislav, et al.
Published: (2026)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
by: Kim, Youbin, et al.
Published: (2026)
by: Kim, Youbin, et al.
Published: (2026)
CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
Axis-level Symmetry Detection with Group-Equivariant Representation
by: Yu, Wongyun, et al.
Published: (2025)
by: Yu, Wongyun, et al.
Published: (2025)
GroupFace: Imbalanced Age Estimation Based on Multi-hop Attention Graph Convolutional Network and Group-aware Margin Optimization
by: Zhang, Yiping, et al.
Published: (2024)
by: Zhang, Yiping, et al.
Published: (2024)
Similar Items
-
Group-DINOmics: Incorporating People Dynamics into DINO for Self-supervised Group Activity Feature Learning
by: Tezuka, Ryuki, et al.
Published: (2026) -
Learning Group Activity Features Through Person Attribute Prediction
by: Nakatani, Chihiro, et al.
Published: (2024) -
End-to-End Shared Attention Estimation via Group Detection with Feedback Refinement
by: Nakatani, Chihiro, et al.
Published: (2026) -
Human-in-the-loop Adaptation in Group Activity Feature Learning for Team Sports Video Retrieval
by: Nakatani, Chihiro, et al.
Published: (2026) -
Size-Variable Virtual Try-On with Physical Clothes Size
by: Yamashita, Yohei, et al.
Published: (2024)