AsyReC: A Multimodal Graph-based Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tang, Wang, Dogan, Fethiye Irmak, Qing, Linbo, Gunes, Hatice |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic Interactions
par: Luo, Cheng, et autres
Publié: (2023)
par: Luo, Cheng, et autres
Publié: (2023)
Does SpatioTemporal information benefit Two video summarization benchmarks?
par: Ganesh, Aashutosh, et autres
Publié: (2024)
par: Ganesh, Aashutosh, et autres
Publié: (2024)
Reimagining Social Robots as Recommender Systems: Foundations, Framework, and Applications
par: Huang, Jin, et autres
Publié: (2026)
par: Huang, Jin, et autres
Publié: (2026)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
par: Ma, Jingtian, et autres
Publié: (2025)
par: Ma, Jingtian, et autres
Publié: (2025)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
par: Meng, Jiahao, et autres
Publié: (2026)
par: Meng, Jiahao, et autres
Publié: (2026)
REMAC: Reference-Based Martian Asymmetrical Image Compression
par: Ding, Qing, et autres
Publié: (2026)
par: Ding, Qing, et autres
Publié: (2026)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
par: Pramanick, Shraman, et autres
Publié: (2025)
par: Pramanick, Shraman, et autres
Publié: (2025)
AGSP-DSA: An Adaptive Graph Signal Processing Framework for Robust Multimodal Fusion with Dynamic Semantic Alignment
par: Karthikeya, KV, et autres
Publié: (2026)
par: Karthikeya, KV, et autres
Publié: (2026)
Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling
par: Li, Xinyu, et autres
Publié: (2025)
par: Li, Xinyu, et autres
Publié: (2025)
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
par: Wang, Tianyu, et autres
Publié: (2026)
par: Wang, Tianyu, et autres
Publié: (2026)
Designing Social Robots with Ethical, User-Adaptive Explainability in the Era of Foundation Models
par: Dogan, Fethiye Irmak, et autres
Publié: (2026)
par: Dogan, Fethiye Irmak, et autres
Publié: (2026)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
par: Meng, Jiahao, et autres
Publié: (2025)
par: Meng, Jiahao, et autres
Publié: (2025)
Diagnosing and Re-learning for Balanced Multimodal Learning
par: Wei, Yake, et autres
Publié: (2024)
par: Wei, Yake, et autres
Publié: (2024)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
par: Wang, Kangsheng, et autres
Publié: (2025)
par: Wang, Kangsheng, et autres
Publié: (2025)
Boosting Temporal Sentence Grounding via Causal Inference
par: Tang, Kefan, et autres
Publié: (2025)
par: Tang, Kefan, et autres
Publié: (2025)
Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
par: Liu, Xinhang, et autres
Publié: (2024)
par: Liu, Xinhang, et autres
Publié: (2024)
WDMIR: Wavelet-Driven Multimodal Intent Recognition
par: Gong, Weiyin, et autres
Publié: (2025)
par: Gong, Weiyin, et autres
Publié: (2025)
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
par: Tang, Hao, et autres
Publié: (2025)
par: Tang, Hao, et autres
Publié: (2025)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
par: Zhao, Yumiao, et autres
Publié: (2024)
par: Zhao, Yumiao, et autres
Publié: (2024)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
par: Moradi, Morteza, et autres
Publié: (2024)
par: Moradi, Morteza, et autres
Publié: (2024)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
par: Chen, Yaru, et autres
Publié: (2025)
par: Chen, Yaru, et autres
Publié: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
par: Wang, Xiao, et autres
Publié: (2024)
par: Wang, Xiao, et autres
Publié: (2024)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
par: Sun, Hao, et autres
Publié: (2024)
par: Sun, Hao, et autres
Publié: (2024)
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
par: Yin, Qilin, et autres
Publié: (2025)
par: Yin, Qilin, et autres
Publié: (2025)
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
par: Yan, Jinbo, et autres
Publié: (2024)
par: Yan, Jinbo, et autres
Publié: (2024)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
par: Cheng, Hongye, et autres
Publié: (2025)
par: Cheng, Hongye, et autres
Publié: (2025)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
par: Wang, Zihua, et autres
Publié: (2025)
par: Wang, Zihua, et autres
Publié: (2025)
DDNet: A Dual-Stream Graph Learning and Disentanglement Framework for Temporal Forgery Localization
par: Zhao, Boyang, et autres
Publié: (2026)
par: Zhao, Boyang, et autres
Publié: (2026)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
par: Ma, Xueqi, et autres
Publié: (2025)
par: Ma, Xueqi, et autres
Publié: (2025)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
par: Muqeet, Abdul, et autres
Publié: (2023)
par: Muqeet, Abdul, et autres
Publié: (2023)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
par: Zhang, Zijian, et autres
Publié: (2025)
par: Zhang, Zijian, et autres
Publié: (2025)
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
par: Diao, Xingjian, et autres
Publié: (2025)
par: Diao, Xingjian, et autres
Publié: (2025)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
par: Qin, Yang, et autres
Publié: (2023)
par: Qin, Yang, et autres
Publié: (2023)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
par: Guan, Jiazhi, et autres
Publié: (2024)
par: Guan, Jiazhi, et autres
Publié: (2024)
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
par: Shu, Ying, et autres
Publié: (2026)
par: Shu, Ying, et autres
Publié: (2026)
Robust Duality Learning for Unsupervised Visible-Infrared Person Re-Identification
par: Li, Yongxiang, et autres
Publié: (2025)
par: Li, Yongxiang, et autres
Publié: (2025)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
par: Huang, Dawei, et autres
Publié: (2025)
par: Huang, Dawei, et autres
Publié: (2025)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
par: Hao, Jing, et autres
Publié: (2025)
par: Hao, Jing, et autres
Publié: (2025)
GANonymization: A GAN-based Face Anonymization Framework for Preserving Emotional Expressions
par: Hellmann, Fabio, et autres
Publié: (2023)
par: Hellmann, Fabio, et autres
Publié: (2023)
Spatial-Temporal Human-Object Interaction Detection
par: Sun, Xu, et autres
Publié: (2025)
par: Sun, Xu, et autres
Publié: (2025)
Documents similaires
-
ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic Interactions
par: Luo, Cheng, et autres
Publié: (2023) -
Does SpatioTemporal information benefit Two video summarization benchmarks?
par: Ganesh, Aashutosh, et autres
Publié: (2024) -
Reimagining Social Robots as Recommender Systems: Foundations, Framework, and Applications
par: Huang, Jin, et autres
Publié: (2026) -
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
par: Ma, Jingtian, et autres
Publié: (2025) -
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
par: Meng, Jiahao, et autres
Publié: (2026)