Saved in:
| Main Authors: | Hanania, Eyal, Kirsch, Nadav, Arkushin, Daniel, Benvenisti, Jonathan, Bercovich, Amos, Zemmour, Elie, Froim, Sahar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.25409 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
by: Hyun, Lee, et al.
Published: (2023)
by: Hyun, Lee, et al.
Published: (2023)
Interactive Multimodal Fusion with Temporal Modeling
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
by: Jung-Mok, Lee, et al.
Published: (2026)
by: Jung-Mok, Lee, et al.
Published: (2026)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
by: Zhao, Fuzheng, et al.
Published: (2024)
by: Zhao, Fuzheng, et al.
Published: (2024)
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
by: Roccabruna, Gabriel, et al.
Published: (2026)
by: Roccabruna, Gabriel, et al.
Published: (2026)
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
by: wang, Shuo, et al.
Published: (2025)
by: wang, Shuo, et al.
Published: (2025)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
by: Sun, Li, et al.
Published: (2024)
by: Sun, Li, et al.
Published: (2024)
Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning
by: Xu, Wenbo, et al.
Published: (2025)
by: Xu, Wenbo, et al.
Published: (2025)
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
by: Anshul, Ashutosh, et al.
Published: (2025)
by: Anshul, Ashutosh, et al.
Published: (2025)
Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
On-Demand Millisecond Storage of Spectro-Temporal Multimode Telecom Photons
by: Sethia, Anuj, et al.
Published: (2025)
by: Sethia, Anuj, et al.
Published: (2025)
Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments
by: Ugai, Takanori, et al.
Published: (2024)
by: Ugai, Takanori, et al.
Published: (2024)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
by: Zhou, Li, et al.
Published: (2025)
by: Zhou, Li, et al.
Published: (2025)
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
by: Chen, Jialin, et al.
Published: (2025)
by: Chen, Jialin, et al.
Published: (2025)
Mesh Based Simulations with Spatial and Temporal awareness
by: Garnier, Paul, et al.
Published: (2026)
by: Garnier, Paul, et al.
Published: (2026)
Cross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning
by: Khorasani, Arian, et al.
Published: (2026)
by: Khorasani, Arian, et al.
Published: (2026)
Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling
by: Choudhary, Yash, et al.
Published: (2025)
by: Choudhary, Yash, et al.
Published: (2025)
A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization
by: Xu, Wenbo, et al.
Published: (2025)
by: Xu, Wenbo, et al.
Published: (2025)
Adaptive Temporal Fusion Transformers for Cryptocurrency Price Prediction
by: Peik, Arash, et al.
Published: (2025)
by: Peik, Arash, et al.
Published: (2025)
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
by: Bercovich, Ivan
Published: (2026)
by: Bercovich, Ivan
Published: (2026)
SNN-Driven Multimodal Human Action Recognition via Sparse Spatial-Temporal Data Fusion
by: Zheng, Naichuan, et al.
Published: (2025)
by: Zheng, Naichuan, et al.
Published: (2025)
Attention-Based Multiscale Temporal Fusion Network for Uncertain-Mode Fault Diagnosis in Multimode Processes
by: Li, Guangqiang, et al.
Published: (2025)
by: Li, Guangqiang, et al.
Published: (2025)
STQuant: Spatio-Temporal Adaptive Framework for Optimizer Quantization in Large Multimodal Model Training
by: Liu, Minglu, et al.
Published: (2026)
by: Liu, Minglu, et al.
Published: (2026)
Temporal Interest-Driven Multimodal Personalized Content Generation
by: Miao, Tian
Published: (2025)
by: Miao, Tian
Published: (2025)
Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens
by: Sun, Dengdi, et al.
Published: (2024)
by: Sun, Dengdi, et al.
Published: (2024)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
Quantum Temporal Fusion Transformer
by: Barik, Krishnakanta, et al.
Published: (2025)
by: Barik, Krishnakanta, et al.
Published: (2025)
Discovering Temporally-Aware Reinforcement Learning Algorithms
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
Localized Dynamic Mode Decomposition with Temporally Adaptive Segmentation
by: Li, Qiuqi, et al.
Published: (2025)
by: Li, Qiuqi, et al.
Published: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
MELON: Multimodal Mixture-of-Experts with Spectral-Temporal Fusion for Long-Term Mobility Estimation in Critical Care
by: Zhang, Jiaqing, et al.
Published: (2025)
by: Zhang, Jiaqing, et al.
Published: (2025)
Detecting Bone Lesions in X-Ray Under Diverse Acquisition Conditions
by: Zimbalist, Tal, et al.
Published: (2022)
by: Zimbalist, Tal, et al.
Published: (2022)
Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision
by: Cao, Yi, et al.
Published: (2023)
by: Cao, Yi, et al.
Published: (2023)
An Adaptive Solid‐State Synapse with Bi‐Directional Relaxation for Multimodal Recognition and Spatio‐Temporal Learning
by: Fang Nie, et al.
Published: (2025)
by: Fang Nie, et al.
Published: (2025)
Efficient and Flexible Multirate Temporal Adaptivity
by: Reynolds, Daniel R., et al.
Published: (2025)
by: Reynolds, Daniel R., et al.
Published: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Similar Items
-
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
by: Hyun, Lee, et al.
Published: (2023) -
Interactive Multimodal Fusion with Temporal Modeling
by: Yu, Jun, et al.
Published: (2025) -
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
by: Cai, Mu, et al.
Published: (2024) -
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
by: Jung-Mok, Lee, et al.
Published: (2026) -
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
by: Zhao, Fuzheng, et al.
Published: (2024)