EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiaoqi, Wang, Yi, Chau, Lap-Pui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
by: Wang, Xiaoqi, et al.
Published: (2024)
by: Wang, Xiaoqi, et al.
Published: (2024)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025)
by: Deng, Kunyuan, et al.
Published: (2025)
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024)
by: Su, Yuejiao, et al.
Published: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
by: Xu, Huaiyuan, et al.
Published: (2024)
by: Xu, Huaiyuan, et al.
Published: (2024)
HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
by: Hu, Qiongyang, et al.
Published: (2025)
by: Hu, Qiongyang, et al.
Published: (2025)
Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025)
by: Su, Yuejiao, et al.
Published: (2025)
Cross-Axis Transformer with 3D Rotary Positional Embeddings
by: Erickson, Lily
Published: (2023)
by: Erickson, Lily
Published: (2023)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
by: Liu, Tianyi, et al.
Published: (2025)
by: Liu, Tianyi, et al.
Published: (2025)
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
by: Ahuja, Rahul, et al.
Published: (2026)
by: Ahuja, Rahul, et al.
Published: (2026)
Weakly-supervised Part-Attention and Mentored Networks for Vehicle Re-Identification
by: Tang, Lisha, et al.
Published: (2021)
by: Tang, Lisha, et al.
Published: (2021)
Evolution-Inspired Sample Competition for Deep Neural Network Optimization
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)
by: Zhang, Erhang, et al.
Published: (2025)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025)
by: Seth, Ashish, et al.
Published: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
by: Li, Junlong, et al.
Published: (2025)
by: Li, Junlong, et al.
Published: (2025)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
by: Su, Yuejiao, et al.
Published: (2026)
by: Su, Yuejiao, et al.
Published: (2026)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024)
by: Zheng, Ying, et al.
Published: (2024)
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
ProCal: Probability Calibration for Neighborhood-Guided Source-Free Domain Adaptation
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
by: Yao, Lei, et al.
Published: (2024)
by: Yao, Lei, et al.
Published: (2024)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
Spatial-Temporal Deep Embedding for Vehicle Trajectory Reconstruction from High-Angle Video
by: D., Tianya T. Zhang Ph., et al.
Published: (2022)
by: D., Tianya T. Zhang Ph., et al.
Published: (2022)
MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
by: Zhu, Haoyu, et al.
Published: (2026)
by: Zhu, Haoyu, et al.
Published: (2026)
Egocentric Bias in Vision-Language Models
by: Wang, Maijunxian, et al.
Published: (2026)
by: Wang, Maijunxian, et al.
Published: (2026)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
by: Chen, Junliang, et al.
Published: (2025)
by: Chen, Junliang, et al.
Published: (2025)
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding
by: Luo, Bingjun, et al.
Published: (2026)
by: Luo, Bingjun, et al.
Published: (2026)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
by: Li, Shicheng, et al.
Published: (2023)
by: Li, Shicheng, et al.
Published: (2023)
Similar Items
-
Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
by: Wang, Xiaoqi, et al.
Published: (2024) -
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
by: Sun, Haoran, et al.
Published: (2025) -
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025) -
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
by: Su, Yuejiao, et al.
Published: (2024) -
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)