AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Patel, Urjitkumar, Yeh, Fang-Chun, Gondhalekar, Chinmay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
by: Gondhalekar, Chinmay, et al.
Published: (2025)
by: Gondhalekar, Chinmay, et al.
Published: (2025)
MultiFinRAG: An Optimized Multimodal Retrieval-Augmented Generation (RAG) Framework for Financial Question Answering
by: Gondhalekar, Chinmay, et al.
Published: (2025)
by: Gondhalekar, Chinmay, et al.
Published: (2025)
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
by: Sharma, Akul, et al.
Published: (2025)
by: Sharma, Akul, et al.
Published: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026)
by: Lentsch, Ted, et al.
Published: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
by: Lentsch, Ted, et al.
Published: (2024)
by: Lentsch, Ted, et al.
Published: (2024)
Enhancing Maritime Object Detection in Real-Time with RT-DETR and Data Augmentation
by: Nemati, Nader
Published: (2025)
by: Nemati, Nader
Published: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
by: Bell-Navas, Andrés, et al.
Published: (2025)
by: Bell-Navas, Andrés, et al.
Published: (2025)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention
by: Roffo, Giorgio, et al.
Published: (2026)
by: Roffo, Giorgio, et al.
Published: (2026)
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
by: Kushal, Koushik Ahmed, et al.
Published: (2025)
by: Kushal, Koushik Ahmed, et al.
Published: (2025)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
by: Patel, Urjitkumar, et al.
Published: (2024)
by: Patel, Urjitkumar, et al.
Published: (2024)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
by: Gad, Eyad, et al.
Published: (2025)
by: Gad, Eyad, et al.
Published: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
by: Svystun, Serhii, et al.
Published: (2025)
by: Svystun, Serhii, et al.
Published: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
by: Baydemir, Poyraz
Published: (2025)
by: Baydemir, Poyraz
Published: (2025)
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
by: Dong, Haohua, et al.
Published: (2025)
by: Dong, Haohua, et al.
Published: (2025)
FANAL -- Financial Activity News Alerting Language Modeling Framework
by: Patel, Urjitkumar, et al.
Published: (2024)
by: Patel, Urjitkumar, et al.
Published: (2024)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
by: Kazemi, Amir, et al.
Published: (2024)
by: Kazemi, Amir, et al.
Published: (2024)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)
by: Zinnen, Mathias, et al.
Published: (2025)
Transforming faces into video stories -- VideoFace2.0
by: Brkljač, Branko, et al.
Published: (2025)
by: Brkljač, Branko, et al.
Published: (2025)
CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
ROI-GS: Interest-based Local Quality 3D Gaussian Splatting
by: Bui, Quoc-Anh, et al.
Published: (2025)
by: Bui, Quoc-Anh, et al.
Published: (2025)
ROI-NeRFs: Hi-Fi Visualization of Objects of Interest within a Scene by NeRFs Composition
by: Bui, Quoc-Anh, et al.
Published: (2025)
by: Bui, Quoc-Anh, et al.
Published: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
by: Komurcu, Kursat, et al.
Published: (2026)
by: Komurcu, Kursat, et al.
Published: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
by: Kambhatla, Akhila, et al.
Published: (2025)
by: Kambhatla, Akhila, et al.
Published: (2025)
Single-Step Reconstruction-Free Anomaly Detection and Segmentation via Diffusion Models
by: Moradi, Mehrdad, et al.
Published: (2025)
by: Moradi, Mehrdad, et al.
Published: (2025)
Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
Image-based Facial Rig Inversion
by: Yang, Tianxiang, et al.
Published: (2025)
by: Yang, Tianxiang, et al.
Published: (2025)
MVTamperBench: Evaluating Robustness of Vision-Language Models
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
Data Augmentation and Resolution Enhancement using GANs and Diffusion Models for Tree Segmentation
by: Ferreira, Alessandro dos Santos, et al.
Published: (2025)
by: Ferreira, Alessandro dos Santos, et al.
Published: (2025)
QSilk: Micrograin Stabilization and Adaptive Quantile Clipping for Detail-Friendly Latent Diffusion
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
by: Kang, Xueyang, et al.
Published: (2026)
by: Kang, Xueyang, et al.
Published: (2026)
Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
by: Meziani, Yani
Published: (2026)
by: Meziani, Yani
Published: (2026)
Machine learning technique for morphological classification of galaxies from SDSS. IV. Visual inspection vs CNN for merging, irregular, edge-on, barred, ringed, and with dust lanes galaxies at 0.02<z<0.1
by: V., Dobrycheva D., et al.
Published: (2026)
by: V., Dobrycheva D., et al.
Published: (2026)
Similar Items
-
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
by: Gondhalekar, Chinmay, et al.
Published: (2025) -
MultiFinRAG: An Optimized Multimodal Retrieval-Augmented Generation (RAG) Framework for Financial Question Answering
by: Gondhalekar, Chinmay, et al.
Published: (2025) -
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
by: Sharma, Akul, et al.
Published: (2025) -
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026) -
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026)