Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Cokelek, Mert, Ozsoy, Halit, Imamoglu, Nevrez, Ozcinar, Cagri, Ayhan, Inci, Erdem, Erkut, Erdem, Aykut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
Attention-Guided Lidar Segmentation and Odometry Using Image-to-Point Cloud Saliency Transfer
by: Ding, Guanqun, et al.
Published: (2023)
by: Ding, Guanqun, et al.
Published: (2023)
Enhanced LULC Segmentation via Lightweight Model Refinements on ALOS-2 SAR Data
by: Caglayan, Ali, et al.
Published: (2026)
by: Caglayan, Ali, et al.
Published: (2026)
A Tutorial on ALOS2 SAR Utilization: Dataset Preparation, Self-Supervised Pretraining, and Semantic Segmentation
by: Imamoglu, Nevrez, et al.
Published: (2026)
by: Imamoglu, Nevrez, et al.
Published: (2026)
SAR-W-MixMAE: SAR Foundation Model Training Using Backscatter Power Weighting
by: Caglayan, Ali, et al.
Published: (2025)
by: Caglayan, Ali, et al.
Published: (2025)
TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
by: Kontostathis, Ioannis, et al.
Published: (2025)
by: Kontostathis, Ioannis, et al.
Published: (2025)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
A Mathematical Framework for AI Singularity: Conditions, Bounds, and Control of Recursive Improvement
by: Jafari, Akbar Anbar, et al.
Published: (2025)
by: Jafari, Akbar Anbar, et al.
Published: (2025)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
by: Hooshanfar, Kiana, et al.
Published: (2025)
by: Hooshanfar, Kiana, et al.
Published: (2025)
Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
by: Jafari, Akbar Anbar, et al.
Published: (2025)
by: Jafari, Akbar Anbar, et al.
Published: (2025)
OmniAudio: Generating Spatial Audio from 360-Degree Video
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
Object and Relation Centric Representations for Push Effect Prediction
by: Tekden, Ahmet E., et al.
Published: (2021)
by: Tekden, Ahmet E., et al.
Published: (2021)
Near-Infrared and Low-Rank Adaptation of Vision Transformers in Remote Sensing
by: Ulku, Irem, et al.
Published: (2024)
by: Ulku, Irem, et al.
Published: (2024)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
by: Gong, Litian, et al.
Published: (2025)
by: Gong, Litian, et al.
Published: (2025)
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2026)
by: Dogan, Mustafa, et al.
Published: (2026)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
by: Xiong, Junwen, et al.
Published: (2024)
by: Xiong, Junwen, et al.
Published: (2024)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
360DVO: Deep Visual Odometry for Monocular 360-Degree Camera
by: Guo, Xiaopeng, et al.
Published: (2026)
by: Guo, Xiaopeng, et al.
Published: (2026)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
Video Question Answering for People with Visual Impairments Using an Egocentric 360-Degree Camera
by: Song, Inpyo, et al.
Published: (2024)
by: Song, Inpyo, et al.
Published: (2024)
Motion-Plane-Adaptive Inter Prediction in 360-Degree Video Coding
by: Regensky, Andy, et al.
Published: (2022)
by: Regensky, Andy, et al.
Published: (2022)
SGFormer: Spherical Geometry Transformer for 360 Depth Estimation
by: Zhang, Junsong, et al.
Published: (2024)
by: Zhang, Junsong, et al.
Published: (2024)
Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM
by: Caglayan, Ali, et al.
Published: (2025)
by: Caglayan, Ali, et al.
Published: (2025)
Similar Items
-
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025) -
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026) -
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024) -
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023) -
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)