A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
Fuente:
arXiv
Saved in:
| Main Authors: | Karanfil, Enes, Imamoglu, Nevrez, Erdem, Erkut, Erdem, Aykut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
A Tutorial on ALOS2 SAR Utilization: Dataset Preparation, Self-Supervised Pretraining, and Semantic Segmentation
by: Imamoglu, Nevrez, et al.
Published: (2026)
by: Imamoglu, Nevrez, et al.
Published: (2026)
Enhanced LULC Segmentation via Lightweight Model Refinements on ALOS-2 SAR Data
by: Caglayan, Ali, et al.
Published: (2026)
by: Caglayan, Ali, et al.
Published: (2026)
SAR-W-MixMAE: SAR Foundation Model Training Using Backscatter Power Weighting
by: Caglayan, Ali, et al.
Published: (2025)
by: Caglayan, Ali, et al.
Published: (2025)
Attention-Guided Lidar Segmentation and Odometry Using Image-to-Point Cloud Saliency Transfer
by: Ding, Guanqun, et al.
Published: (2023)
by: Ding, Guanqun, et al.
Published: (2023)
Robust Multispectral Semantic Segmentation under Missing or Full Modalities via Structured Latent Projection
by: Ulku, Irem, et al.
Published: (2026)
by: Ulku, Irem, et al.
Published: (2026)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
Scene Change Detection with Vision-Language Representation Learning
by: Sheng, Diwei, et al.
Published: (2026)
by: Sheng, Diwei, et al.
Published: (2026)
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
by: Wu, Guile, et al.
Published: (2026)
by: Wu, Guile, et al.
Published: (2026)
Infrared Domain Adaptation with Zero-Shot Quantization
by: Sevsay, Burak, et al.
Published: (2024)
by: Sevsay, Burak, et al.
Published: (2024)
How to Augment for Atmospheric Turbulence Effects on Thermal Adapted Object Detection Models?
by: Uzun, Engin, et al.
Published: (2024)
by: Uzun, Engin, et al.
Published: (2024)
Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM
by: Caglayan, Ali, et al.
Published: (2025)
by: Caglayan, Ali, et al.
Published: (2025)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Near-Infrared and Low-Rank Adaptation of Vision Transformers in Remote Sensing
by: Ulku, Irem, et al.
Published: (2024)
by: Ulku, Irem, et al.
Published: (2024)
Dynamic Scene Understanding from Vision-Language Representations
by: Pruss, Shahaf, et al.
Published: (2025)
by: Pruss, Shahaf, et al.
Published: (2025)
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
by: Ling, Lu, et al.
Published: (2025)
by: Ling, Lu, et al.
Published: (2025)
Beyond the Visible: Multispectral Vision-Language Learning for Earth Observation
by: Marimo, Clive Tinashe, et al.
Published: (2025)
by: Marimo, Clive Tinashe, et al.
Published: (2025)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
by: Huang, Jiangyong, et al.
Published: (2025)
by: Huang, Jiangyong, et al.
Published: (2025)
Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
TransAdapter: Vision Transformer for Feature-Centric Unsupervised Domain Adaptation
by: Doruk, A. Enes, et al.
Published: (2024)
by: Doruk, A. Enes, et al.
Published: (2024)
Deep Learning-Based Real-Time Sequential Facial Expression Analysis Using Geometric Features
by: Koksal, Talha Enes, et al.
Published: (2025)
by: Koksal, Talha Enes, et al.
Published: (2025)
ACE-LoRA: Graph-Attentive Context Enhancement for Parameter-Efficient Adaptation of Medical Vision-Language Models
by: Aydın, M. Arda, et al.
Published: (2026)
by: Aydın, M. Arda, et al.
Published: (2026)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
by: Liu, Yuanyuan, et al.
Published: (2025)
by: Liu, Yuanyuan, et al.
Published: (2025)
Φeat: Physically-Grounded Feature Representation
by: Vecchio, Giuseppe, et al.
Published: (2025)
by: Vecchio, Giuseppe, et al.
Published: (2025)
Similar Items
-
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025) -
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025) -
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024) -
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026) -
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)