Gespeichert in:
| Hauptverfasser: | Kolner, Oleh, Ortner, Thomas, Woźniak, Stanisław, Pantazi, Angeliki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.20213 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Event-based Optical Identification and Communication
von: von Arnim, Axel, et al.
Veröffentlicht: (2023)
von: von Arnim, Axel, et al.
Veröffentlicht: (2023)
FlowState: Sampling Rate Invariant Time Series Forecasting
von: Graf, Lars, et al.
Veröffentlicht: (2025)
von: Graf, Lars, et al.
Veröffentlicht: (2025)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
von: Pardyl, Adam, et al.
Veröffentlicht: (2024)
von: Pardyl, Adam, et al.
Veröffentlicht: (2024)
Unraveling the geometry of visual relational reasoning
von: Shang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Shang, Jiaqi, et al.
Veröffentlicht: (2025)
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
GAP: Gaussianize Any Point Clouds with Text Guidance
von: Zhang, Weiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Weiqi, et al.
Veröffentlicht: (2025)
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
von: Wu, Zhixuan, et al.
Veröffentlicht: (2026)
von: Wu, Zhixuan, et al.
Veröffentlicht: (2026)
Perception Encoder: The best visual embeddings are not at the output of the network
von: Bolya, Daniel, et al.
Veröffentlicht: (2025)
von: Bolya, Daniel, et al.
Veröffentlicht: (2025)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2025)
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2025)
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
Mind the Shape Gap: A Benchmark and Baseline for Deformation-Aware 6D Pose Estimation of Agricultural Produce
von: Chatzis, Nikolas, et al.
Veröffentlicht: (2026)
von: Chatzis, Nikolas, et al.
Veröffentlicht: (2026)
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Align the GAP: Prior-based Unified Multi-Task Remote Physiological Measurement Framework For Domain Generalization and Personalization
von: Wang, Jiyao, et al.
Veröffentlicht: (2025)
von: Wang, Jiyao, et al.
Veröffentlicht: (2025)
Active Visual Perception: Opportunities and Challenges
von: Li, Yian, et al.
Veröffentlicht: (2025)
von: Li, Yian, et al.
Veröffentlicht: (2025)
Incremental dimension reduction for efficient and accurate visual anomaly detection
von: Lee, Teng-Yok
Veröffentlicht: (2026)
von: Lee, Teng-Yok
Veröffentlicht: (2026)
Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D Glimpses
von: Lee, Inhee, et al.
Veröffentlicht: (2024)
von: Lee, Inhee, et al.
Veröffentlicht: (2024)
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
Affine transformation estimation improves visual self-supervised learning
von: Torpey, David, et al.
Veröffentlicht: (2024)
von: Torpey, David, et al.
Veröffentlicht: (2024)
Gradient events: improved acquisition of visual information in event cameras
von: Lehtonen, Eero, et al.
Veröffentlicht: (2024)
von: Lehtonen, Eero, et al.
Veröffentlicht: (2024)
UnGAP: Uncertainty-Guided Affine Prompting for Real-Time Crack Segmentation
von: Li, Conghui, et al.
Veröffentlicht: (2026)
von: Li, Conghui, et al.
Veröffentlicht: (2026)
Relaxed forced choice improves performance of visual quality assessment methods
von: Jenadeleh, Mohsen, et al.
Veröffentlicht: (2023)
von: Jenadeleh, Mohsen, et al.
Veröffentlicht: (2023)
BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
von: Khoche, Ajinkya, et al.
Veröffentlicht: (2025)
von: Khoche, Ajinkya, et al.
Veröffentlicht: (2025)
Learning Underwater Active Perception in Simulation
von: Cardaillac, Alexandre, et al.
Veröffentlicht: (2025)
von: Cardaillac, Alexandre, et al.
Veröffentlicht: (2025)
Glimpse: Generalized Locality for Scalable and Robust CT
von: Khorashadizadeh, AmirEhsan, et al.
Veröffentlicht: (2024)
von: Khorashadizadeh, AmirEhsan, et al.
Veröffentlicht: (2024)
VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
von: Kuzyk, Oleh, et al.
Veröffentlicht: (2025)
von: Kuzyk, Oleh, et al.
Veröffentlicht: (2025)
Active Perception Agent for Omnimodal Audio-Video Understanding
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
von: Jiang, Jianwen, et al.
Veröffentlicht: (2025)
von: Jiang, Jianwen, et al.
Veröffentlicht: (2025)
HeadGAP: Few-Shot 3D Head Avatar via Generalizable Gaussian Priors
von: Zheng, Xiaozheng, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaozheng, et al.
Veröffentlicht: (2024)
Mind the Hitch: Dynamic Calibration and Articulated Perception for Autonomous Trucks
von: Zhu, Morui, et al.
Veröffentlicht: (2026)
von: Zhu, Morui, et al.
Veröffentlicht: (2026)
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
von: Zheng, Weicheng, et al.
Veröffentlicht: (2025)
von: Zheng, Weicheng, et al.
Veröffentlicht: (2025)
Splatter Image: Ultra-Fast Single-View 3D Reconstruction
von: Szymanowicz, Stanislaw, et al.
Veröffentlicht: (2023)
von: Szymanowicz, Stanislaw, et al.
Veröffentlicht: (2023)
Through the PRISm: Importance-Aware Scene Graphs for Image Retrieval
von: Georgoulopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Georgoulopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
Novel View Synthesis from A Few Glimpses via Test-Time Natural Video Completion
von: Xu, Yan, et al.
Veröffentlicht: (2025)
von: Xu, Yan, et al.
Veröffentlicht: (2025)
Unmasking the Uniqueness: A Glimpse into Age-Invariant Face Recognition of Indigenous African Faces
von: Ajewole, Fakunle, et al.
Veröffentlicht: (2024)
von: Ajewole, Fakunle, et al.
Veröffentlicht: (2024)
ActFormer: Scalable Collaborative Perception via Active Queries
von: Huang, Suozhi, et al.
Veröffentlicht: (2024)
von: Huang, Suozhi, et al.
Veröffentlicht: (2024)
Audio-visual training for improved grounding in video-text LLMs
von: Sagare, Shivprasad, et al.
Veröffentlicht: (2024)
von: Sagare, Shivprasad, et al.
Veröffentlicht: (2024)
Fourier or Wavelet bases as counterpart self-attention in spikformer for efficient visual classification
von: Wang, Qingyu, et al.
Veröffentlicht: (2024)
von: Wang, Qingyu, et al.
Veröffentlicht: (2024)
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
von: Gkotsi, Polytimi Anna, et al.
Veröffentlicht: (2026)
von: Gkotsi, Polytimi Anna, et al.
Veröffentlicht: (2026)
Explaining Vision GNNs: A Semantic and Visual Analysis of Graph-based Image Classification
von: Chaidos, Nikolaos, et al.
Veröffentlicht: (2025)
von: Chaidos, Nikolaos, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dynamic Event-based Optical Identification and Communication
von: von Arnim, Axel, et al.
Veröffentlicht: (2023) -
FlowState: Sampling Rate Invariant Time Series Forecasting
von: Graf, Lars, et al.
Veröffentlicht: (2025) -
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
von: Pardyl, Adam, et al.
Veröffentlicht: (2024) -
Unraveling the geometry of visual relational reasoning
von: Shang, Jiaqi, et al.
Veröffentlicht: (2025) -
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)