HyenaPixel: Global Image Context with Convolutions
Fuente:
arXiv
Salvato in:
| Autori principali: | Spravil, Julian, Houben, Sebastian, Behnke, Sven |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
di: Spravil, Julian, et al.
Pubblicazione: (2025)
di: Spravil, Julian, et al.
Pubblicazione: (2025)
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
di: Quenzel, Jan, et al.
Pubblicazione: (2024)
di: Quenzel, Jan, et al.
Pubblicazione: (2024)
Iterative Motion Compensation for Canonical 3D Reconstruction from UAV Plant Images Captured in Windy Conditions
di: Rochow, Andre, et al.
Pubblicazione: (2025)
di: Rochow, Andre, et al.
Pubblicazione: (2025)
FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression Features
di: Rochow, Andre, et al.
Pubblicazione: (2024)
di: Rochow, Andre, et al.
Pubblicazione: (2024)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
PDCFNet: Enhancing Underwater Images through Pixel Difference Convolution
di: Zhang, Song, et al.
Pubblicazione: (2024)
di: Zhang, Song, et al.
Pubblicazione: (2024)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
di: Wang, Yihao, et al.
Pubblicazione: (2025)
di: Wang, Yihao, et al.
Pubblicazione: (2025)
DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models
di: Cao, Helin, et al.
Pubblicazione: (2024)
di: Cao, Helin, et al.
Pubblicazione: (2024)
SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-Net
di: Cao, Helin, et al.
Pubblicazione: (2024)
di: Cao, Helin, et al.
Pubblicazione: (2024)
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
di: Cao, Helin, et al.
Pubblicazione: (2025)
di: Cao, Helin, et al.
Pubblicazione: (2025)
Marker-free Human Gait Analysis using a Smart Edge Sensor System
di: Bauer, Eva Katharina, et al.
Pubblicazione: (2024)
di: Bauer, Eva Katharina, et al.
Pubblicazione: (2024)
Learning from SAM: Harnessing a Foundation Model for Sim2Real Adaptation by Regularization
di: Bonani, Mayara E., et al.
Pubblicazione: (2023)
di: Bonani, Mayara E., et al.
Pubblicazione: (2023)
TextOCVP: Object-Centric Video Prediction with Language Guidance
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
di: Pätzold, Bastian, et al.
Pubblicazione: (2025)
di: Pätzold, Bastian, et al.
Pubblicazione: (2025)
Feature-Preserving Mesh Decimation for Normal Integration
di: Heep, Moritz, et al.
Pubblicazione: (2025)
di: Heep, Moritz, et al.
Pubblicazione: (2025)
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2024)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2024)
Taming the Randomness: Towards Label-Preserving Cropping in Contrastive Learning
di: Hassan, Mohamed, et al.
Pubblicazione: (2025)
di: Hassan, Mohamed, et al.
Pubblicazione: (2025)
Towards Context-aware Convolutional Network for Image Restoration
di: Hao, Fangwei, et al.
Pubblicazione: (2024)
di: Hao, Fangwei, et al.
Pubblicazione: (2024)
Person Segmentation and Action Classification for Multi-Channel Hemisphere Field of View LiDAR Sensors
di: Seliunina, Svetlana, et al.
Pubblicazione: (2024)
di: Seliunina, Svetlana, et al.
Pubblicazione: (2024)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
di: Cao, Helin, et al.
Pubblicazione: (2025)
di: Cao, Helin, et al.
Pubblicazione: (2025)
Learning Embeddings with Centroid Triplet Loss for Object Identification in Robotic Grasping
di: Gouda, Anas, et al.
Pubblicazione: (2024)
di: Gouda, Anas, et al.
Pubblicazione: (2024)
Pyramid Pixel Context Adaption Network for Medical Image Classification with Supervised Contrastive Learning
di: Zhang, Xiaoqing, et al.
Pubblicazione: (2023)
di: Zhang, Xiaoqing, et al.
Pubblicazione: (2023)
PixelDiT: Pixel Diffusion Transformers for Image Generation
di: Yu, Yongsheng, et al.
Pubblicazione: (2025)
di: Yu, Yongsheng, et al.
Pubblicazione: (2025)
Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation
di: Tutevych, Vitalii, et al.
Pubblicazione: (2026)
di: Tutevych, Vitalii, et al.
Pubblicazione: (2026)
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer
di: Elharrouss, Omar, et al.
Pubblicazione: (2025)
di: Elharrouss, Omar, et al.
Pubblicazione: (2025)
Pixel-wise Gradient Uncertainty for Convolutional Neural Networks applied to Out-of-Distribution Segmentation
di: Maag, Kira, et al.
Pubblicazione: (2023)
di: Maag, Kira, et al.
Pubblicazione: (2023)
PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization
di: Guichemerre, Alexis, et al.
Pubblicazione: (2025)
di: Guichemerre, Alexis, et al.
Pubblicazione: (2025)
Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation
di: Jiang, Wentao, et al.
Pubblicazione: (2026)
di: Jiang, Wentao, et al.
Pubblicazione: (2026)
Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
di: Hui, Tianrui, et al.
Pubblicazione: (2023)
di: Hui, Tianrui, et al.
Pubblicazione: (2023)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
di: Zhang, Lei, et al.
Pubblicazione: (2026)
di: Zhang, Lei, et al.
Pubblicazione: (2026)
Inter-Image Pixel Shuffling for Multi-focus Image Fusion
di: Lin, Huangxing, et al.
Pubblicazione: (2026)
di: Lin, Huangxing, et al.
Pubblicazione: (2026)
Scale-Invariant Object Detection by Adaptive Convolution with Unified Global-Local Context
di: Singh, Amrita, et al.
Pubblicazione: (2024)
di: Singh, Amrita, et al.
Pubblicazione: (2024)
T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
di: Zhang, Shijie, et al.
Pubblicazione: (2024)
di: Zhang, Shijie, et al.
Pubblicazione: (2024)
PCIM: Learning Pixel Attributions via Pixel-wise Channel Isolation Mixing in High Content Imaging
di: Siegismund, Daniel, et al.
Pubblicazione: (2024)
di: Siegismund, Daniel, et al.
Pubblicazione: (2024)
Enhancing Video-Based Robot Failure Detection Using Task Knowledge
di: Thoduka, Santosh, et al.
Pubblicazione: (2025)
di: Thoduka, Santosh, et al.
Pubblicazione: (2025)
Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Images
di: Chen, Jiuchen, et al.
Pubblicazione: (2025)
di: Chen, Jiuchen, et al.
Pubblicazione: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
From Image- to Pixel-level: Label-efficient Hyperspectral Image Reconstruction
di: Leng, Yihong, et al.
Pubblicazione: (2025)
di: Leng, Yihong, et al.
Pubblicazione: (2025)
Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
di: Palit, Sanchar, et al.
Pubblicazione: (2025)
di: Palit, Sanchar, et al.
Pubblicazione: (2025)
Mapping Image Transformations Onto Pixel Processor Arrays
di: Bose, Laurie, et al.
Pubblicazione: (2024)
di: Bose, Laurie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
di: Spravil, Julian, et al.
Pubblicazione: (2025) -
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
di: Quenzel, Jan, et al.
Pubblicazione: (2024) -
Iterative Motion Compensation for Canonical 3D Reconstruction from UAV Plant Images Captured in Windy Conditions
di: Rochow, Andre, et al.
Pubblicazione: (2025) -
FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression Features
di: Rochow, Andre, et al.
Pubblicazione: (2024) -
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)