HyenaPixel: Global Image Context with Convolutions
Fuente:
arXiv
Saved in:
| Main Authors: | Spravil, Julian, Houben, Sebastian, Behnke, Sven |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
by: Spravil, Julian, et al.
Published: (2025)
by: Spravil, Julian, et al.
Published: (2025)
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
by: Quenzel, Jan, et al.
Published: (2024)
by: Quenzel, Jan, et al.
Published: (2024)
Iterative Motion Compensation for Canonical 3D Reconstruction from UAV Plant Images Captured in Windy Conditions
by: Rochow, Andre, et al.
Published: (2025)
by: Rochow, Andre, et al.
Published: (2025)
FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression Features
by: Rochow, Andre, et al.
Published: (2024)
by: Rochow, Andre, et al.
Published: (2024)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
PDCFNet: Enhancing Underwater Images through Pixel Difference Convolution
by: Zhang, Song, et al.
Published: (2024)
by: Zhang, Song, et al.
Published: (2024)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models
by: Cao, Helin, et al.
Published: (2024)
by: Cao, Helin, et al.
Published: (2024)
SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-Net
by: Cao, Helin, et al.
Published: (2024)
by: Cao, Helin, et al.
Published: (2024)
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
by: Cao, Helin, et al.
Published: (2025)
by: Cao, Helin, et al.
Published: (2025)
Marker-free Human Gait Analysis using a Smart Edge Sensor System
by: Bauer, Eva Katharina, et al.
Published: (2024)
by: Bauer, Eva Katharina, et al.
Published: (2024)
Learning from SAM: Harnessing a Foundation Model for Sim2Real Adaptation by Regularization
by: Bonani, Mayara E., et al.
Published: (2023)
by: Bonani, Mayara E., et al.
Published: (2023)
TextOCVP: Object-Centric Video Prediction with Language Guidance
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
by: Pätzold, Bastian, et al.
Published: (2025)
by: Pätzold, Bastian, et al.
Published: (2025)
Feature-Preserving Mesh Decimation for Normal Integration
by: Heep, Moritz, et al.
Published: (2025)
by: Heep, Moritz, et al.
Published: (2025)
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
by: Villar-Corrales, Angel, et al.
Published: (2024)
by: Villar-Corrales, Angel, et al.
Published: (2024)
Taming the Randomness: Towards Label-Preserving Cropping in Contrastive Learning
by: Hassan, Mohamed, et al.
Published: (2025)
by: Hassan, Mohamed, et al.
Published: (2025)
Towards Context-aware Convolutional Network for Image Restoration
by: Hao, Fangwei, et al.
Published: (2024)
by: Hao, Fangwei, et al.
Published: (2024)
Person Segmentation and Action Classification for Multi-Channel Hemisphere Field of View LiDAR Sensors
by: Seliunina, Svetlana, et al.
Published: (2024)
by: Seliunina, Svetlana, et al.
Published: (2024)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
by: Cao, Helin, et al.
Published: (2025)
by: Cao, Helin, et al.
Published: (2025)
Learning Embeddings with Centroid Triplet Loss for Object Identification in Robotic Grasping
by: Gouda, Anas, et al.
Published: (2024)
by: Gouda, Anas, et al.
Published: (2024)
Pyramid Pixel Context Adaption Network for Medical Image Classification with Supervised Contrastive Learning
by: Zhang, Xiaoqing, et al.
Published: (2023)
by: Zhang, Xiaoqing, et al.
Published: (2023)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation
by: Tutevych, Vitalii, et al.
Published: (2026)
by: Tutevych, Vitalii, et al.
Published: (2026)
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer
by: Elharrouss, Omar, et al.
Published: (2025)
by: Elharrouss, Omar, et al.
Published: (2025)
Pixel-wise Gradient Uncertainty for Convolutional Neural Networks applied to Out-of-Distribution Segmentation
by: Maag, Kira, et al.
Published: (2023)
by: Maag, Kira, et al.
Published: (2023)
PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization
by: Guichemerre, Alexis, et al.
Published: (2025)
by: Guichemerre, Alexis, et al.
Published: (2025)
Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation
by: Jiang, Wentao, et al.
Published: (2026)
by: Jiang, Wentao, et al.
Published: (2026)
Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
by: Hui, Tianrui, et al.
Published: (2023)
by: Hui, Tianrui, et al.
Published: (2023)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Inter-Image Pixel Shuffling for Multi-focus Image Fusion
by: Lin, Huangxing, et al.
Published: (2026)
by: Lin, Huangxing, et al.
Published: (2026)
Scale-Invariant Object Detection by Adaptive Convolution with Unified Global-Local Context
by: Singh, Amrita, et al.
Published: (2024)
by: Singh, Amrita, et al.
Published: (2024)
T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
by: Zhang, Shijie, et al.
Published: (2024)
by: Zhang, Shijie, et al.
Published: (2024)
PCIM: Learning Pixel Attributions via Pixel-wise Channel Isolation Mixing in High Content Imaging
by: Siegismund, Daniel, et al.
Published: (2024)
by: Siegismund, Daniel, et al.
Published: (2024)
Enhancing Video-Based Robot Failure Detection Using Task Knowledge
by: Thoduka, Santosh, et al.
Published: (2025)
by: Thoduka, Santosh, et al.
Published: (2025)
Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Images
by: Chen, Jiuchen, et al.
Published: (2025)
by: Chen, Jiuchen, et al.
Published: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
From Image- to Pixel-level: Label-efficient Hyperspectral Image Reconstruction
by: Leng, Yihong, et al.
Published: (2025)
by: Leng, Yihong, et al.
Published: (2025)
Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
by: Palit, Sanchar, et al.
Published: (2025)
by: Palit, Sanchar, et al.
Published: (2025)
Mapping Image Transformations Onto Pixel Processor Arrays
by: Bose, Laurie, et al.
Published: (2024)
by: Bose, Laurie, et al.
Published: (2024)
Similar Items
-
Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
by: Spravil, Julian, et al.
Published: (2025) -
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
by: Quenzel, Jan, et al.
Published: (2024) -
Iterative Motion Compensation for Canonical 3D Reconstruction from UAV Plant Images Captured in Windy Conditions
by: Rochow, Andre, et al.
Published: (2025) -
FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression Features
by: Rochow, Andre, et al.
Published: (2024) -
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)