MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Changho, Kim, Minho, Kim, Jinkyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SemanticControl: A Training-Free Approach for Handling Loosely Aligned Visual Conditions in ControlNet
von: Joung, Woosung, et al.
Veröffentlicht: (2025)
von: Joung, Woosung, et al.
Veröffentlicht: (2025)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models
von: Chae, Daewon, et al.
Veröffentlicht: (2025)
von: Chae, Daewon, et al.
Veröffentlicht: (2025)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025)
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025)
Addressing Diverging Training Costs using BEVRestore for High-resolution Bird's Eye View Map Construction
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
GreenEye: Development of Real-Time Traffic Signal Recognition System for Visual Impairments
von: Kim, Danu
Veröffentlicht: (2024)
von: Kim, Danu
Veröffentlicht: (2024)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
von: Chae, Daewon, et al.
Veröffentlicht: (2023)
von: Chae, Daewon, et al.
Veröffentlicht: (2023)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
Clustering-based Image-Text Graph Matching for Domain Generalization
von: Park, Nokyung, et al.
Veröffentlicht: (2023)
von: Park, Nokyung, et al.
Veröffentlicht: (2023)
FPANet: Frequency-based Video Demoireing using Frame-level Post Alignment
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2023)
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2023)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
von: Kim, Daeyoung
Veröffentlicht: (2026)
von: Kim, Daeyoung
Veröffentlicht: (2026)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding
von: Lu, Jingyu, et al.
Veröffentlicht: (2025)
von: Lu, Jingyu, et al.
Veröffentlicht: (2025)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
Finetuning Pre-trained Model with Limited Data for LiDAR-based 3D Object Detection by Bridging Domain Gaps
von: Jang, Jiyun, et al.
Veröffentlicht: (2024)
von: Jang, Jiyun, et al.
Veröffentlicht: (2024)
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2025)
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2025)
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
von: Kim, Jihyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jihyeon, et al.
Veröffentlicht: (2026)
Spanning Tree Autoregressive Visual Generation
von: Lee, Sangkyu, et al.
Veröffentlicht: (2025)
von: Lee, Sangkyu, et al.
Veröffentlicht: (2025)
Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball Rotation
von: Choi, YoungChan, et al.
Veröffentlicht: (2025)
von: Choi, YoungChan, et al.
Veröffentlicht: (2025)
LatentSwap: An Efficient Latent Code Mapping Framework for Face Swapping
von: Choi, Changho, et al.
Veröffentlicht: (2024)
von: Choi, Changho, et al.
Veröffentlicht: (2024)
Image-Guided Semantic Pseudo-LiDAR Point Generation for 3D Object Detection
von: Lee, Minseung, et al.
Veröffentlicht: (2024)
von: Lee, Minseung, et al.
Veröffentlicht: (2024)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
von: Lee, Jaeseong, et al.
Veröffentlicht: (2024)
von: Lee, Jaeseong, et al.
Veröffentlicht: (2024)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
von: Kwon, Mincheol, et al.
Veröffentlicht: (2026)
von: Kwon, Mincheol, et al.
Veröffentlicht: (2026)
Selective Visual Prompting in Vision Mamba
von: Yao, Yifeng, et al.
Veröffentlicht: (2024)
von: Yao, Yifeng, et al.
Veröffentlicht: (2024)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
Rethinking Visual Information Processing in Multimodal LLMs
von: Kim, Dongwan, et al.
Veröffentlicht: (2025)
von: Kim, Dongwan, et al.
Veröffentlicht: (2025)
SEDEG:Sequential Enhancement of Decoder and Encoder's Generality for Class Incremental Learning with Small Memory
von: Chen, Hongyang, et al.
Veröffentlicht: (2025)
von: Chen, Hongyang, et al.
Veröffentlicht: (2025)
VideoPrism: A Foundational Visual Encoder for Video Understanding
von: Zhao, Long, et al.
Veröffentlicht: (2024)
von: Zhao, Long, et al.
Veröffentlicht: (2024)
LD-Pruner: Efficient Pruning of Latent Diffusion Models using Task-Agnostic Insights
von: Castells, Thibault, et al.
Veröffentlicht: (2024)
von: Castells, Thibault, et al.
Veröffentlicht: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
Patch-wise Auto-Encoder for Visual Anomaly Detection
von: Cui, Yajie, et al.
Veröffentlicht: (2023)
von: Cui, Yajie, et al.
Veröffentlicht: (2023)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
von: Park, Jiho, et al.
Veröffentlicht: (2025)
von: Park, Jiho, et al.
Veröffentlicht: (2025)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
von: Back, Kyungryul, et al.
Veröffentlicht: (2025)
von: Back, Kyungryul, et al.
Veröffentlicht: (2025)
MoSSDA: A Semi-Supervised Domain Adaptation Framework for Multivariate Time-Series Classification using Momentum Encoder
von: Kim, Seonyoung, et al.
Veröffentlicht: (2025)
von: Kim, Seonyoung, et al.
Veröffentlicht: (2025)
MATHENA: Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy
von: Kim, Kyeonghun, et al.
Veröffentlicht: (2026)
von: Kim, Kyeonghun, et al.
Veröffentlicht: (2026)
Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
von: Choi, Jae Young, et al.
Veröffentlicht: (2026)
von: Choi, Jae Young, et al.
Veröffentlicht: (2026)
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
von: Kim, Daeyoung
Veröffentlicht: (2025)
von: Kim, Daeyoung
Veröffentlicht: (2025)
CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder
von: Dao, Duy-Phuong, et al.
Veröffentlicht: (2026)
von: Dao, Duy-Phuong, et al.
Veröffentlicht: (2026)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SemanticControl: A Training-Free Approach for Handling Loosely Aligned Visual Conditions in ControlNet
von: Joung, Woosung, et al.
Veröffentlicht: (2025) -
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
von: Park, Minjeong, et al.
Veröffentlicht: (2025) -
DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models
von: Chae, Daewon, et al.
Veröffentlicht: (2025) -
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025) -
Addressing Diverging Training Costs using BEVRestore for High-resolution Bird's Eye View Map Construction
von: Kim, Minsu, et al.
Veröffentlicht: (2024)