REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Khosla, Savya, TV, Sethuraman, Lee, Barnett, Schwing, Alexander, Hoiem, Derek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026)
by: Khosla, Savya, et al.
Published: (2026)
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024)
by: Khosla, Savya, et al.
Published: (2024)
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025)
by: TV, Sethuraman, et al.
Published: (2025)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024)
by: Wu, Yuqun, et al.
Published: (2024)
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
by: T V, Sethuraman, et al.
Published: (2026)
by: T V, Sethuraman, et al.
Published: (2026)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
Visual Program Distillation with Template-Based Augmentation
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
Anytime Continual Learning for Open Vocabulary Classification
by: Zhu, Zhen, et al.
Published: (2024)
by: Zhu, Zhen, et al.
Published: (2024)
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems
by: Zhu, Zhen, et al.
Published: (2023)
by: Zhu, Zhen, et al.
Published: (2023)
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
by: Lee, Jae Yong, et al.
Published: (2024)
by: Lee, Jae Yong, et al.
Published: (2024)
The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based Generation
by: Cheng, Ho Kei, et al.
Published: (2025)
by: Cheng, Ho Kei, et al.
Published: (2025)
Patch-enhanced Mask Encoder Prompt Image Generation
by: Xu, Shusong, et al.
Published: (2024)
by: Xu, Shusong, et al.
Published: (2024)
SimpliHuMoN: Simplifying Human Motion Prediction
by: Agrawal, Aadya, et al.
Published: (2026)
by: Agrawal, Aadya, et al.
Published: (2026)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2024)
by: Wen, Jing, et al.
Published: (2024)
REN: Anatomically-Informed Mixture-of-Experts for Interstitial Lung Disease Diagnosis
by: Peltekian, Alec K., et al.
Published: (2025)
by: Peltekian, Alec K., et al.
Published: (2025)
Patch-Level Glioblastoma Subregion Classification with a Contrastive Learning-Based Encoder
by: Zhang, Juexin, et al.
Published: (2025)
by: Zhang, Juexin, et al.
Published: (2025)
LIFe-GoM: Generalizable Human Rendering with Learned Iterative Feedback Over Multi-Resolution Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
RegionDrag: Fast Region-Based Image Editing with Diffusion Models
by: Lu, Jingyi, et al.
Published: (2024)
by: Lu, Jingyi, et al.
Published: (2024)
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
by: Takmaz, Ece, et al.
Published: (2025)
by: Takmaz, Ece, et al.
Published: (2025)
WISE-FUSE: Efficient Whole Slide Image Encoding via Coarse-to-Fine Patch Selection with VLM and LLM Knowledge Fusion
by: Shin, Yonghan, et al.
Published: (2025)
by: Shin, Yonghan, et al.
Published: (2025)
Variational Rectified Flow Matching
by: Guo, Pengsheng, et al.
Published: (2025)
by: Guo, Pengsheng, et al.
Published: (2025)
Efficient Image Synthesis with Sphere Latent Encoder
by: Do, Tung, et al.
Published: (2026)
by: Do, Tung, et al.
Published: (2026)
PatchCraft: Exploring Texture Patch for Efficient AI-generated Image Detection
by: Zhong, Nan, et al.
Published: (2023)
by: Zhong, Nan, et al.
Published: (2023)
PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution
by: Liu, Yong, et al.
Published: (2024)
by: Liu, Yong, et al.
Published: (2024)
HairFastGAN: Realistic and Robust Hair Transfer with a Fast Encoder-Based Approach
by: Nikolaev, Maxim, et al.
Published: (2024)
by: Nikolaev, Maxim, et al.
Published: (2024)
Fast Diffeomorphic Image Registration using Patch based Fully Convolutional Networks
by: Wu, Jiong, et al.
Published: (2024)
by: Wu, Jiong, et al.
Published: (2024)
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
by: Wu, Yuqun, et al.
Published: (2025)
by: Wu, Yuqun, et al.
Published: (2025)
On Inductive Biases That Enable Generalization of Diffusion Transformers
by: An, Jie, et al.
Published: (2024)
by: An, Jie, et al.
Published: (2024)
Distilled Pooling Transformer Encoder for Efficient Realistic Image Dehazing
by: Tran, Le-Anh, et al.
Published: (2024)
by: Tran, Le-Anh, et al.
Published: (2024)
SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
by: Zhang, Yuxuan, et al.
Published: (2023)
by: Zhang, Yuxuan, et al.
Published: (2023)
Patch-Based Stochastic Attention for Image Editing
by: Cherel, Nicolas, et al.
Published: (2022)
by: Cherel, Nicolas, et al.
Published: (2022)
Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
by: Mehrotra, Ayushi, et al.
Published: (2025)
by: Mehrotra, Ayushi, et al.
Published: (2025)
Fast Encoder-Based 3D from Casual Videos via Point Track Processing
by: Kasten, Yoni, et al.
Published: (2024)
by: Kasten, Yoni, et al.
Published: (2024)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
Putting the Object Back into Video Object Segmentation
by: Cheng, Ho Kei, et al.
Published: (2023)
by: Cheng, Ho Kei, et al.
Published: (2023)
Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective
by: Zhao, Xiaoming, et al.
Published: (2025)
by: Zhao, Xiaoming, et al.
Published: (2025)
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)
by: Cui, Yajie, et al.
Published: (2023)
Similar Items
-
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026) -
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024) -
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025) -
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024) -
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)