Visual Accommodation: Rethinking Image Scale as a Learnable Variable for Object Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Seo, Daeun, Yang, Hoeseok, Park, Sihyeong, Kim, Hyungshin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DyRA: Portable Dynamic Resolution Adjustment Network for Existing Detectors
di: Seo, Daeun, et al.
Pubblicazione: (2023)
di: Seo, Daeun, et al.
Pubblicazione: (2023)
Mixed Non-linear Quantization for Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2024)
di: Kim, Gihwan, et al.
Pubblicazione: (2024)
Simplifying Two-Stage Detectors for On-Device Inference in Remote Sensing
di: Kang, Jaemin, et al.
Pubblicazione: (2024)
di: Kang, Jaemin, et al.
Pubblicazione: (2024)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2025)
di: Kim, Gihwan, et al.
Pubblicazione: (2025)
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
di: Lee, Yujin, et al.
Pubblicazione: (2024)
di: Lee, Yujin, et al.
Pubblicazione: (2024)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
di: Han, Donghoon, et al.
Pubblicazione: (2026)
di: Han, Donghoon, et al.
Pubblicazione: (2026)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
di: Lee, Jemin, et al.
Pubblicazione: (2023)
di: Lee, Jemin, et al.
Pubblicazione: (2023)
Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images
di: Li, Jinfu, et al.
Pubblicazione: (2025)
di: Li, Jinfu, et al.
Pubblicazione: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
VizECGNet: Visual ECG Image Network for Cardiovascular Diseases Classification with Multi-Modal Training and Knowledge Distillation
di: Nam, Ju-Hyeon, et al.
Pubblicazione: (2024)
di: Nam, Ju-Hyeon, et al.
Pubblicazione: (2024)
BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers
di: Seo, Jungwook, et al.
Pubblicazione: (2026)
di: Seo, Jungwook, et al.
Pubblicazione: (2026)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
di: Park, Joonhyung, et al.
Pubblicazione: (2025)
di: Park, Joonhyung, et al.
Pubblicazione: (2025)
Object Aware Egocentric Online Action Detection
di: An, Joungbin, et al.
Pubblicazione: (2024)
di: An, Joungbin, et al.
Pubblicazione: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
di: Lee, Daeun, et al.
Pubblicazione: (2026)
di: Lee, Daeun, et al.
Pubblicazione: (2026)
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
di: Park, Jiho, et al.
Pubblicazione: (2025)
di: Park, Jiho, et al.
Pubblicazione: (2025)
Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM
di: Park, Hyobin, et al.
Pubblicazione: (2026)
di: Park, Hyobin, et al.
Pubblicazione: (2026)
Investigating Long-term Training for Remote Sensing Object Detection
di: Park, JongHyun, et al.
Pubblicazione: (2024)
di: Park, JongHyun, et al.
Pubblicazione: (2024)
ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects
di: Lee, Woojin, et al.
Pubblicazione: (2025)
di: Lee, Woojin, et al.
Pubblicazione: (2025)
Enhancing Few-Shot Image Classification through Learnable Multi-Scale Embedding and Attention Mechanisms
di: Askari, Fatemeh, et al.
Pubblicazione: (2024)
di: Askari, Fatemeh, et al.
Pubblicazione: (2024)
Uncertainty Quantification in Detection Transformers: Object-Level Calibration and Image-Level Reliability
di: Park, Young-Jin, et al.
Pubblicazione: (2024)
di: Park, Young-Jin, et al.
Pubblicazione: (2024)
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
LEAP:D -- A Novel Prompt-based Approach for Domain-Generalized Aerial Object Detection
di: Park, Chanyeong, et al.
Pubblicazione: (2024)
di: Park, Chanyeong, et al.
Pubblicazione: (2024)
Rethinking Visual Information Processing in Multimodal LLMs
di: Kim, Dongwan, et al.
Pubblicazione: (2025)
di: Kim, Dongwan, et al.
Pubblicazione: (2025)
SFUOD: Source-Free Unknown Object Detection
di: Park, Keon-Hee, et al.
Pubblicazione: (2025)
di: Park, Keon-Hee, et al.
Pubblicazione: (2025)
Towards Predicting Temporal Changes in a Patient's Chest X-ray Images based on Electronic Health Records
di: Kyung, Daeun, et al.
Pubblicazione: (2024)
di: Kyung, Daeun, et al.
Pubblicazione: (2024)
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
di: Kim, Dong-Hee, et al.
Pubblicazione: (2026)
di: Kim, Dong-Hee, et al.
Pubblicazione: (2026)
Empirical Analysis of Anomaly Detection on Hyperspectral Imaging Using Dimension Reduction Methods
di: Kim, Dongeon, et al.
Pubblicazione: (2024)
di: Kim, Dongeon, et al.
Pubblicazione: (2024)
Implicit Deformable Medical Image Registration with Learnable Kernels
di: Fogarollo, Stefano, et al.
Pubblicazione: (2025)
di: Fogarollo, Stefano, et al.
Pubblicazione: (2025)
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
di: Park, Konyul, et al.
Pubblicazione: (2024)
di: Park, Konyul, et al.
Pubblicazione: (2024)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
di: An, Sojung, et al.
Pubblicazione: (2025)
di: An, Sojung, et al.
Pubblicazione: (2025)
Re-Scoring Using Image-Language Similarity for Few-Shot Object Detection
di: Jung, Min Jae, et al.
Pubblicazione: (2023)
di: Jung, Min Jae, et al.
Pubblicazione: (2023)
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
di: Park, Junsung, et al.
Pubblicazione: (2024)
di: Park, Junsung, et al.
Pubblicazione: (2024)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
di: Kim, Jiwan, et al.
Pubblicazione: (2025)
di: Kim, Jiwan, et al.
Pubblicazione: (2025)
Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
di: Vlachogiannis, Dimitrios N., et al.
Pubblicazione: (2025)
di: Vlachogiannis, Dimitrios N., et al.
Pubblicazione: (2025)
Efficient Learnable Collaborative Attention for Single Image Super-Resolution
di: Zheng, Yigang Zhao Chaowei, et al.
Pubblicazione: (2024)
di: Zheng, Yigang Zhao Chaowei, et al.
Pubblicazione: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)
di: Lee, Daeun, et al.
Pubblicazione: (2024)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
di: Luo, Yuanhao, et al.
Pubblicazione: (2026)
di: Luo, Yuanhao, et al.
Pubblicazione: (2026)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
di: Chi, Donghwan, et al.
Pubblicazione: (2025)
di: Chi, Donghwan, et al.
Pubblicazione: (2025)
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
di: Yu, Yi, et al.
Pubblicazione: (2025)
di: Yu, Yi, et al.
Pubblicazione: (2025)
An Efficient Aerial Image Detection with Variable Receptive Fields
di: Wenbin, Liu
Pubblicazione: (2025)
di: Wenbin, Liu
Pubblicazione: (2025)
Documenti analoghi
-
DyRA: Portable Dynamic Resolution Adjustment Network for Existing Detectors
di: Seo, Daeun, et al.
Pubblicazione: (2023) -
Mixed Non-linear Quantization for Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2024) -
Simplifying Two-Stage Detectors for On-Device Inference in Remote Sensing
di: Kang, Jaemin, et al.
Pubblicazione: (2024) -
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2025) -
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
di: Lee, Yujin, et al.
Pubblicazione: (2024)