Less yet robust: crucial region selection for scene recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jianqi, Wang, Mengxuan, Wang, Jingyao, Si, Lingyu, Zheng, Changwen, Xu, Fanjiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
by: Song, Fei, et al.
Published: (2025)
by: Song, Fei, et al.
Published: (2025)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
by: Yang, Sheng, et al.
Published: (2025)
by: Yang, Sheng, et al.
Published: (2025)
Learning Novel View Synthesis from Heterogeneous Low-light Captures
by: Zheng, Quan, et al.
Published: (2024)
by: Zheng, Quan, et al.
Published: (2024)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
Enhancing Large Language Models for Time-Series Forecasting via Vector-Injected In-Context Learning
by: Zhang, Jianqi, et al.
Published: (2026)
by: Zhang, Jianqi, et al.
Published: (2026)
Configural processing as an optimized strategy for robust object recognition in neural networks
by: Jang, Hojin, et al.
Published: (2024)
by: Jang, Hojin, et al.
Published: (2024)
Adversarial Testing for Visual Grounding via Image-Aware Property Reduction
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
by: Son, Dongwon, et al.
Published: (2024)
by: Son, Dongwon, et al.
Published: (2024)
Skeleton-Based Action Recognition with Spatial-Structural Graph Convolution
by: Wang, Jingyao, et al.
Published: (2024)
by: Wang, Jingyao, et al.
Published: (2024)
Generating metamers of human scene understanding
by: Raina, Ritik, et al.
Published: (2026)
by: Raina, Ritik, et al.
Published: (2026)
Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration
by: Xu, Jinglin, et al.
Published: (2026)
by: Xu, Jinglin, et al.
Published: (2026)
From Pixels to Predicates Structuring urban perception with scene graphs
by: Liu, Yunlong, et al.
Published: (2025)
by: Liu, Yunlong, et al.
Published: (2025)
Image-based Freeform Handwriting Authentication with Energy-oriented Self-Supervised Learning
by: Wang, Jingyao, et al.
Published: (2024)
by: Wang, Jingyao, et al.
Published: (2024)
Swift4D:Adaptive divide-and-conquer Gaussian Splatting for compact and efficient reconstruction of dynamic scene
by: Wu, Jiahao, et al.
Published: (2025)
by: Wu, Jiahao, et al.
Published: (2025)
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025)
by: Luo, Jiamin, et al.
Published: (2025)
A Physical Model-Guided Framework for Underwater Image Enhancement and Depth Estimation
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
by: Jiang, Ruixiang, et al.
Published: (2025)
by: Jiang, Ruixiang, et al.
Published: (2025)
Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism
by: Lv, Haotian, et al.
Published: (2025)
by: Lv, Haotian, et al.
Published: (2025)
SEDEG:Sequential Enhancement of Decoder and Encoder's Generality for Class Incremental Learning with Small Memory
by: Chen, Hongyang, et al.
Published: (2025)
by: Chen, Hongyang, et al.
Published: (2025)
Hybrid guided variational autoencoder for visual place recognition
by: Wang, Ni, et al.
Published: (2026)
by: Wang, Ni, et al.
Published: (2026)
BCFPL: Binary classification ConvNet based Fast Parking space recognition with Low resolution image
by: Zhang, Shuo, et al.
Published: (2024)
by: Zhang, Shuo, et al.
Published: (2024)
UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment
by: Wang, Zheng, et al.
Published: (2026)
by: Wang, Zheng, et al.
Published: (2026)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
by: Ma, Junxiao, et al.
Published: (2025)
by: Ma, Junxiao, et al.
Published: (2025)
A comprehensive survey of oracle character recognition: challenges, benchmarks, and beyond
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
by: Sanchez, Cristhian, et al.
Published: (2024)
by: Sanchez, Cristhian, et al.
Published: (2024)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
by: Song, Fei, et al.
Published: (2025)
by: Song, Fei, et al.
Published: (2025)
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
by: Lian, Jingchun, et al.
Published: (2024)
by: Lian, Jingchun, et al.
Published: (2024)
Self-driving cars: Are we there yet?
by: Atasever, Merve, et al.
Published: (2025)
by: Atasever, Merve, et al.
Published: (2025)
LPT: Less-overfitting Prompt Tuning for Vision-Language Model
by: Ding, Chenhao, et al.
Published: (2024)
by: Ding, Chenhao, et al.
Published: (2024)
On Semiotic-Grounded Interpretive Evaluation of Generative Art
by: Jiang, Ruixiang, et al.
Published: (2026)
by: Jiang, Ruixiang, et al.
Published: (2026)
BadHMP: Backdoor Attack against Human Motion Prediction
by: Xu, Chaohui, et al.
Published: (2024)
by: Xu, Chaohui, et al.
Published: (2024)
Continual-learning-based framework for structural damage recognition
by: Shu, Jiangpeng, et al.
Published: (2024)
by: Shu, Jiangpeng, et al.
Published: (2024)
Similar Items
-
Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
by: Wang, Jingyao, et al.
Published: (2025) -
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024) -
Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation
by: Wang, Jingyao, et al.
Published: (2025) -
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
by: Song, Fei, et al.
Published: (2025) -
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
by: Yang, Sheng, et al.
Published: (2025)