Enhancing Visual Classification using Comparative Descriptors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Hankyeol, Seo, Gawon, Choi, Wonseok, Jung, Geunyoung, Song, Kyungwoo, Jung, Jiyoung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
von: Jung, Geunyoung, et al.
Veröffentlicht: (2026)
von: Jung, Geunyoung, et al.
Veröffentlicht: (2026)
Robust Adaptation of Foundation Models with Black-Box Visual Prompting
von: Oh, Changdae, et al.
Veröffentlicht: (2024)
von: Oh, Changdae, et al.
Veröffentlicht: (2024)
APC: Transferable and Efficient Adversarial Point Counterattack for Robust 3D Point Cloud Recognition
von: Jung, Geunyoung, et al.
Veröffentlicht: (2026)
von: Jung, Geunyoung, et al.
Veröffentlicht: (2026)
Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments
von: Kim, Jaehyun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehyun, et al.
Veröffentlicht: (2025)
REVIVE 3D: Refinement via Encoded Voluminous Inflated prior for Volume Enhancement
von: Lee, Hankyeol, et al.
Veröffentlicht: (2026)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2026)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2025)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2025)
VizECGNet: Visual ECG Image Network for Cardiovascular Diseases Classification with Multi-Modal Training and Knowledge Distillation
von: Nam, Ju-Hyeon, et al.
Veröffentlicht: (2024)
von: Nam, Ju-Hyeon, et al.
Veröffentlicht: (2024)
Online Continuous Generalized Category Discovery
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
Visual Detector Compression via Location-Aware Discriminant Analysis
von: Lan, Qizhen, et al.
Veröffentlicht: (2025)
von: Lan, Qizhen, et al.
Veröffentlicht: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts
von: Nguyen, Khanh Binh, et al.
Veröffentlicht: (2026)
von: Nguyen, Khanh Binh, et al.
Veröffentlicht: (2026)
Performance Evaluation of 3D Keypoint Detectors and Descriptors on Coloured Point Clouds in Subsea Environments
von: Jung, Kyungmin, et al.
Veröffentlicht: (2022)
von: Jung, Kyungmin, et al.
Veröffentlicht: (2022)
Amnesia as a Catalyst for Enhancing Black Box Pixel Attacks in Image Classification and Object Detection
von: Song, Dongsu, et al.
Veröffentlicht: (2025)
von: Song, Dongsu, et al.
Veröffentlicht: (2025)
Automated Classification of Cell Shapes: A Comparative Evaluation of Shape Descriptors
von: Vadori, Valentina, et al.
Veröffentlicht: (2024)
von: Vadori, Valentina, et al.
Veröffentlicht: (2024)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks
von: Jung, Aecheon, et al.
Veröffentlicht: (2025)
von: Jung, Aecheon, et al.
Veröffentlicht: (2025)
UAVTwin: Neural Digital Twins for UAVs using Gaussian Splatting
von: Choi, Jaehoon, et al.
Veröffentlicht: (2025)
von: Choi, Jaehoon, et al.
Veröffentlicht: (2025)
IM360: Large-scale Indoor Mapping with 360 Cameras
von: Jung, Dongki, et al.
Veröffentlicht: (2025)
von: Jung, Dongki, et al.
Veröffentlicht: (2025)
RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
von: Jung, Dongki, et al.
Veröffentlicht: (2025)
von: Jung, Dongki, et al.
Veröffentlicht: (2025)
OASIS: Online Sample Selection for Continual Visual Instruction Tuning
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
A Comparative Analysis of Image Descriptors for Histopathological Classification of Gastric Cancer
von: Usai, Marco, et al.
Veröffentlicht: (2025)
von: Usai, Marco, et al.
Veröffentlicht: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
von: Ko, Jungmin, et al.
Veröffentlicht: (2026)
von: Ko, Jungmin, et al.
Veröffentlicht: (2026)
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
von: Liu, Chengxin, et al.
Veröffentlicht: (2026)
von: Liu, Chengxin, et al.
Veröffentlicht: (2026)
A Text-Guided Vision Model for Enhanced Recognition of Small Instances
von: Jung, Hyun-Ki
Veröffentlicht: (2026)
von: Jung, Hyun-Ki
Veröffentlicht: (2026)
Exploring Image Representation with Decoupled Classical Visual Descriptors
von: Qu, Chenyuan, et al.
Veröffentlicht: (2025)
von: Qu, Chenyuan, et al.
Veröffentlicht: (2025)
GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results
von: Won, Jiye, et al.
Veröffentlicht: (2026)
von: Won, Jiye, et al.
Veröffentlicht: (2026)
Integrating Query-aware Segmentation and Cross-Attention for Robust VQA
von: Choi, Wonjun, et al.
Veröffentlicht: (2024)
von: Choi, Wonjun, et al.
Veröffentlicht: (2024)
PBVS 2024 Solution: Self-Supervised Learning and Sampling Strategies for SAR Classification in Extreme Long-Tail Distribution
von: Kim, Yuhyun, et al.
Veröffentlicht: (2024)
von: Kim, Yuhyun, et al.
Veröffentlicht: (2024)
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
von: Jung, Seongmin, et al.
Veröffentlicht: (2025)
von: Jung, Seongmin, et al.
Veröffentlicht: (2025)
DALDA: Data Augmentation Leveraging Diffusion Model and LLM with Adaptive Guidance Scaling
von: Jung, Kyuheon, et al.
Veröffentlicht: (2024)
von: Jung, Kyuheon, et al.
Veröffentlicht: (2024)
CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image
von: Roh, Wonseok, et al.
Veröffentlicht: (2024)
von: Roh, Wonseok, et al.
Veröffentlicht: (2024)
WWW: Where, Which and Whatever Enhancing Interpretability in Multimodal Deepfake Detection
von: Jung, Juho, et al.
Veröffentlicht: (2024)
von: Jung, Juho, et al.
Veröffentlicht: (2024)
Spectrum Translation for Refinement of Image Generation (STIG) Based on Contrastive Learning and Spectral Filter Profile
von: Lee, Seokjun, et al.
Veröffentlicht: (2024)
von: Lee, Seokjun, et al.
Veröffentlicht: (2024)
Descriptive Image-Text Matching with Graded Contextual Similarity
von: Jang, Jinhyun, et al.
Veröffentlicht: (2025)
von: Jang, Jinhyun, et al.
Veröffentlicht: (2025)
From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation
von: Çetinkaya, Evren, et al.
Veröffentlicht: (2026)
von: Çetinkaya, Evren, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025) -
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
von: Jung, Geunyoung, et al.
Veröffentlicht: (2026) -
Robust Adaptation of Foundation Models with Black-Box Visual Prompting
von: Oh, Changdae, et al.
Veröffentlicht: (2024) -
APC: Transferable and Efficient Adversarial Point Counterattack for Robust 3D Point Cloud Recognition
von: Jung, Geunyoung, et al.
Veröffentlicht: (2026) -
Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments
von: Kim, Jaehyun, et al.
Veröffentlicht: (2025)