Learning through Creation: A Hash-Free Framework for On-the-Fly Category Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Bohan, Tang, Weidong, Chi, Zhixiang, Jin, Yi, Li, Zhenbo, Wang, Yang, Wu, Yanan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BID: Boundary-Interior Decoding for Unsupervised Temporal Action Localization Pre-Trainin
by: Fang, Qihang, et al.
Published: (2024)
by: Fang, Qihang, et al.
Published: (2024)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)
by: Zinnen, Mathias, et al.
Published: (2025)
Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
by: Çetiner, Kemal Alperen, et al.
Published: (2026)
Detecting 3D Line Segments for 6DoF Pose Estimation with Limited Data
by: Mok, Matej, et al.
Published: (2026)
by: Mok, Matej, et al.
Published: (2026)
GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation
by: Catalini, Riccardo, et al.
Published: (2026)
by: Catalini, Riccardo, et al.
Published: (2026)
Diffusion Features for Zero-Shot 6DoF Object Pose Estimation
by: Von Gimborn, Bernd, et al.
Published: (2024)
by: Von Gimborn, Bernd, et al.
Published: (2024)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
by: Su, Yuetong, et al.
Published: (2025)
by: Su, Yuetong, et al.
Published: (2025)
ReFlow6D: Refraction-Guided Transparent Object 6D Pose Estimation via Intermediate Representation Learning
by: Gupta, Hrishikesh, et al.
Published: (2024)
by: Gupta, Hrishikesh, et al.
Published: (2024)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
by: Jin, Xiaofeng, et al.
Published: (2025)
by: Jin, Xiaofeng, et al.
Published: (2025)
RailSafeNet: Visual Scene Understanding for Tram Safety
by: Valach, Ondřej, et al.
Published: (2025)
by: Valach, Ondřej, et al.
Published: (2025)
Event-ECC: Asynchronous Tracking of Events with Continuous Optimization
by: Zafeiri, Maria, et al.
Published: (2024)
by: Zafeiri, Maria, et al.
Published: (2024)
VDPP: Video Depth Post-Processing for Speed and Scalability
by: Yoon, Daewon, et al.
Published: (2026)
by: Yoon, Daewon, et al.
Published: (2026)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
by: Chen, Jingkun, et al.
Published: (2025)
by: Chen, Jingkun, et al.
Published: (2025)
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
by: Hu, Yangliu, et al.
Published: (2025)
by: Hu, Yangliu, et al.
Published: (2025)
N-DriverMotion: Driver motion learning and prediction using an event-based camera and directly trained spiking neural networks on Loihi 2
by: Chung, Hyo Jong, et al.
Published: (2024)
by: Chung, Hyo Jong, et al.
Published: (2024)
The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
by: Ömercikoğlu, Ahmet Can, et al.
Published: (2025)
by: Ömercikoğlu, Ahmet Can, et al.
Published: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
by: Mohammad, Noor Islam S., et al.
Published: (2025)
by: Mohammad, Noor Islam S., et al.
Published: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
by: Mohammad, Noor Islam S.
Published: (2025)
by: Mohammad, Noor Islam S.
Published: (2025)
Improving Object Detection for Time-Lapse Imagery Using Temporal Features in Wildlife Monitoring
by: Jenkins, Marcus, et al.
Published: (2024)
by: Jenkins, Marcus, et al.
Published: (2024)
Decoder Generates Manufacturable Structures: A Framework for 3D-Printable Object Synthesis
by: Kumar, Abhishek
Published: (2026)
by: Kumar, Abhishek
Published: (2026)
TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery
by: Wu, Yanan, et al.
Published: (2026)
by: Wu, Yanan, et al.
Published: (2026)
Reference-based Category Discovery: Unsupervised Object Detection with Category Awareness
by: Li, Yichen, et al.
Published: (2026)
by: Li, Yichen, et al.
Published: (2026)
Unlocking UML Class Diagram Understanding in Vision Language Models
by: Naboichenko, Artem, et al.
Published: (2026)
by: Naboichenko, Artem, et al.
Published: (2026)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
by: Zhang, Junbin, et al.
Published: (2022)
by: Zhang, Junbin, et al.
Published: (2022)
Corn Ear Detection and Orientation Estimation Using Deep Learning
by: Sprague, Nathan, et al.
Published: (2024)
by: Sprague, Nathan, et al.
Published: (2024)
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
by: Ko, Hyun-kyu, et al.
Published: (2026)
by: Ko, Hyun-kyu, et al.
Published: (2026)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
by: Bruno, Pierangela, et al.
Published: (2024)
by: Bruno, Pierangela, et al.
Published: (2024)
Camera Pose Revisited
by: Skarbek, Władysław, et al.
Published: (2026)
by: Skarbek, Władysław, et al.
Published: (2026)
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
by: Lu, Hongyu, et al.
Published: (2026)
by: Lu, Hongyu, et al.
Published: (2026)
AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics
by: Tencent HY Team
Published: (2026)
by: Tencent HY Team
Published: (2026)
NeuroWrite: Predictive Handwritten Digit Classification using Deep Neural Networks
by: Asish, Kottakota, et al.
Published: (2023)
by: Asish, Kottakota, et al.
Published: (2023)
Archival Faces: Detection of Faces in Digitized Historical Documents
by: Vaško, Marek, et al.
Published: (2025)
by: Vaško, Marek, et al.
Published: (2025)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
by: Jung, Seoik, et al.
Published: (2025)
by: Jung, Seoik, et al.
Published: (2025)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
Experimental Evaluation of Road-Crossing Decisions by Autonomous Wheelchairs against Environmental Factors
by: Corradini, Franca, et al.
Published: (2024)
by: Corradini, Franca, et al.
Published: (2024)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
A large-scale, physically-based synthetic dataset for satellite pose estimation
by: Velkei, Szabolcs, et al.
Published: (2025)
by: Velkei, Szabolcs, et al.
Published: (2025)
Similar Items
-
BID: Boundary-Interior Decoding for Unsupervised Temporal Action Localization Pre-Trainin
by: Fang, Qihang, et al.
Published: (2024) -
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025) -
Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
by: Çetiner, Kemal Alperen, et al.
Published: (2026) -
Detecting 3D Line Segments for 6DoF Pose Estimation with Limited Data
by: Mok, Matej, et al.
Published: (2026) -
GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation
by: Catalini, Riccardo, et al.
Published: (2026)