VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ahmad, Niaz, Lee, Youngmoon, Wang, Guanghui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Keypoints as Dynamic Centroids for Unified Human Pose and Segmentation
di: Ahmad, Niaz, et al.
Pubblicazione: (2025)
di: Ahmad, Niaz, et al.
Pubblicazione: (2025)
AT-SNN: Adaptive Tokens for Vision Transformer on Spiking Neural Network
di: Kang, Donghwa, et al.
Pubblicazione: (2024)
di: Kang, Donghwa, et al.
Pubblicazione: (2024)
NCDD: Nearest Centroid Distance Deficit for Out-Of-Distribution Detection in Gastrointestinal Vision
di: Pokhrel, Sandesh, et al.
Pubblicazione: (2024)
di: Pokhrel, Sandesh, et al.
Pubblicazione: (2024)
Steerable Visual Representations
di: Ruthardt, Jona, et al.
Pubblicazione: (2026)
di: Ruthardt, Jona, et al.
Pubblicazione: (2026)
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
di: Zhang, Xianqi
Pubblicazione: (2026)
di: Zhang, Xianqi
Pubblicazione: (2026)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
di: Gandhi, Sanket, et al.
Pubblicazione: (2024)
di: Gandhi, Sanket, et al.
Pubblicazione: (2024)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
di: Lee, Eunsoo, et al.
Pubblicazione: (2026)
di: Lee, Eunsoo, et al.
Pubblicazione: (2026)
Representation Learning with Adaptive Superpixel Coding
di: Khalil, Mahmoud, et al.
Pubblicazione: (2025)
di: Khalil, Mahmoud, et al.
Pubblicazione: (2025)
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
di: Paruchuri, Akshay, et al.
Pubblicazione: (2026)
di: Paruchuri, Akshay, et al.
Pubblicazione: (2026)
A Survey on Mamba Architecture for Vision Applications
di: Ibrahim, Fady, et al.
Pubblicazione: (2025)
di: Ibrahim, Fady, et al.
Pubblicazione: (2025)
Beyond ZOH: Advanced Discretization Strategies for Vision Mamba
di: Ibrahim, Fady, et al.
Pubblicazione: (2026)
di: Ibrahim, Fady, et al.
Pubblicazione: (2026)
Comparing Computational Pathology Foundation Models using Representational Similarity Analysis
di: Mishra, Vaibhav, et al.
Pubblicazione: (2025)
di: Mishra, Vaibhav, et al.
Pubblicazione: (2025)
Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
di: Wang, Wentao, et al.
Pubblicazione: (2025)
di: Wang, Wentao, et al.
Pubblicazione: (2025)
Mitigating Bias in Facial Recognition Systems: Centroid Fairness Loss Optimization
di: Conti, Jean-Rémy, et al.
Pubblicazione: (2025)
di: Conti, Jean-Rémy, et al.
Pubblicazione: (2025)
Aligning Machine and Human Visual Representations across Abstraction Levels
di: Muttenthaler, Lukas, et al.
Pubblicazione: (2024)
di: Muttenthaler, Lukas, et al.
Pubblicazione: (2024)
Cluster Contrast for Unsupervised Visual Representation Learning
di: Giakoumoglou, Nikolaos, et al.
Pubblicazione: (2025)
di: Giakoumoglou, Nikolaos, et al.
Pubblicazione: (2025)
Privacy-Concealing Cooperative Perception for BEV Scene Segmentation
di: Wang, Song, et al.
Pubblicazione: (2026)
di: Wang, Song, et al.
Pubblicazione: (2026)
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
di: Lin, Han, et al.
Pubblicazione: (2026)
di: Lin, Han, et al.
Pubblicazione: (2026)
GTMA: Dynamic Representation Optimization for OOD Vision-Language Models
di: Zhang, Jensen, et al.
Pubblicazione: (2025)
di: Zhang, Jensen, et al.
Pubblicazione: (2025)
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
di: Zheng, Sixiao, et al.
Pubblicazione: (2024)
di: Zheng, Sixiao, et al.
Pubblicazione: (2024)
Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
di: Ahmad, Kaleem
Pubblicazione: (2025)
di: Ahmad, Kaleem
Pubblicazione: (2025)
Human-Like Coarse Object Representations in Vision Models
di: Gizdov, Andrey, et al.
Pubblicazione: (2026)
di: Gizdov, Andrey, et al.
Pubblicazione: (2026)
Feature Fusion for Human Activity Recognition using Parameter-Optimized Multi-Stage Graph Convolutional Network and Transformer Models
di: Belal, Mohammad, et al.
Pubblicazione: (2024)
di: Belal, Mohammad, et al.
Pubblicazione: (2024)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2023)
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2023)
Target-Dependent Multimodal Sentiment Analysis Via Employing Visual-to Emotional-Caption Translation Network using Visual-Caption Pairs
di: Pandey, Ananya, et al.
Pubblicazione: (2024)
di: Pandey, Ananya, et al.
Pubblicazione: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
di: Li, Siyuan, et al.
Pubblicazione: (2025)
di: Li, Siyuan, et al.
Pubblicazione: (2025)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video Reconstruction
di: Wang, Haonan, et al.
Pubblicazione: (2025)
di: Wang, Haonan, et al.
Pubblicazione: (2025)
LLMs in Political Science: Heralding a New Era of Visual Analysis
di: Wang, Yu
Pubblicazione: (2024)
di: Wang, Yu
Pubblicazione: (2024)
PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
di: Ibrahem, Hatem, et al.
Pubblicazione: (2025)
di: Ibrahem, Hatem, et al.
Pubblicazione: (2025)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
di: Wang, Haoming, et al.
Pubblicazione: (2026)
di: Wang, Haoming, et al.
Pubblicazione: (2026)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
di: Li, Hengzhuang, et al.
Pubblicazione: (2025)
di: Li, Hengzhuang, et al.
Pubblicazione: (2025)
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
di: Xie, Roy, et al.
Pubblicazione: (2026)
di: Xie, Roy, et al.
Pubblicazione: (2026)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
di: Miller, Luke James, et al.
Pubblicazione: (2026)
di: Miller, Luke James, et al.
Pubblicazione: (2026)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models
di: Kim, Jeongho, et al.
Pubblicazione: (2024)
di: Kim, Jeongho, et al.
Pubblicazione: (2024)
MedEyes: Learning Dynamic Visual Focus for Medical Progressive Diagnosis
di: Zhu, Chunzheng, et al.
Pubblicazione: (2025)
di: Zhu, Chunzheng, et al.
Pubblicazione: (2025)
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
di: Li, Ruihang, et al.
Pubblicazione: (2026)
di: Li, Ruihang, et al.
Pubblicazione: (2026)
ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding
di: Sun, Dongwei, et al.
Pubblicazione: (2026)
di: Sun, Dongwei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Keypoints as Dynamic Centroids for Unified Human Pose and Segmentation
di: Ahmad, Niaz, et al.
Pubblicazione: (2025) -
AT-SNN: Adaptive Tokens for Vision Transformer on Spiking Neural Network
di: Kang, Donghwa, et al.
Pubblicazione: (2024) -
NCDD: Nearest Centroid Distance Deficit for Out-Of-Distribution Detection in Gastrointestinal Vision
di: Pokhrel, Sandesh, et al.
Pubblicazione: (2024) -
Steerable Visual Representations
di: Ruthardt, Jona, et al.
Pubblicazione: (2026) -
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
di: Zhang, Xianqi
Pubblicazione: (2026)