DisenQ: Disentangling Q-Former for Activity-Biometrics
Fuente:
arXiv
Salvato in:
| Autori principali: | Azad, Shehreen, Rawat, Yogesh S |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
di: Azad, Shehreen, et al.
Pubblicazione: (2025)
di: Azad, Shehreen, et al.
Pubblicazione: (2025)
Activity-Biometrics: Person Identification from Daily Activities
di: Azad, Shehreen, et al.
Pubblicazione: (2024)
di: Azad, Shehreen, et al.
Pubblicazione: (2024)
StreamReady: Learning What to Answer and When in Long Streaming Videos
di: Azad, Shehreen, et al.
Pubblicazione: (2026)
di: Azad, Shehreen, et al.
Pubblicazione: (2026)
Understanding Depth and Height Perception in Large Visual-Language Models
di: Azad, Shehreen, et al.
Pubblicazione: (2024)
di: Azad, Shehreen, et al.
Pubblicazione: (2024)
Probing Conceptual Understanding of Large Visual-Language Models
di: Schiappa, Madeline, et al.
Pubblicazione: (2023)
di: Schiappa, Madeline, et al.
Pubblicazione: (2023)
Robustness Analysis on Foundational Segmentation Models
di: Schiappa, Madeline Chantry, et al.
Pubblicazione: (2023)
di: Schiappa, Madeline Chantry, et al.
Pubblicazione: (2023)
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
di: Mitra, Sirshapan, et al.
Pubblicazione: (2025)
di: Mitra, Sirshapan, et al.
Pubblicazione: (2025)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
di: Liang, Xin, et al.
Pubblicazione: (2025)
di: Liang, Xin, et al.
Pubblicazione: (2025)
Navigating Hallucinations for Reasoning of Unintentional Activities
di: Grover, Shresth, et al.
Pubblicazione: (2024)
di: Grover, Shresth, et al.
Pubblicazione: (2024)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
di: Chen, Hong, et al.
Pubblicazione: (2023)
di: Chen, Hong, et al.
Pubblicazione: (2023)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
di: Chen, Hong, et al.
Pubblicazione: (2024)
di: Chen, Hong, et al.
Pubblicazione: (2024)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
di: Mitra, Sirshapan, et al.
Pubblicazione: (2026)
di: Mitra, Sirshapan, et al.
Pubblicazione: (2026)
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
di: Dalmonte, Francesco, et al.
Pubblicazione: (2025)
di: Dalmonte, Francesco, et al.
Pubblicazione: (2025)
CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities
di: Liu, Bo, et al.
Pubblicazione: (2026)
di: Liu, Bo, et al.
Pubblicazione: (2026)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
di: Garg, Aaryan, et al.
Pubblicazione: (2025)
di: Garg, Aaryan, et al.
Pubblicazione: (2025)
Scaling Open-Vocabulary Action Detection
di: Sia, Zhen Hao, et al.
Pubblicazione: (2025)
di: Sia, Zhen Hao, et al.
Pubblicazione: (2025)
MolVision: Molecular Property Prediction with Vision Language Models
di: Adak, Deepan, et al.
Pubblicazione: (2025)
di: Adak, Deepan, et al.
Pubblicazione: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
di: Abdullah, Raiyaan, et al.
Pubblicazione: (2025)
di: Abdullah, Raiyaan, et al.
Pubblicazione: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
di: Kumar, Akash, et al.
Pubblicazione: (2025)
di: Kumar, Akash, et al.
Pubblicazione: (2025)
Stable Mean Teacher for Semi-supervised Video Action Detection
di: Kumar, Akash, et al.
Pubblicazione: (2024)
di: Kumar, Akash, et al.
Pubblicazione: (2024)
Asynchronous Perception Machine For Efficient Test-Time-Training
di: Modi, Rajat, et al.
Pubblicazione: (2024)
di: Modi, Rajat, et al.
Pubblicazione: (2024)
Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance
di: Rawat, Dhruvraj Singh, et al.
Pubblicazione: (2025)
di: Rawat, Dhruvraj Singh, et al.
Pubblicazione: (2025)
DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision
di: Liao, Jiashu, et al.
Pubblicazione: (2025)
di: Liao, Jiashu, et al.
Pubblicazione: (2025)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
di: Kumar, Akash, et al.
Pubblicazione: (2025)
di: Kumar, Akash, et al.
Pubblicazione: (2025)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
di: Pathak, Priyank, et al.
Pubblicazione: (2025)
OmViD: Omni-supervised active learning for video action detection
di: Rana, Aayush, et al.
Pubblicazione: (2025)
di: Rana, Aayush, et al.
Pubblicazione: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
di: Ahmad, Shahzad, et al.
Pubblicazione: (2023)
di: Ahmad, Shahzad, et al.
Pubblicazione: (2023)
MolSight: Molecular Property Prediction with Images
di: Baranwal, Aaditya, et al.
Pubblicazione: (2026)
di: Baranwal, Aaditya, et al.
Pubblicazione: (2026)
VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models
di: Chen, Hong, et al.
Pubblicazione: (2023)
di: Chen, Hong, et al.
Pubblicazione: (2023)
Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images
di: Jha, Abhishek, et al.
Pubblicazione: (2024)
di: Jha, Abhishek, et al.
Pubblicazione: (2024)
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
di: Grover, Shresth, et al.
Pubblicazione: (2025)
di: Grover, Shresth, et al.
Pubblicazione: (2025)
RobustGait: Robustness Analysis for Appearance Based Gait Recognition
di: Sayera, Reeshoon, et al.
Pubblicazione: (2025)
di: Sayera, Reeshoon, et al.
Pubblicazione: (2025)
Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
di: Wang, Zengyan, et al.
Pubblicazione: (2026)
di: Wang, Zengyan, et al.
Pubblicazione: (2026)
Task-adaptive Q-Face
di: Sun, Haomiao, et al.
Pubblicazione: (2024)
di: Sun, Haomiao, et al.
Pubblicazione: (2024)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
di: Modi, Rajat, et al.
Pubblicazione: (2024)
di: Modi, Rajat, et al.
Pubblicazione: (2024)
Foundation Models for Video Understanding: A Survey
di: Madan, Neelu, et al.
Pubblicazione: (2024)
di: Madan, Neelu, et al.
Pubblicazione: (2024)
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
di: Xu, Peiran, et al.
Pubblicazione: (2025)
di: Xu, Peiran, et al.
Pubblicazione: (2025)
Re:Verse -- Can Your VLM Read a Manga?
di: Baranwal, Aaditya, et al.
Pubblicazione: (2025)
di: Baranwal, Aaditya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
di: Azad, Shehreen, et al.
Pubblicazione: (2025) -
Activity-Biometrics: Person Identification from Daily Activities
di: Azad, Shehreen, et al.
Pubblicazione: (2024) -
StreamReady: Learning What to Answer and When in Long Streaming Videos
di: Azad, Shehreen, et al.
Pubblicazione: (2026) -
Understanding Depth and Height Perception in Large Visual-Language Models
di: Azad, Shehreen, et al.
Pubblicazione: (2024) -
Probing Conceptual Understanding of Large Visual-Language Models
di: Schiappa, Madeline, et al.
Pubblicazione: (2023)